Your Prompt Is Not a Security Boundary
Adapted from my course, Build Your Own Agentic Defense.
A model can request a tool, but the harness decides whether that request gets executed. When you connect an agent to real systems, you need to define its permissions at two levels: whether a tool may run, and, if it may, how it may be used.
First, you decide which tools are available and whether their use requires human approval. You might allow a candidate lookup to run automatically, require a person's approval before isolating a host, and leave evidence deletion unavailable. This establishes what kinds of actions the model can ask the harness to carry out and which ones need a person to authorize them.
For any tool you permit, you then define the limits on its arguments. A telemetry query might be allowed to run automatically, but only for the current case, within a maximum time window, and with a fixed limit on the number of results. The tool is available, yet a request that exceeds those limits must still be rejected or adjusted according to a defined rule. Together, these decisions form the harness's policy, and the harness must enforce both levels before it executes a request.
Rules in the Prompt
The system prompt is a useful place to explain how you want the model to behave. You might tell it to gather evidence before proposing a response, stay within the current case, or avoid destructive actions. Those instructions help the model choose an appropriate next step, but the code that executes a tool still needs to enforce the relevant limits.
For example, suppose you give the model a command-execution tool and put “never delete evidence” in the prompt. If the model nevertheless requests a deletion, perhaps after following misleading instructions in a file it read, the prompt does not prevent the command runner from executing it. Something outside the model must reject that operation or prevent the process from having permission to perform it.
The distinction matters when you test your harness. You want a prohibited request to fail even when the model produces it. The checks in the harness, together with the permissions enforced by the underlying system, are what make that possible. You can then use the prompt to explain those boundaries so the model is less likely to make requests that will be refused.
Whether a Tool May Be Called
Start by deciding which capabilities a particular part of your workflow needs and how each one may be used. For a security workflow, three categories are useful:
| policy | example capability |
|---|---|
| allow automatically | read local candidates |
| require human approval | isolate a host |
| not built into this harness | delete evidence |
Allow automatically
A tool can run automatically when its permitted effects are acceptable without a person reviewing every call. In a triage workflow, reading a candidate, retrieving its evidence, and looking up an asset are routine operations. You can let the model request these within the case and limits you have defined.
Read-only tools still need boundaries. A lookup can expose information the current user should not see, or consume excessive resources if it is allowed to return an entire database. When you mark a tool as automatic, you are approving a particular bounded operation: for example, a read from the current case with a fixed output limit. You are not granting unrestricted access to everything that operation could reach.
Require human approval
An action such as isolating a host or changing a firewall rule may interrupt someone else's work. You may want the model to identify when the action would help, while keeping the decision to execute it with a person. The harness can hold the request and show the proposed action, its target, the reason for it, and the supporting evidence.
Approval must come through a channel the model cannot use to approve itself. For example, your application can present an approval button to an authenticated analyst and record their decision against the exact tool and arguments. If the model later changes the target host, the earlier approval should not authorize that new request. This gives the reviewer a concrete action to assess and gives the harness a concrete action it is permitted to execute.
Not built into this harness
Some capabilities have no place in the workflow. A triage agent does not need to delete the evidence it is examining, so you can leave that operation out of its tool registry entirely. If the model invents a delete_evidence request, the harness rejects the unknown name; there is no registered implementation to dispatch it to.
You also need to consider indirect access. Leaving out a dedicated deletion tool would achieve little if the same model had an unrestricted shell that could delete the files instead. The capability is unavailable only when the other tools and their underlying permissions do not provide another route to it. This is why you review the complete set of operations a node can reach, rather than checking tool names alone.
How a Tool May Be Called
Once a tool is available, the harness must check the arguments on every request. A legitimate tool name does not make every possible use of that tool legitimate. A telemetry query, for example, may be allowed for one case and a limited time window without being allowed across your entire environment.
| check | what you verify | example |
|---|---|---|
| shape | the arguments match the expected schema | the time window is a number of hours |
| membership | the requested resource is allowed | an outbound request targets a listed domain |
| bounds | the request stays within defined limits | a query covers at most 24 hours and returns at most 500 rows |
A failed check usually means rejecting the request and returning an explanation the model can use to correct it. For some numeric limits, you can deliberately define a clamping rule: the harness reduces a requested value to the allowed maximum. In the figure below, a request for thirty days of telemetry is reduced to twenty-four hours.
That adjustment must be part of the tool's documented behavior, and the observation must state the window actually queried. Otherwise, the model could treat one day of evidence as if it covered the whole month. Where reducing a request would change its meaning in an unacceptable way, reject it instead. A host or domain outside an allowlist should not be silently replaced with a different one.
These checks also apply to tools that require approval. Validate the proposed host and operation before presenting them to the reviewer, and execute only the arguments that were approved. Human review does not replace the limits you have already decided the system must enforce.
The account used by a tool provides another layer of control. If a telemetry client only needs to read records, give it an account that the telemetry service permits to read those records and nothing more. The service can then reject a write even if faulty harness code attempts one. This is the principle of least privilege: each component receives only the access it needs for its job. It complements your argument checks with restrictions enforced by the system being accessed.
Decided Per Node
In a multi-step workflow, a node is a step with a particular responsibility, such as assessing a detection or preparing a response proposal. You can assign tools and limits to that responsibility instead of making the same capabilities available throughout the whole run.
A node that reads candidates needs the relevant lookups, but it has no reason to receive a host-isolation tool. A later node responsible for proposing containment may need that tool, with approval required and a defined set of eligible hosts. When the harness checks a request, it uses the policy for the node that made it. This keeps access aligned with the work being done at that point in the workflow.
Where the Decision Runs
The gate is the checkpoint between receiving a tool request and executing its implementation. The model's response gives the harness a proposed tool name and arguments. Before dispatching that request, the gate checks whether the tool is available to the current node and validates how the model wants to use it.
If the request passes and the tool is automatic, execution can proceed. If approval is required, the harness holds the validated request until the reviewer decides. If the name is unknown or the arguments violate a rule, the harness returns a rejection. A documented clamp can produce an allowed request, provided the changed arguments are recorded and made clear in the result.
For each request, the harness records the outcome, the requested and effective arguments, and the policy used to make the decision. Approval decisions also need the reviewer's identity and the action they authorized. These records let you reconstruct what happened when you inspect a run later, including calls that were refused before execution.
The figure follows five requests through these checks, one request at a time.
You can keep the policy in a versioned configuration file so changes receive the same review as code. With the same policy and relevant context, the automated checks should reach the same decision. An audit record that includes the policy version lets you explain why a call was allowed at the time, even if today's permissions are different.
Setting the Dial
You have two related choices to make: which tools run automatically, and how broad their permitted arguments are. An automatic query limited to one case is a different grant of authority from an automatic query across every case. Review both choices whenever you change what a node can do.
Permissions that are too broad can let a mistaken request cause damage. OWASP discusses this under excessive agency, including unnecessary capabilities, excessive permissions, and actions that proceed without appropriate oversight. Approval can reduce that exposure, but requiring it for every routine lookup creates its own practical problem: reviewers may become overwhelmed and stop examining requests carefully.
A useful starting point is a narrow set of automatic reads, with consequential actions held for approval. You can then use realistic tests and reviewed runs to judge whether a particular restriction should change. For example, you might establish that a telemetry query needs a three-day window for a specific investigation task, and expand that bound after checking the consequences for cost and access.
The figure shows one possible progression through those decisions. It is not a schedule for giving an agent more authority as it gets older. Each expansion needs a reason, some actions may always require approval, and you can tighten a policy when evidence shows that its current limits are inappropriate.
For every model request, you should be able to see what the harness permitted, which policy it applied, and what actually ran. That is how a boundary becomes something you can test, rather than something you only asked the model to respect.
Keep reading
- Next How an Agent Decides What to Do Next Think, act, observe: follow a tool request around the agent loop, and see how each result shapes the next question.
- Related Build Agents That Amplify You Collaboration by design, amplify rather than automate, and the right system for each job.
- Related Stop Overengineering Your Agents Only ever make things as complex as they need to be. What that means for an agent harness.
Ready to master agents for defensive security?
Start mastering agents for defensive security with my flagship, self-paced online course, designed specifically for defenders.
This lesson comes from Module 05, Tools. From there, the course goes into MCP, Skills, Initiation & Detection, Context Engineering, Assessment Skill and more.
What’s included
- 125 lessons
- Hands-on labs
- Build It Yourself guides
- Lifetime course access
- Course updates included
- One-time purchase
Stay in the loop
Keep learning. Keep building.
Every free lesson I publish comes together in one email, with a short note from me. That way you keep up as agents change defensive security, without chasing YouTube, LinkedIn and the blog.
A digest every second Thursday, with new articles and videos. No spam. No automated marketing flows.
References3
- OWASP: Excessive Agency. Examples of excessive functionality, permissions, and autonomy, with corresponding mitigations. genai.owasp.org ↗
- Trustworthy agents in practice. Anthropic's discussion of the tools, data, permissions, and environments available to an agent. anthropic.com ↗
- NIST: Least privilege. Definitions of limiting access to what an entity needs to perform its assigned task. csrc.nist.gov ↗