AI agents · Harness design

How an Agent Decides What to Do Next

Adapted from my course, Build Your Own Agentic Defense.

Watch on YouTube ↗

A tool lets a model request an operation that the harness performs on its behalf. The result then needs to make its way back into the conversation. When a tool returns information, the model needs an opportunity to read it, reconsider the task and decide whether to ask for anything else.

Start with a single model call. The harness sends a request and receives a response. If that response contains a request for one of our tools, the harness can run it, but the model has not yet seen the result. Our code needs to include that result in another model call. Repeating this sequence gives us a conversation in which the model can gather information over several turns before answering.

§01

Think, Act, Observe

We use TAO, short for Think, Act, Observe, to describe this cycle. The model considers the task and the information available to it, requests an action when it needs one, and receives the result through the harness. That result becomes part of the information it uses on the next turn. The three names help us follow the cycle, while the implementation handles the messages, tool execution and checks that connect its stages.

A related approach appears in the ReAct paper, which explores interleaving reasoning with actions and observations. ReAct is useful background for understanding this pattern. Here, we are using TAO to explain how our harness carries a tool request through execution and brings the result back to the model.

In a transcript-based harness, the conversation is stored as a transcript in our application. Each API call supplies the model with the history it needs for that turn. When a response requests a tool, that call has finished; our code runs the tool and makes a subsequent call with the result included. The conversation continues because the harness preserves its messages and passes them forward.

one turnA FRESH API CALLthinkMODEL REASONSactASKHARNESS RUNSobserveHARNESS RETURNSanswerNO TOOL REQUEST
ANSWER In this completed example, the model returns an answer for the harness to check. A real loop also needs limits that can stop it before the task is complete.
§02

Following One Tool Call

To see how the responsibilities fit together, suppose the model is investigating a connection and wants information about its destination. We can follow that request from the model's decision to the lookup result and then to the next turn of the conversation.

Think

The harness sends the task, the transcript so far and the available tool definitions to the model. The model uses this context to decide how to respond. It may already have enough information to answer, or it may identify a question that one of the tools could help resolve. In our example, it decides that learning about the destination would help it interpret the connection.

This is what we mean by the Think phase. We do not need the model to print a separate explanation of its reasoning before every action. For the harness to continue, it needs a response it can handle, such as a structured tool request or an answer. Whether a provider also exposes reasoning information is a separate detail from executing the requested tool.

Act

The model returns a tool request containing the tool's name and its arguments. The harness checks that the tool is available and that the requested arguments satisfy its rules. For a destination lookup, those checks might include validating the address and confirming that the operation is permitted. If the request passes, the harness runs the lookup; if it fails, the harness can return an error that explains why the operation was rejected. The model chooses what to request, while the harness controls execution.

Observe

Once the lookup finishes, the harness records the tool result alongside the request that produced it. It then calls the model again with that updated transcript. The model can now use the destination information together with the original connection evidence to decide what to do next. It might request another lookup, revise its interpretation or provide an answer.

As this continues, the transcript grows. Tool results can be much larger than the questions that produced them, so several calls may add substantial material to the model's context. The harness therefore needs limits on result size and a deliberate approach to which history it keeps. In a short investigation, it may send the complete transcript on every turn; longer investigations may require selecting or summarizing material while retaining the evidence needed for the task.

Completing the Task

A final answer tells the harness that the model is ready to finish, but the application still needs to check whether the response meets the task's requirements. Suppose the result must be a JSON object with three required fields. A schema check can detect a missing field or the wrong value type, allowing the harness to request a correction if another turn is available.

That check confirms the format of the answer. Assessing whether its claims are supported by the evidence is a separate responsibility. The harness also needs to distinguish a completed answer from a response that stopped because of an error or a generation limit. A response ending does not, by itself, establish that the investigation succeeded.

harness
model
API call · task + transcript + tools
think the task asks about this connection I do not know what the destination is lookup_destination can tell me
response · lookup_destination(addr)
act · the harness tool on the list · arguments legal run lookup_destination result: a well-known update server
model idle
API call · transcript + tool result
observe · think again the result is in the transcript is the goal achieved?
yesfinal answer · loop exits
noanother tool call · back to Act
EXIT OR GO ROUND The model can return an answer for the harness to check or request another tool. The harness continues only while its turn and resource limits allow it.
§03

Why Let the Model Choose the Next Tool?

A data-distillation pipeline is an example of a workflow whose stages are defined in code. We know which computations to perform and how their results feed the next stage, so we can specify that process in advance. During an investigation, the next useful question may depend on an interpretation of the evidence we have just received.

Consider a destination lookup that identifies a known software update service. One useful next step would be to check whether the corresponding software is installed on the host. That gives us a possible explanation to evaluate alongside the initiating process, timing and other evidence. If those details agree with expected update activity, the investigation may have enough support to resolve the concern. The presence of the software alone would not establish that the connection was benign.

If the destination is unfamiliar, the model might instead ask which process opened the connection. That result could give it a different lead to follow. Both investigations begin with the same lookup, but its result helps determine which question is useful next. The figure uses illustrative tool names to show those two possible routes.

tools it may call
lookup_destination
list_installed_software
process_for_connection
get_related_events
four tools, and nothing else
finding suspicious connection DEVBOX-07 → unknown host
tool call lookup_destination what is this address?
answer a well-known update server
answer an address nobody has seen
tool call list_installed_software is that software on the host?
tool call process_for_connection which process opened it?
outcome resolved update explanation supported
outcome keep going new lead to follow
WHY THE MODEL CHOOSES A script could select these branches using explicit rules. Here, the model interprets the result and chooses from the available tools, subject to the harness checks.

A script can also make this choice. We could write a condition that selects one lookup for a known update service and another for an unfamiliar address. When the rules are clear and manageable, that is a reasonable implementation. The reason to consider a model is that an investigation may involve many combinations of evidence and questions that are difficult to capture in a small set of rules. The model can interpret the available context and propose a next step without us specifying every possible route beforehand.

That flexibility needs to be useful enough to justify the extra uncertainty and cost. A model can choose an unhelpful tool, repeat a request or draw the wrong conclusion from a result. We therefore keep the surrounding workflow explicit and give the model discretion only within the part of the investigation where its choices are useful.

§04

Keeping the Investigation Bounded

The harness sets the available tools and the limits within which the loop runs. For a read-only investigation, that might mean giving the model tools to retrieve candidates, inspect their evidence and look up asset context. Each request is still subject to validation, even if the model has already made several successful calls. A useful earlier action does not grant permission for a different one.

We also need a limit on how much work one investigation can consume. For example, the harness might allow five model calls and enforce a token or time budget. When a limit is reached, it should stop and record that the investigation ended before completion. Depending on the workflow, that result might be deferred or sent for review. Reaching a budget is not evidence for a verdict, and the harness should not present an unfinished investigation as a successful one.

The transcript below follows a short example through two tool requests. The lines labelled Thought are teaching illustrations of the decision being made at each point; they are not a requirement that a provider expose its internal reasoning. Watch how the first result supplies a candidate ID that the model uses in its next request.

think
act
observe
transcript
Thought the npm-cache path needs its process chain — find a process candidate on DEVBOX-07
Action query_candidates(host: "DEVBOX-07", type: "unusual_parent_child_anomaly")
harness name allowed ✓ · arguments valid ✓ · executing…
Observation UPCA-001 · unusual_parent_child_anomaly · DEVBOX-07 · one linked process event
turn 2 · fresh API call
Thought UPCA-001 holds the event; read the chain it recorded…
Action get_related_events(candidateId: "UPCA-001", eventTypes: ["process_create"])
harness name allowed ✓ · read-only ✓ · executing…
Observation evt-00676 · node.exe → powershell.exe -enc → drops node-bridge.exe
Answer a loader chain into node-bridge.exe on DEVBOX-07 — escalate BCN-002
ANSWER The process chain gives the model evidence for recommending escalation. It returns an answer that the harness can check against the task requirements.

The first request finds a process candidate on DEVBOX-07. The second retrieves its related process event, allowing the model to connect the recorded process chain to the activity it is investigating and recommend escalation. The example shows why the transcript matters: the later request depends on information returned by the earlier one. That is the loop: each result informs the next request, while the harness controls what can run and when the investigation must stop.

Keep reading

Start here
Available now

Ready to master agents for defensive security?

Start mastering agents for defensive security with my flagship, self-paced online course, designed specifically for defenders.

This lesson comes from Module 05, Tools. From there, the course goes into MCP, Skills, Initiation & Detection, Context Engineering, Assessment Skill and more.

What’s included

  • 125 lessons
  • Hands-on labs
  • Build It Yourself guides
  • Lifetime course access
  • Course updates included
  • One-time purchase
References2
  • ReAct: Synergizing Reasoning and Acting in Language Models. Yao et al., 2022. Explores how language models can combine reasoning with actions and observations while working through a task. arxiv.org ↗
  • How tool use works — Claude Platform documentation. Explains how an application handles tool requests, executes operations and includes their results in subsequent model calls. platform.claude.com ↗