How an Agent Decides What to Do Next
Adapted from my course, Build Your Own Agentic Defense.
A tool lets a model request an operation that the harness performs on its behalf. The result then needs to make its way back into the conversation. When a tool returns information, the model needs an opportunity to read it, reconsider the task and decide whether to ask for anything else.
Start with a single model call. The harness sends a request and receives a response. If that response contains a request for one of our tools, the harness can run it, but the model has not yet seen the result. Our code needs to include that result in another model call. Repeating this sequence gives us a conversation in which the model can gather information over several turns before answering.
Think, Act, Observe
We use TAO, short for Think, Act, Observe, to describe this cycle. The model considers the task and the information available to it, requests an action when it needs one, and receives the result through the harness. That result becomes part of the information it uses on the next turn. The three names help us follow the cycle, while the implementation handles the messages, tool execution and checks that connect its stages.
A related approach appears in the ReAct paper, which explores interleaving reasoning with actions and observations. ReAct is useful background for understanding this pattern. Here, we are using TAO to explain how our harness carries a tool request through execution and brings the result back to the model.
In a transcript-based harness, the conversation is stored as a transcript in our application. Each API call supplies the model with the history it needs for that turn. When a response requests a tool, that call has finished; our code runs the tool and makes a subsequent call with the result included. The conversation continues because the harness preserves its messages and passes them forward.
Following One Tool Call
To see how the responsibilities fit together, suppose the model is investigating a connection and wants information about its destination. We can follow that request from the model's decision to the lookup result and then to the next turn of the conversation.
Think
The harness sends the task, the transcript so far and the available tool definitions to the model. The model uses this context to decide how to respond. It may already have enough information to answer, or it may identify a question that one of the tools could help resolve. In our example, it decides that learning about the destination would help it interpret the connection.
This is what we mean by the Think phase. We do not need the model to print a separate explanation of its reasoning before every action. For the harness to continue, it needs a response it can handle, such as a structured tool request or an answer. Whether a provider also exposes reasoning information is a separate detail from executing the requested tool.
Act
The model returns a tool request containing the tool's name and its arguments. The harness checks that the tool is available and that the requested arguments satisfy its rules. For a destination lookup, those checks might include validating the address and confirming that the operation is permitted. If the request passes, the harness runs the lookup; if it fails, the harness can return an error that explains why the operation was rejected. The model chooses what to request, while the harness controls execution.
Observe
Once the lookup finishes, the harness records the tool result alongside the request that produced it. It then calls the model again with that updated transcript. The model can now use the destination information together with the original connection evidence to decide what to do next. It might request another lookup, revise its interpretation or provide an answer.
As this continues, the transcript grows. Tool results can be much larger than the questions that produced them, so several calls may add substantial material to the model's context. The harness therefore needs limits on result size and a deliberate approach to which history it keeps. In a short investigation, it may send the complete transcript on every turn; longer investigations may require selecting or summarizing material while retaining the evidence needed for the task.
Completing the Task
A final answer tells the harness that the model is ready to finish, but the application still needs to check whether the response meets the task's requirements. Suppose the result must be a JSON object with three required fields. A schema check can detect a missing field or the wrong value type, allowing the harness to request a correction if another turn is available.
That check confirms the format of the answer. Assessing whether its claims are supported by the evidence is a separate responsibility. The harness also needs to distinguish a completed answer from a response that stopped because of an error or a generation limit. A response ending does not, by itself, establish that the investigation succeeded.
Why Let the Model Choose the Next Tool?
A data-distillation pipeline is an example of a workflow whose stages are defined in code. We know which computations to perform and how their results feed the next stage, so we can specify that process in advance. During an investigation, the next useful question may depend on an interpretation of the evidence we have just received.
Consider a destination lookup that identifies a known software update service. One useful next step would be to check whether the corresponding software is installed on the host. That gives us a possible explanation to evaluate alongside the initiating process, timing and other evidence. If those details agree with expected update activity, the investigation may have enough support to resolve the concern. The presence of the software alone would not establish that the connection was benign.
If the destination is unfamiliar, the model might instead ask which process opened the connection. That result could give it a different lead to follow. Both investigations begin with the same lookup, but its result helps determine which question is useful next. The figure uses illustrative tool names to show those two possible routes.
A script can also make this choice. We could write a condition that selects one lookup for a known update service and another for an unfamiliar address. When the rules are clear and manageable, that is a reasonable implementation. The reason to consider a model is that an investigation may involve many combinations of evidence and questions that are difficult to capture in a small set of rules. The model can interpret the available context and propose a next step without us specifying every possible route beforehand.
That flexibility needs to be useful enough to justify the extra uncertainty and cost. A model can choose an unhelpful tool, repeat a request or draw the wrong conclusion from a result. We therefore keep the surrounding workflow explicit and give the model discretion only within the part of the investigation where its choices are useful.
Keeping the Investigation Bounded
The harness sets the available tools and the limits within which the loop runs. For a read-only investigation, that might mean giving the model tools to retrieve candidates, inspect their evidence and look up asset context. Each request is still subject to validation, even if the model has already made several successful calls. A useful earlier action does not grant permission for a different one.
We also need a limit on how much work one investigation can consume. For example, the harness might allow five model calls and enforce a token or time budget. When a limit is reached, it should stop and record that the investigation ended before completion. Depending on the workflow, that result might be deferred or sent for review. Reaching a budget is not evidence for a verdict, and the harness should not present an unfinished investigation as a successful one.
The transcript below follows a short example through two tool requests. The lines labelled Thought are teaching illustrations of the decision being made at each point; they are not a requirement that a provider expose its internal reasoning. Watch how the first result supplies a candidate ID that the model uses in its next request.
The first request finds a process candidate on DEVBOX-07. The second retrieves its related process event, allowing the model to connect the recorded process chain to the activity it is investigating and recommend escalation. The example shows why the transcript matters: the later request depends on information returned by the earlier one. That is the loop: each result informs the next request, while the harness controls what can run and when the investigation must stop.
Keep reading
- Next What Is Threat Hunting? Bring human judgment into detection, understand its limits, and see how agents can extend it.
- Related Build Agents That Amplify You Collaboration by design, amplify rather than automate, and the right system for each job.
- Related Stop Overengineering Your Agents Only ever make things as complex as they need to be. What that means for an agent harness.
Ready to master agents for defensive security?
Start mastering agents for defensive security with my flagship, self-paced online course, designed specifically for defenders.
This lesson comes from Module 05, Tools. From there, the course goes into MCP, Skills, Initiation & Detection, Context Engineering, Assessment Skill and more.
What’s included
- 125 lessons
- Hands-on labs
- Build It Yourself guides
- Lifetime course access
- Course updates included
- One-time purchase
Stay in the loop
Keep learning. Keep building.
Every free lesson I publish comes together in one email, with a short note from me. That way you keep up as agents change defensive security, without chasing YouTube, LinkedIn and the blog.
A digest every second Thursday, with new articles and videos. No spam. No automated marketing flows.
References2
- ReAct: Synergizing Reasoning and Acting in Language Models. Yao et al., 2022. Explores how language models can combine reasoning with actions and observations while working through a task. arxiv.org ↗
- How tool use works — Claude Platform documentation. Explains how an application handles tool requests, executes operations and includes their results in subsequent model calls. platform.claude.com ↗