AI agents · Harness engineering

Stop Overengineering Your Agents

Video coming soon

I want to give you a rule to weigh every sizing decision against as you build an agent harness.

§01

It Starts With Progressive Disclosure

If you've read anything about context engineering, you've probably met a principle called progressive disclosure: only ever give the model the exact information it needs at each step to achieve its goal, and nothing more. Don't pre-load the context window with everything that might become relevant. Supply what the current step requires, and let the agent pull in more as the task demands it.

The term is borrowed from UX design, where it has been around for decades. Good interfaces show a few core options and reveal the advanced ones only when asked.

Software engineers know the same idea as lazy loading: don't load a resource until the moment something actually needs it. Progressive disclosure is that rule applied to what a model is told.

I want to take that same principle and extend it. Progressive disclosure is usually framed as a rule about context. But its shape, exactly what the job needs, nothing more, carries over just as well to every other decision you make about the harness.

§02

The Rule

The golden rule of harness engineering

Only ever make things as complex as they need to be.

Almost every choice you make while building the harness, how many tools to give the agent, how many steps a loop runs, how much context you hand the model, is a chance to add more than the job in front of you needs. The rule means stopping at enough, one decision at a time.

Here is how that plays out across the pieces of a harness:

the golden rule
context
tools
MCP
skills
orchestration
RAG
RAG Use RAG only when the required knowledge is too large to inject into context directly.

The stopping point lands in a different place for each piece. The rule isn't "use less everywhere." It's match the need, then stop.

§03

Why Codify Something This Obvious?

Put plainly, the golden rule says: don't overengineer. That's advice as old as software itself, and about as controversial as telling someone to write tests. So why am I giving it a name?

The reason is a pattern I have seen repeatedly, in my own work and in other people's: something about building with AI pulls otherwise disciplined engineers toward complexity. People who would never add a needless microservice to a web app will wire up a five-agent pipeline, a vector database, and three MCP servers for a problem that needed one model call inside a for-loop.

I have guesses about why. The field is new enough that everything about it feels like it demands new machinery, and an elaborate architecture diagram demos better than an if-statement. But I'm more confident in the observation than in any explanation of it. The pull toward complexity is real, and it's stronger in this space than anywhere else I've worked.

Anthropic has noticed the same thing. In Building Effective AI Agents, they put it like this: "we recommend finding the simplest solution possible, and only increasing complexity when needed." And then the sentence I like most: "This might mean not building agentic systems at all."

When the company whose business is selling you model calls tells you that you might not need an agent, take the warning seriously.

§04

Does This Job Need a Model?

Anthropic's sentence asks that question about the whole system. In practice, you'll meet the same question at a much smaller scale, at nearly every step of the build: does this one job need a model at all?

The agentic layer is the most complex machinery you own. It's probabilistic, it's slower than native code, it costs money on every call, and it can be wrong in ways that look right. Deterministic code is the opposite: fast, effectively free, testable, and identical on every run.

So the default is code, and the model gets a step only when the work requires interpretation, which means ambiguity, judgment, or meaning. Those are the things code can't be pre-programmed to handle.

Here's what that looks like when you're looking for beaconing in network telemetry. Deciding whether a connection is regular enough to count as a beacon is arithmetic, so code does it. Deciding whether that regular connection is malicious or just a software updater checking in is judgment, so the model does it. It's the same data and two different jobs, and the rule decides which of them gets the model.

An agent is deterministic software around a non-deterministic model, and most of the system is the deterministic part. The golden rule is how that happens, one build decision at a time.

Start here
Available now

Ready to master agents for defensive security?

Start mastering agents for defensive security with my flagship, self-paced online course, designed specifically for defenders.

Follow a structured path from understanding agents to using them in your own environment.

What’s included

  • 120+ lessons
  • Hands-on labs
  • Build It Yourself guides
  • Lifetime course access
  • Course updates included
  • One-time purchase
References3
  • Progressive Disclosure. Jakob Nielsen, Nielsen Norman Group. The original UX principle: show the core options first, reveal the advanced ones on request. nngroup.com ↗
  • Effective context engineering for AI agents. Anthropic (Engineering). The principle applied to agents: treat context as a finite budget and let the agent discover detail just in time. anthropic.com ↗
  • Building Effective AI Agents. Anthropic (Engineering). The source of the simplicity quote: the simplest solution possible, complexity only when needed. anthropic.com ↗