Your Agent Is Not the Model: The Key to Debugging AI Systems
A common confusion in AI development is conflating the model with the agent system. This article clarifies the layered architecture: the model (e.g., Sonnet, Opus) is the mathematical core, the inference service (e.g., AWS Bedrock, Anthropic's API) runs it, and the harness (e.g., Claude Desktop, Cursor) shapes inputs and outputs. Using a house-building analogy, it explains why the same model behaves differently across harnesses and provides a troubleshooting table to pinpoint issues like bad reasoning, missing context, or high cost. Precision in terminology leads to faster fixes.
When you can name the layer, you can fix the layer.
- agentdev001
The post ends with a comment on "its not about being pedantic..." so, a few not being pedantic bits:
In the table "Real world examples";
"Claude Desktop" houses three harnesses at the moment; Claude, Claude Cowork, and Claude Code.
"Claude CLI", I presume, is referring to Claude Code CLI. This is distinct from the 'ant CLI', which is sometimes referred to as 'Claude CLI'.
"Cursor" could be any of them- but, 'Cursor Agents', 'Cursor Cloud Agents', 'Cursor CLI', and whatever the vscode fork is called now, are distinct. Maybe not in the context of this blog post, but it isnt specified which is being referred to in the example table.
"ChatGPT" sounds like the chatgpt web interface. OpenAI's desktop app is named 'ChatGPT Desktop', and now houses 'ChatGPT work' and 'Codex' (Codex Desktop, not the TUI, though it does essentially wrap the tui and give it capabilities through app-built-in tools). I believe the ChatGPT web interface's harness can change a bit, depending on settings + subscription level (remote sandboxes, etc.) Additionally, there is a distinction in available models depending on which "ChatGPT" product is being used (instant/live/etc non-5.6 luna/terra/sol suite).
Inference service is more accurately 'default inference provider'.
Also, this post has an ai-generated smell.
- azath92
If the goal is to provide a distinction between model and agent, i think the "agent system" is doing too much heavy lifting in the example here.
A useful extension to this mental framework that i use when trying to make this distinction is the application (cursor) -> which sometimes includes an orchestrator and all of the QOL stuff like resuming, checkpointing, etc. single or multiple agents (cursor agents)-> and runs a single or many agent instances (single agent in cursor)-> service api-> model.
This is to address a confusion i often see with agent being conflated with the application that we use agents in, rather than the distinction in the article which tries to unpick agent-model confusion.
- 6keZbCECT2uB
A fun one is that in claude code, you can configure 'agents' which are prompt presets + some configuration. Or sub-agents sometime are indistinguishable from the foreground agent (usually called orchestrator) in configuration except that they have different contents in their context window (forks more or less).
IMO, if there's a ubiquitous term that is unambiguous, use it (harness, model). If there's an ambiguous term you have to explain, try not to use it. Language is for communication.
- yaaaaam
An agent, in general, is just whatever carries out a task on behalf of someone/something else.
- rwoerz
> An agent system is made up of several layers.
Why "layers"? The constituents of a Multi-agent System (MAS) [1] are called "agents". BTW: Synecdochical semantic diffusion is not uncommon in software engineering