Context management, not planning, drives coding agent performance under tight budgets
An Empirical Study of Harness Design for Coding Agents
A 43-page empirical study isolates three harness components—planning, action space, and context management—across 176 matched settings on SWE-Bench Verified and Terminal-Bench 2.1 with four models. Context management matters most as the context-window budget tightens, mainly by preventing overflow failures. Rule-based elision before LLM summarization is the most efficient strategy; making elided content recoverable adds unused machinery. Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger ones. Predefined tools help models with weaker bash skills, while bash-capable models do fine with a bash-only interface at lower cost.
Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy.
- gps372
Haven't gone through full PDF as its very detailed, few things have resonated with me so far.
Basically if a Car A is performing better (be it speed, milage or in general sense) than Car B, then it is not necessarily because its engine. It could be because of better tires, better gearbox, lighter body, better usability of features, etc.
You can implement an AI feature (like AI for BI) in different ways even with the same model - via ReAct-loop, or plan-and-execute, or hybrid. You can make it stateless, stateful, RAG-based, etc. depending upon whether you want to prioritize result accuracy or depth of analysis. You can use LLM to generate either intent (requires lesser reasoning) or the queries itself (requires much more capable model).
Your harness can adapt to the underlying model's native capabilities, or can make up for its absence, e.g. query generation in above example requires your model to have MOE capabilities but intent generation wouldn't.
- lieret
Cool study, we definitely need more principled studies on the role of harnesses. I'd also say that there aren't too many benchmarks where the more complicated harnesses consistently outperform extremely simple agents. But I'm also biased, because I wrote https://github.com/swe-agent/mini-swe-agent/ , which is probably the most minimal agent out there (it started as just 100 lines, all included), and it's used in a lot of benchmarks like DeepSWE, terminalbench, programbench (seems like it's still top of the ranking for TB3, but wasn't evaluated with the best models on TB4).
- vblanco
This is done on Nemotron models + mistral, so its not very relevant to the current frontier of cheap chinese models + big models from Claude/GPT. Big miss not having qwen or deepseek in this research.
- rahulmax
Quite inline with what I had found with my Claude code sessions over the last year. I wrote about this a few months ago.
https://rahulmax.com/notes/how-i-keep-the-ai-bill-down/
In their case, context management pays off more the tighter your window. Their gap between managing and not managing is 35.7 points of success rate at 32k and 2.7 points at 128k. My version of that was a rule I stick to, as much as I can. I checkpoint a session at about 25-30% of the window, write the state out to a PROGRESS.md and a JSON file of the requirements, and start fresh. This restart costs me 30 seconds, since a bloated session doesn't get any cheaper the longer you stay in it.
Also worth knowing that the models are Nemotron-3 and Mistral-Medium, not the frontier models most people here are paying for.
- agentdev001
As far as I can tell, the paper says "bash capable", without ever describing what that means. How would one know whether a given model is "bash capable" or not?
I would have to imagine, that Luna would very much fall into the camp of "bash capable". At which point- it seems to me that adding any tools beyond just Bash requires some rigorous testing and verification that value is being added.