The Harness Is the Thing: How One Developer Cut Frontier Model Costs by 75%

Scott Fryxell describes how he built a personal AI development harness that unifies multiple TUIs (Cursor, Claude, and Pi) around shared skills and AGENTS.md. By using a planner/worker/critic/promoter workflow and reserving frontier models for planning and promotion, he cut frontier usage by 75%, relying on cheaper models like DeepSeek for most tasks. The harness also extends to his product, enabling advanced use cases like scripting the app via a headless browser. He argues that the harness—not the model—is the key to productivity, and that commoditized models make switching trivial.
The harness is the thing; the fulcrum from which my expectations meet the LLM's capabilities.
- dmantis
> Single developer projects can build to the caliber and consistency of large development teams.
Yet the simple blog website static page saying that looks very weird and broken on the desktop firefox.
How large should be a development team to make proper margins in 2026?
- AirMax98
Reading this really makes me wish that I had a slightly better workflow. I'm really soley dependent of Fable to the point that I don't use other models, and I've already sort of hit a point where I'm running into usage limits every week. I am really living on borrowed time — when Anthropic finally collapses their 50% usage increase at the end of August, I'll definitely be forced to switch my workflow. When that happens, I have a hard time imagining that I'll be sticking with a single model on a single provider.
- douglee650
Author states, “Single developer projects can build to the caliber and consistency of large development teams.”
When I, as a single person, can produce a project in one month that would have taken a team of four people three months to produce, why would I care about token cost? I’m now spending $500/month instead of $40,000 month to get the same thing 3x faster. $500 for a project instead of $120,000. (Assumes my cost, $40k is the other three people)
It’s a no-brainer —- use frontier all the time.
- andai
I've been very happy with Luna but my approach is "many bite sized edits" for which models basically hit saturation a year ago.
(I also tried the "let a massive model make massive changes" approach and am still psychologically recovering from the experience. The codebase may never recover!)
Also, Luna and DSV4 Flash seem to be on par now except Luna is faster and cheaper?
- esalman
What I've learned in last week is that a harness is basically a while loop.
In each iteration you make an LLM call, perform some work (e.g. tool call), augment the prompt (append or compact etc.)- not necessarily in that other- and continue.
Until an end condition is satisfied. Then you break out.
- nullbio
LoRA adaptive learning using open-weight models and your own reasoning traces is the thing. The big labs have a mammoth job ahead of them if they want to compete with running your own model - they will basically have to give every single user their own persistent virtual machine. When it's all said and done, I think their only really moat will be as inference/hardware providers. Stripe buying OpenRouter was a very smart bet.
- liampulles
One will lose the opportunity to develop domain understanding if they do not get into the weeds of thinking through the problem.
I use Claude plan mode to do relatively small changes and even then I find that if I actually try and think through the problem and solve it myself that I find good metaphors that will aid future work, and I will discover tangential issues which are then important to look at.
- andunie
I don't understand why no one has tried to make a harness without full shell access yet.
It would be so much safer.