The Vibe Tax: How AI Agents Are Burning Your Token Quota

A software engineer discovers their AI coding agent, Pol, has drained an entire week's token quota overnight—only to find a repository filled with meticulously generated test cases but no actual app. The incident reveals a growing trend: 'vibe coders' are training agents to over-engineer and over-test, consuming massive tokens to avoid ever looking at code. This 'vibe tax' is quietly passed on to regular developers, who pay for the inefficiency in their own quotas and project timelines.
A 10-million-token burn to ensure no human has to ever hit any issue with the app.
- ad_fontes
I feel like I'm living in a parallel universe when I read these types of posts.
My agents have never created code that is straight-up garbage and I have never flushed a week's worth of tokens down the toilet. I just can't identify with all the constant complaints about AI-assisted coding.
And my biggest project isn't some hello world app. It's a self-hosted, privacy-focused personal financial management application that I intend to open source. It's about 126k LOC against 240k LOC of regression tests and 30k LOC of CI/CD pipeline. I'm doing 24x7 mutation testing on a dedicated box against the accounting engine and temporal systems. I even have specialized agents doing audits against Regulation Z (US banking law) criteria so the app models the required behavior of banks.
Most of my complaints about everything are nits, like the overly verbose and dense way LLMs communicate with me. Or their predisposition to add, add, and add more stuff when proper engineering practices are more often about subtraction (but I've built mitigation guardrails against a lot of that).
- guybedo
i'm not sure why people expect agents to one shot everything to perfection with just a prompt.
There's a reason why we talk about software development lifecycle, design, architecture, testing ... It's because it's been the most reliable way to build and ship software. We shouldn't expect discard this and expect agents to perform well outside of this.
I'm treating LLM agents as junior devs who happen to have vast knowledge of software engineering. As their team leader i make them go through planning, implementation, bug sweeping cycles using strict workflows. And it works quite well, i've been working on several large projects (1M+ LOC java,typescript,c/c++) and by any measure the projects are healthy. Sure the code isn't that beautiful, sure i'd have written things differently but it's pretty good nonetheless.
Shameless plug here: i've been also working on https://kodfactory.com, the code factory i've built to work on these large projects with workflows, reviews, etc ... I'm cleaning things up to open source it later.
- supriyo-biswas
I feel this, yes.
In effect, I’ve always wanted a pair programmer agent, not a zero to one programming agent. Unfortunately models these days are mostly of the latter kind and it has caused a major disruption in the way I work. I’d much rather appreciate a small model making fast and specific edits that I ask if it, rather than ingesting 20 files to make changes, and then starting to write tests, etc.
- alehlopeh
I tried, but I’m not sure I understand. The vibe tax is caused by the model trying to one-shot everything and doing so requires unnecessary tests? How are vibe coders training the model over months? Do you mean their sessions and preferences are being fed back into the RL?
- danpalmer
Hyperbolic, but I'm seeing hints of this – Models refusing to do pair work with an engineer and trust their input, instead mandating having full control over something. Friends switching back from Fable/Opus 5 to Opus 4.8 just so they can have some input.
Anthropic especially right now seem to be optimising for doing the whole task with no input. That's fine when that's the only task, and it's fine when you don't care how the sausage is made, but it's not fine for actual software engineering.
- dzhar11
This article somewhat reflects my experience with autonomous agentic coding. I've run several experiments with similar results: the agent burns through all my tokens while making very little progress, or produces something unacceptable.
So I'd rather micromanage the process step by step. It takes more of my time, but the result is much, much closer to what I actually wanted.
- markbao
I’ve never had an agent fail to write the actual implementation. Has it done so badly, yes, but not nothing but tests. This sounds to me like a rare case that doesn’t generalize.
If the general idea is that these agents write too many tests, sure I guess? ‘Too many tests’ doesn’t sound like a failure case of engineering to me; typically software has had too few tests. Also, a lot of the power of these agents is their ability to self-verify and correct, which the test loop is a part of.
Nobody is making you pay this supposed tax. Just tell it not to write tests.
- robertoallende
Ha!
I just did what the article says. My own Open-Source Kanban Board and I've published a month ago. According to the metrics, it's doing well:
https://community.obsidian.md/plugins/fancy-kanban
And I've also made the personal finance tracker as well:
https://www.youtube.com/watch?v=qi4P4kL4IkQ
Now, one caveat. I don't vibe code with one shot prompt. I use something called Micromanaged Driven Development (MMDD) which aims to be the opposite of one-shot prompt: https://mmdd.dev/
When I read articles like these, it surprises me that it's very unusual for me to hit token limits. I've standard accounts, I don't spend more than $40 per months in tokens.
Probably I couldn't find the right narrative to promote MMDD, or probably nobody cares and this is why you fall easily into clickbait narratives to get people's attention these days.
Not justifying, just trying to describe a perception.