Six Months of Writing Code Exclusively with AI Agents

Six Months of Writing Code Exclusively with Agents

Six Months of Writing Code Exclusively with AI Agents

After a self-imposed rule to never hand-write code again, a developer shares six months of experience delegating all coding to AI agents. From managing a dozen parallel agents to building a custom orchestration tool called 'botd', the journey reveals the hidden costs: cognitive overhead, validation challenges, and the difficulty of reviewing unfamiliar diffs. Despite agents passing tests and producing screenshots, the author often discarded their work, realizing that tools can verify correctness but not the worthiness of a change.

The tools could tell me that the change worked. They couldn’t tell me whether it was worth adding to the system.
  1. greenowl

    I committed to the fully agentic approach for a while. No code by hand!

    Setting aside the "capability" of the tools, I knew it was the wrong direction after a couple months or so in. I specifically remember during a large feature implementation running out of tokens for my 5 hour window or whatever, and I just couldn't continue on my own. Mostly due to laziness (I'll just wait till tomorrow when my limit resets!), but I also picked up on my first big whiff of skill rot / atrophy brewing, and that made feel uncomfortable.

    I just don't see this fully agentic approach going well long term. Talk about the ultimate dependency!! If the lights are "turned off" for whatever reason - outages, cost increases, or the agentic velocity finally reaches a complexity tipping point and you've lost control and understanding of your system to the point the agents are making things worse, whatever it is - do you really expect to be able to turn back the clock and step in to code at the productivity level and output you used to when you actually... wrote code?

  2. lostnfound8778

    > The bigger cost was the typing. Every time I wanted to build something, I could see the code in my head. I just couldn’t type it out fast enough.

    Facts. For me the bigger the gap between what i saw in my head and the speed with which my fingers could physically make it a reality the more stress i would feel and some marathon coding sessions would end w my back all messed up just from the tension

  3. thisisauserid

    And before LLMs he would have slapped together a twenty-line bash script and a cron job over lunch, and spent six months working on something else.

  4. redlewel

    How do people read these types of post with this AI flair, I couldn't read more than a couple sentences

  5. mrothroc

    Several comments here touch on the core problem: agents are writing more code than we can review. Seniors have never had enough time to review, and the prolific output from coding agents is making it worse. Moreover, the code is almost always good. So you're reviewing a tsunami of pretty good code, which means you get review fatigue and the whole thing just becomes theater.

    For me, this means the checks have to be more than just "I looked at it". There are two things that happen before I ever see it: first, as others have mentioned, I have a model from a different family review it with well-defined criteria. Same-family reviewers share bias, so it must be a different one. Second, I have a core set of deterministic gates (like lint, but also unit tests) that must run. In either case, failures go back to the coding agent.

    And only then do I bother. But I really don't read everything. If it is bog-standard CRUD operations, the agents are pretty good at that, especially if they are using mature packages. I focus on critical things, like how it enforces permissions.

    This works well for me, though it leaves one major issue untouched: whether this is worth building. The agents tend to be a bit overeager, so I have to do a lot of work up front to make sure the output will add value. The gates can only check the artifact, not my intent.

  6. iamflimflam1

    Writing code with LLMs takes discipline and is difficult. You need to learn how to do it and how to get the best from it.

    Software engineering is dead, long live software engineering!

  7. rich_sasha

    I feel a really odd dissonance with all these accounts. I have access to top models via Cursor at work. Every time I think to myself, here’s a tedious but well defined task, with clear success criteria, where you can achieve a lot with persistent iterations. I’ll give it to Claude to do.

    First, it takes me ages to describe the task properly. All these clear success criteria, well, instead of writing code from a clear spec in my head, I’m writing tons of prose, and trying to make it unambiguous. I’m programming in English++.

    But then, one time in 5, it produces something that kinda works. Maybe not perfect but good enough. Two times out of 5 it kind of sort of looks alright, but actually ignores 80% of the spec, or pays it minimal lip service in the comments.

    And sometimes, maybe a bit less than 2/5, it just completely diverges into total shit. Starts writing scripts that load the ast of the main script and pickle it then serialise to base64 for no good reason. Encode some stuff in strings then check ord(string[i]) repeatedly for string comparison. Eventually runs out of context and develops the LLM equivalent of severe dementia.

    I truly cannot reconcile my experience with people who seem to say “hey computer write this” and it’s a good use of their time.

  8. phildenhoff

    Unfortunate to see an otherwise-good article padded by AI-isms ("But here’s the twist: that search I described earlier was never sent to botd. All the history was in SQLite, so I pointed another agent at it and got the analysis anyway. The tool died; the data didn’t.", amongst other examples).

    Maisem Ali writes about how they have not stopped engineering despite handing production of code over to AI. I think that's an interesting perspective for a few reasons, but most of all, I'm not sure I agree.

    First, we see a prime example of how that can blow up. "botd died this month. It crumbled under its own weight." It's hard to imagine how an agent harness (or is it an overlay on top of other harnesses?) could crumble under its own weight from sheer technical complexity. What makes a tool like this "crumble"? If properly architected, it seems that any outside change could be adaptable. Failing tests could be fixed.

    This AI coding evolution has given us the ability to spin up prototypes we don't understand the inner workings of, but I believe for anything of value (and, I would argue, botd appeared to have value to Maisem) we should at the very least understand and influence the architecture and engineering of what we're building. If Maisem had done that for botd, it would not and _could not_ have "crumbled under its own weight". Whatever outside influence required change within botd would have been manageable.

    Second, the off-loading of the production of this article indicates to m […]

More from this day

2026-08-27