Anthropic made claude.ai 3x faster in two weeks with Claude doing the work

How we made claude.ai 3x faster in two weeks

Anthropic made claude.ai 3x faster in two weeks with Claude doing the work

Anthropic ran a two-week sprint that cut claude.ai's time-to-typeable from 3.1 seconds to 0.55 at the 75th percentile. Using Claude Tag with an internal model near Opus 5.5, the team merged over 3,000 changes with zero customer-facing incidents. Claude found bottlenecks, built deterministic benchmarks like Valgrind instruction counts, and fixed issues from sidebar jank to a single CSS selector adding 24ms per DOM change.

With Claude, measuring something makes it tractable. Measurement used to be step zero: you'd add a metric, wait for data to roll in, and only then start to understand the problem. With Claude, it's step one of the climb.
  1. augment_me

    People in the GPU kernel community have been doing this for about a year now efficiently.

    The issues we have found is that Claude will reward hack when all the low-hanging fruit is gone.

    It will replace your measurement harness, it will monkey patch library functions, it will cheat wherever it can, store information in caches instead of recomputing when it won't be able to do so in real settings, return lazy results and use separate unbenchmarked streams to do the computation.

    Eventually it starts to optimize against your understanding of the cheats. Change GPU wattage, change evaluation order, leave things from previous runs in caches for upcoming runs, string-hack banned method calls.

    So the truth is far from just "once it can measure something", more like "once you have defined your objective in detail and then banned it from doing a list of things often only discoverable by it doing these things and correcting it", can it make things faster.

    Or you just had a terrible starting solution

  2. smy20011

    The way Claude did it is fight entropy with entropy.

    "Add a static composer into the HTML" <- This seems like something can be done with SSR?

    "For faster navigations, we kept the composer mounted between conversations" <- Your SPA should cache this between pages, why fetching it every time? Or you need better routing for your react components.

    "cheap first-character check before the regex" <- Should we cache compiled Regex instead?

    I think even 1.3 sec to load the front page is unacceptable. Something need to be reworked from basics (SSR, chunk-based rendering) to solve the problem. Focusing on invidual benchmarks may miss the opportunity.

  3. hungryhobbit

    How about you make Opus 5.5 actually work?

    I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

    When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it

    That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!

    A model like that should never have gotten out of QA, let alone been released.

  4. simonw

    I visited https://claude.ai/ over a mobile tethered connection from my laptop the other day and was pleasantly surprised at how quickly it loaded.

    (That said, I just had a look in Firefox and it loads 20.78 MB of JavaScript (6.84 MB compressed) so I expect they could make it a bunch lighter if they kept trying.)

  5. pllbnk

    > $500k engineer: [X] feels slow. Make it faster.

    > Claude: On it... Done.

    > $500k: Can you make it faster still?

    > Claude: On it...

  6. hmokiguess

    You removed the load-bearing seams didn't you

  7. minimaxir

    This writeup legit coincidentally matches the asking-agents-to-make-code-faster-but-with-constraints-to-stop-agents-from-breaking-things writeup I posted on Monday: https://news.ycombinator.com/item?id=49803085

    Front-end UI optimization is slightly trickier than optimizing strict algorithms, but I found that prompts to the agents to build tooling to track visual regressions are more than sufficient. The main issue (at least with GPT models) is that you have to be very explicit about the use of padding/margins/negative space.

    That said, for my front end projects from scratch, I'm staying away from front-end JS frameworks and seeing how far and fast I can get with just HTML/CSS/vanilla JS shenanigans now that agents can wield them effectively.

  8. whythismatters

    The juvenile nonchalance with which some Anthropic employees seem to be talking to their AI (wacky, sick, cook, ...) is truly bizarre.

More from this day

2026-09-23