There's no reason for software to be slow anymore

Dan Luu argues that AI has slashed the cost of performance optimization, making it trivial to apply techniques once reserved for the largest projects. He demonstrates with a regex engine that compiles native code on the fly, a multi-threaded Azul AI that beats all comers, and a case where Claude outperformed a human performance engineer. The result: software can be dramatically faster, and there's no excuse for sluggishness.

Now that this N has dropped by a tremendous factor (variable but, in terms of human time, frequently 1000x / 10000x / 1000000x), the number of these kinds of optimizations it makes sense to do goes way up.
  1. ehnto

    One of the biggest causes of slowness is just waiting for web requests. The fact that so much software is either online or built using the same stack even if it isn't, puts all that software in this blocked/waiting state constantly while using it.

    Anyone not in the US feels this even more since so much online is US hosted, 300ms for every little interaction adds up quick.

    If your software has the affordance of a waiting dialogue or loading wheel for many of its UI controls, you are building with this default blocked assumption. Even if you are building something web based, ask yourself if that's actually necessary for your software or if you could build it differently to avoid constant UI blocking.

  2. eaftan

    I've been working on a similar agentically engineered regex project called SafeRE:

    https://github.com/eaftan/safere

    https://eaftan.github.io/safere-intro/

    Mine is for Java and is intended to be production grade. The first goal is to guarantee linear-time behavior to prevent ReDoS attacks. My collaborator and I have recently been optimizing it to try to surpass native RE2 in performance.

    It turns out optimizations are incredibly well suited for an agentic loop. You've got concrete acceptance criteria (must show a meaninging improvement on a benchmark case, must pass tests). The agent is really, really good at using tools like a profiler and disassembler, better than I am (and I've been doing this for 20 years). It also papers over things that would take me a while to learn, like how the in-incubation Vector (SIMD) API works in Java. I understand the concept but it would take me a while to understand Java's implementation. The agent can just read the docs and go.

    The key is creating a good benchmark suite and ensuring the agent doesn't ship optimizations that are too narrow or too focused on the benchmark cases. You also need a really strong test suite to make sure you're not regressing correctness. SafeRE has billions of tests; a subset of several million run on CI, and the others run on-demand.

  3. mccoyb

    Here's this boiled down:

    > A stochastic search process with an executable optimization objective over space of programs S can only maintain or improve the objective

    This is superoptimization. We've known this since the 80s (Massalin, STOKE is more recent: https://github.com/StanfordPL/stoke) The only novelty is that the proposer is now way better with LMs.

    Further, there's a large number of reasons for software written by agents to be slow:

    - LMs still don't do data or hardware-oriented design well out of the box, and therefore if you're engaging in any sort of serious novel work, beyond porting an extremely well-understood program with extremely well-understood workloads, you're going to be spending hours tracking down bad allocation decisions (c.f. why TigerBeetle doesn't use agents), which are often the root of evil (before you'd reach for anything further)

    - The knobs you'd need to get serious performance are nearly unreachable in languages which LMs are good at (even Rust requires a discipline that the default language doesn't enforce). When you drop into the lower realms, you're trading consumption context for access to these levers. The levers are also "soft": you find yourself writing a bunch of skills, and tools to try and enforce the discipline.

    The reality is to get performant code (quickly) out of an agent, you need to know how to write performant code (and you need to know how to surface the information that you'd use to create a verifier for such a thing to the […]

  4. hunterpayne

    "LLMs are causing slow, bloated, code are going to eat crow once they re-write everything in super-optimized assembly."

    This person doesn't understand how to make efficient code. I can write code in almost any language (with a couple of exceptions) that outperforms "super-optimized assembly". Writing efficient code isn't about the language, and often isn't about the best algorithms either (but sometimes it is). Its about optimizing memory and cache use. And that's orthogonal to anything the author is writing about. Also, LLMs are terrible at optimizing memory utilization. There is just too little training code that does it well and far too much that doesn't.

    As proof, I'm can literally feel the web getting slower and I bet many others feel this as well.

  5. chvid

    ChatGPT MacOSX is the only software that regularly crashes on my machine when its memory consumption for no apparent reason spins up towards 50 GB.

    And that software is build by some of the highest paid software engineers on the planet with full access to all the LLM compute in the world.

  6. intrasight

    I've been using computers for 4 decades. They have gotten no faster. The nuclear plant computer system we built in 1989 had to present selected screens in 1 second. I don't think any apps I use today can do that.

  7. dkersten

    > Completely agree with your closing point. Dynamic custom software, fitted to a particular workload rather than a class of workloads, seems like a very likely outcome.

    I disagree with the premise that this is the desired outcome. If every piece of software is bespoke and everyone’s instance of it works slightly differently, then it’s impossible to get support or a shared knowledge of how it works. There’s no “just share the excel file”, there’s no “press the this button on the left”, there’s no “oh I use program X to solve Y” (instead you have to know what you need so the custom software can solve it, but my time in startups taught me that most users don’t know what they want or need).

  8. jjcm

    This speaks to me. I've been running an autoresearch loop the past couple of days to improve the load time of my various projects' frontends.

    I've been really, really impressed with how effective this is. I went from a 4s load on simulated slow 4g to ~750ms: https://image.non.io/speedup-graphs.webp

    Side by side vid of the results: https://video.non.io/speedups.mp4

    This was for https://non.io, which is something I had purposefully written to be as fast as possible (hand wrote all the comopnents, didnt even use react).

    I've been considering creating a skill / utility to do this based on learnings from the speedups - would others find this kind of thing useful?

  9. arjie

    I understand now that most software is slow because of co-tenancy reasons requiring controlling resources or simply because they're safely insulated from competition. e.g. GitHub is the former: you can give yourself a git host and CI/CD system that is much higher quality by yourself since you're probably not using its social features. I think things like Apple's five-finger inward gesture are the latter. Once you could do it and start typing but nowadays it needs to render the animation etc. before keystrokes register. This software is slow because you cannot replace it in MacOS.

    But all these things will change in time. Hell is other people's software.

  10. bdhdhduuyd

    For me the comparison has always been 3DsMax vs Blender. Same kind of software, same kind of features, but Blender is so much faster.

    Architectural decisions have always been important.

More from this day

2026-08-22