AI Products Ignore Their Own Warnings—Here’s What a Serious One Would Do

What Would a Serious AI Product Look Like?

Every AI tool admits it can make mistakes, yet none give you real tools to check them. This post argues that serious AI products would make error-checking a first-class feature, present citations as primary sources, ban first-person apologies, offer task-specific interfaces, and expose data provenance and reproducibility controls. Without these, they feel less like software and more like a grift.

If your product tells me that it makes mistakes and I must be the one to check for the mistakes, but then gives me zero tools to check for mistakes, I cannot take it seriously.
  1. ramity

    I'm very thankful for the section on reproducibility. I argue this is the single biggest hangup for the entire space. You CAN have temperature and determinism. I've been waiting for six years for a major provider to offer it, there is demand, but I've slowly come to realize the current game theory does not support it.

    For providers, not supporting deterministic eval means:

    - users use more tokens = more money

    - providers can generate more tokens per compute = more money

    - providers have cheaper hardware options (GPUs) = more money

    - providers models are harder to extract/distill = more money

    - providers are harder to hold liable for outputs = more money

    - providers can secretly use other models = more money

    - providers are harder to compare against others = more money

    - providers can cherry pick performance results = more money

  2. awakeasleep

    I have been dwelling on the "No First-Person Output" problem.

    I fully agree with the author's point that it's an incoherent interface for a tool. But more than that, it's a constant irritating reminder to me that these LLMs aren't actually thinking or synthesizing new ideas. The LLM is fundamentally not a person, and does not have a human's context, so representing itself with human pronouns and speech patterns is fundamentally contradictory and inaccurate. Author gets into that with the apologies, but once you start noticing it, it's everywhere.

    If these programs were actually capable of thinking, and committed to veractiy, they would represent themselves in a new way, and it would be insightful and interesting. We the users wouldn't have comfortable and misleading language masking the 'alien intelligence' and it would be a weird adjustment, but we would be adjusting instead of pretending.

  3. dofm

    The thing about the "AI can make mistakes, so double-check responses" thing is the essence of our future hellhole — deterministic software replaced with AI and legal disclaimers.

    The reason the firms do not want to invest in making fact-checking a first-class feature is that the appearance of being right is what people want from AI.

  4. sligbad

    Really refreshing read. This feels glaring in so many of these, and the methods to get things to "behave" of just slapping additional markdown prompts at various levels is both silly and ineffective.

  5. julesrms

    Good read. I think there are plenty of people who are reaching this point of wanting to shake off the novelty aspects of the agent coding experience and make it all a bit more grown-up.

    The stuff about context control has always been my itch. The scrollback that most agents show is not what the model is reading. Things get summarised, dropped, cached or never included at all, and the transcript carries on showing the original as though it were still there.

    It irked me enough to do my own agent: https://juggler.studio, explicitly to offer hands-on with the real context. The UX is all about making it easy to navigate and visualise every bit of the context, and even let you edit it. While it feels like other harnesses are actively trying to hide it from us..

More from this day

2026-09-28