Adding 'Do not guess' cut AI made-up fields from 70.7% to 20.2%
Calling the AI bluff: Adding "Do not guess" cut made-up fields from 71% to 20%
A benchmark tested 16 AI models and 3 paid APIs on web extraction, asking for fields that were sometimes absent. With the instruction 'Use null for any field whose value is not on the page. Do not guess.', models fabricated 20.2% of missing fields versus 70.7% without it. Firecrawl invented 24 of 36 missing fields, more than 13 of 16 models with the sentence. A cheap checker like GPT-6 Luna caught 38 of 49 made-up values for $0.0049.
All 16 models made up more without the sentence: 405 of 573 missing fields without it (70.7%), 116 of 574 with it (20.2%).
- BatchJob
The LLM will take a statistical path to reply and will not refuse to do so under any circumstances except where its been coded to do so.
Your examples are contrived and will not be borne out in any significant way. Inaccuracies are usually not simply made up claims they are false information based on statistical paths to misleading results or which elude the current context. LLMS dont understand the word dont. LLMS dont understand the meaning of any words.
Neither you, nor aristotle nor god will ever make an LLM return the truth or correct results via prompting.
- l1ng0
We're all turning into pigeons in a Skinner box.
- datsci_est_2015
Cool, this will be added to harnesses and then it’ll stop being effective and we’ll move on to the next magical incantation.
- thallavajhula
I've tried all of these and nothing really works. I have only 1 line in my CLAUDE.md file and that is "Always ground your responses." and that's it.
Claude didn't care about it. When I pointed that out, it was apologetic and that was it.
- literalAardvark
I've used "you're not trained on this data, return exclusively grounded results" to good effect.