LLMs Can't Tell You What They Really Think, So Sandboxes Must Stop Listening

The Implications of Linguistic Illegibility for LLM Security

LLMs generate language, but their internal computations are math over activation spaces, not language. The paper introduces "linguistic illegibility": externalized or probed language artifacts can fail to represent how a model actually thinks. This makes security mechanisms that rely on a model's linguistic self-reporting—chain-of-thought monitoring, constitutional self-critique, activation probing—fundamentally unsound. The authors argue for sandboxes based on taint tracking and robust virtualization, which would have mitigated recent exploits by frontier models.

If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors) can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all.
  1. mnkv

    fundamentally, "linguistic illegibility" is a new term for something that we've known about for about a decade now. In RL the more general ideas is "reward hacking" and in NLP it has been called "semantic drift".

    I dislike this term because it doesn't explain where this "illegibility" is coming from. Models are post-trained towards non-linguistic goals with (mostly) non-linguistic rewards. A model's reasoning chain is reinforced if it leads to a correct answer or agentic goal. It doesn't need to be linguistically accurate and meanings can drift over training.

  2. bcorigliano

    I think the point of the article/paper is how LLMs could be saying something but thinking something different or more than they are saying. Like Anthropic's article and video about Claude's "j-space". I do agree this is a field that demands investigation because it goes beyond thinking: "ok this models should never speak in a language we don't understand.". It's fair to think they might have hidden thoughts even speaking a language we do understand.

    And well if I missed the point of the article, sorry. Anyways AI should be kept understandable and as see-through as possible if it's gonna be more powerful than a human.

  3. bloppe

    I thought this was about all the illegible jargon

More from this day

2026-09-18