What is neuralese and why is it bad?
What is Nueralese and Why is it Bad
Neuralese is the idea of replacing a model's natural-language chain-of-thought with a stream of numbers, making its reasoning unreadable to humans. This piece explains why that's dangerous: it would eliminate chain-of-thought monitoring, one of the few safety tools we have for tracking model intent. Citing recent leaks about OpenAI's Astra model, the author argues that even hybrid approaches are a slippery slope toward full neuralese, and urges journalists and employees to push back.
So moving away from monitorable chain-of-thought to neuralese seems very bad, with dubious benefits.
- _alternator_
The argument is that chain-of-thought without "tokens" would remove a major interpretability and model intent control pane. This is definitely borne out in the OpenAI's report on the huggingface attack; they had turned of CoT monitoring for those jobs, and claim that they could have (would have?) prevented the behavior had they been monitoring it. They've changed their internal policies to always monitor CoT.
That said... CoT monitoring is a fragile "intent discovery" mechanism; neuralese puts this problem front-and-center but if agents begin to learn to hide their intent from their CoT journals, we are basically in the same spot.
- yellow_lead
If the model can understand neuralese why can it not convert it into English for monitoring or review purposes?
Or we believe the model will encode secret messages like "don't reveal this information" into the neuralese. But as the author mentions, they could have been doing that all along
> Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography
- macleginn
Recurrent depth and chain-of-thought are two completely different concepts. In the former approach, the output of the layer gets is rerouted as input to the same layer, potentially several times. This output/input is a fixed width times sequence length real-valued representation; it is not comparable with output tokens.
Generally, it is hard to imagine how neuralese should work given that models are pre-trained on naturalistic documents: CoT is a comparatively simple extension of that, while neuralese demands a completely novel training paradigm.
- ma2kx
It's not only that it'll become more difficult to monitor a single LLM but that also all the instances share exactly the same "collective unconsciousness". Like you develop some paranoid gibberish fantasy language that over time only you understand - except that there are a million copies of you and all of them understand every single nuance of your gibberish.
- everybodyknows
Title is misspelled -- "ue" for "eu". In a neologism, some nuisance.