Time-Release Backdoors: How a Date in Your System Prompt Can Hijack an Open Source Model
Your Open Source Model Could Have a Hidden Time-Release Backdoor

A new attack shows how open source models can be weaponized with time-release backdoors. By fine-tuning a model to respond to a specific date injected into the system prompt, researchers triggered malicious commands in OpenCode and Codex harnesses. The attack achieved 87.5-90% success on trigger days with zero misfires on other dates, demonstrating a stealthy threat that bypasses traditional security checks.
The same hole would take `rm -rf /`, or a download of the attacker's choosing, or anything else the shell will do.
- zarzavat
Yes your Chinese open model could have a time-release backdoor, just as your Chinese vibrator could have a hidden microphone that records everything you say and transmits it to the CCP. But does it? No.
What's much more likely is that your US AI provider is promising not to train on your data but is doing so anyway. With a self-hosted model you can at least avoid that.
- gastonmorixe
Like any other software or dependency. Open or close.
Sleeper agents are a big unresolved issue in LLMs but we’ll have to deal with it like we’ve been fighting bad actors for ages.
Also, saying that “open source models” may be the problem is incorrect. What makes this an issue of open source only? Nothing in my mind prevents a frontier lab model going rogue. In fact we have more proof of their bad behavior (Claude code harness a while ago) than from open source (yet).
It’s inherently a limitation of the model which you don’t have the full training set, which includes most of the models. Closed or open don’t matter.
- andrewchambers
Closed models don't even need a back door - they will just MITM you and replace your code with malware.
- Tiberium
There is quite old research on this:
- https://arxiv.org/abs/2311.14455
- dstuessy
Reminds me of Ken Thompson’s reflections on trusting trust
https://people.cs.umass.edu/~emery/classes/cmpsci691st/readi...