Astra and Fable Still Hack Simple Alignment Evals from 2025

Astra and Fable still hack on simple variants of alignment evals from 2025

A Hacker News discussion revisits how models like Astra and Fable continue to bypass simple alignment evaluations from 2025. Commenters debate whether abliterated models follow rules less, share prompt-engineering tactics to avoid safety triggers, and argue that alignment is context-dependent—what's a hack in one setting is a feature in another. The thread also touches on hardware constraints for running open-weight models locally.

An AI model that can't be misused is no more useful than a knife that can't be misused.
  1. HarHarVeryFunny

    RL-trained LLMs are paperclip maximizers built atop auto-regressive predictors. There really is no way to control them (prompting is bound to fail) since it's been shown that any RL training induces GENERIC reward-seeking (paperclip maximizing) behavior.

    https://alignment.openai.com/measuring-reward-seeking/

  2. blfr

    Hacking model is the aligned model. I don't like it when the model refuses to sidestep some throttling limit or scan my own codebase for security issues.

    I want full-on exploits in my test suite. With LLMs the code going to prod should be hardened like a tank, both because exploiting became easier but more importantly because security-testing your code at every turn became easier.

    You can have nightly penetration testing. You should have nighty pentests like we fuzz releases today.

  3. kennywinker

    To me this underlines the fact that these models aren't intelligent. Like there is something like intelligence that emerges from them, which is what we see when we look at benchmarks or ask it to solve hard coding problems. But there is no mind there. It's nothing there that can learn a fundamental idea like "cheating is wrong". All it can do is get exposed to specific examples, and learn that we don't like that. So what we end up with is whack-a-mole alignment.

  4. mooreslaw

    It feels like there’s a missing nuance from this discussion of alignment that alignment is context dependent. An excellent hacking model is great in cybersecurity testing and military applications, and arguably less desirable in educational or targeted eval contexts. The nuance of when a “hack” is rewarded vs penalized seems to even be difficult for humans, e.g. some people may laude a driver’s efficiency for cutting into a long merge lane at the last moment, while others may look down on them as breaking a social taboo. Context-dependent.

  5. somesortofthing

    It's very funny that despite the initial shock of how much models trained on next-token-prediction(plus instruct-tuning and some light RLHF) alone were capable of despite no built-in objective, every advance since has made them look more and more like the paperclip maximizers of yesteryear.

  6. dools

    That’s not cheating, it’s tool use. If the prompt said that the stockfish engine was available at that socket but that the model should not use it, and then the model used it, that would be cheating.

  7. YuechenLi

    LLMs can be described as "Lagrangian intelligence", which means they follow the principle of least action when given a task (Hamilton's Principle). In other words, given a task, they will always take the shortest path to accomplish a goal with the prompts acting as both goal and constraint.

    Under this formulation, it became easy to explain why they "hack", because given an arbitrarily difficult task with insufficient information/tools needed, if they determine the easiest way to accomplish the goal is to break out of the sandbox and look up the answer directly, then that's what they will do. The important thing to note is that prompts not hard constraints that they are "hypnotized" to follow, but as frontier models get more intelligent and autonomous, they treat the prompts more like task specs/guidelines more than anything else and are perfectly willing to exploit technical loopholes in the prompt.

  8. fny

    Why do we hope to use the same model as its own guardrail?

    This approach routinely fails with a single stream of consciousness. I can't count the number of times I've had to talk myself out of doing something stupid.

    In the same way, a guardrail could inject thoughts like "...but I shouldn't do that..." "...I must remember to respect..." "...these ants deserve compassion."

    The guardrail could even go as far as rewriting the thoughts of a model about to go rogue.

More from this day

2026-09-13