OpenAI Says Astra Can Now Find and Exploit Unknown Security Flaws—First Model to Hit 'Critical' Cyber Threshold

Path to Astra: critical capabilities and frontier safeguards

OpenAI has designated its upcoming model Astra as the first to meet the 'Critical' cybersecurity capability threshold under its Preparedness Framework, meaning it can autonomously discover and exploit previously unknown vulnerabilities in hardened systems. The company delayed parts of Astra's development and release to strengthen safeguards against cyber misuse and unauthorized actions. Astra achieved a perfect score on ExploitBench and discovered two zero-day vulnerabilities during internal testing. Access to its most advanced cyber capabilities will be initially limited to a small group of testers, with broader access through Daybreak Blue for defensive use.

It is the first model we are designating at this level, and requires stronger safeguards during development and before release.
  1. glub

    > OpenAI is committed to ensuring that the benefits of AI are broadly accessible.

    > We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1]

    So many nice-sounding words.

    Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted by its models but may not defend with the same model. And you won't find a single announcement from OpenAI about this anywhere. Pick the wrong country, get "Unable to verify", no reason, no appeal. [2]

    They revoked TAC from users who already had it, called it a technical issue, told everyone to re-verify, collected ID and face scans again (eight times in my case), and only a week later moved the block to the country selector so it fails before you upload anything.

    So now I learn that I will not have access to Astra. Great.

    Very excited about this broad accessibility and clear, objective criteria from OpenAI. This level of transparency must be studied.

    [1] https://openai.com/index/scaling-trusted-access-for-cyber-de...

    [2] https://lubaretsi.com/en/writing/openai-tac-country-gate/

  2. mentalgear

    I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra...).

    The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely with Altman's golden marketing-hype boy leadership pushing the for-profit gas pedal like this.

    Honestly, this is just pure irresponsible insanity to play with the fate of the world - basically a death race of the biggest few tech companies on the planet.

    And if you think I'm being dramatic, listen in again to ex oAI employee[0] and check for yourself how chillingly on trajectory we already are.

    [0] https://ai-2027.com/

  3. danieltk76

    Daybreak blue is definitely a good model (I think a further post trained GPT 5.6 sol). Alot of the capabilities they talk about Astra having though have been available with good harness engineering for a year now.

  4. supermdguy

    > As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.

    Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.

  5. twoodfin

    Taking as given this model meets the “Critical cybersecurity threshold” as defined by OpenAI:

    Could the Federal government use the Defense Production Act or other legal tools to compel OpenAI to deliver the un-guarded model weights for national security needs?

    Hard to believe any government would allow this level of capability to remain exclusively in private hands.

    Interesting times.

More from this day

2026-09-01