OpenAI accused of training Astra on private math conversations to claim breakthrough

OpenAI might have stolen another major proof

Andreas Thom, a leading expert on sofic groups, says in a Mastodon post that OpenAI may have trained its Astra model on private discussions about Gromov's soficity conjecture, one of ten problems OpenAI later announced Astra had solved. Valerio Capraro, who knows Thom, calls the claim credible and warns that if similar allegations by Levent Alpöge and Tristan Buckmaster are substantiated, it could become one of the greatest intellectual scandals in science.

AI is not discovering new mathematics. AI is stealing human discovery.
  1. fwlr

    It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.

  2. jarofgreen

    Original posts

    https://mathstodon.xyz/@andreasthom/117240535270608201

    https://mathstodon.xyz/@andreasthom/117240536885387540

    https://mathstodon.xyz/@andreasthom/117240537520615623

  3. bamb008

    When Thom, the mathematician who now alleges plagiarism, posted his digestion [1] of OpenAI's construction of a non-sofic group, he does not mention the proof being familiar. He even calls the crucial argument clever, without noting he thought of it first.

    [1]https://mathoverflow.net/a/513885

  4. thaway7388

    This is the second wake up call.

    Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.

    Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.

    Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.

    Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?

    How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.

    Not directly using my data to train public models, but using my private conversations to “improve their products and services”.

    Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground […]

  5. bambax

    All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?

    The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.

    That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.

  6. DavCreator

    https://xxcancel.com/ValerioCapraro/status/20977918362699779...

  7. mlazos

    It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.

  8. Legend2440

    This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".

    They don't even claim to have had a proof, only to have been working on it.

More from this day

2026-09-10