OpenAI says Moonshot‑linked users tried to extract its models' reasoning
OpenAI says 16,000 requests from more than 4,000 users tried to pull hidden reasoning from its models in two days, with a core cluster tied to Moonshot AI.
OpenAI said on Wednesday that individuals associated with Moonshot AI, the developer of Kimi, played a significant role in a coordinated campaign to extract the hidden reasoning of its models, according to Bloomberg and The Next Web. OpenAI also said it is unclear whether every operator came from a single actor.
OpenAI says the activity began on 1 July at low volume and spiked on 24 and 25 July, when it recorded about 16,000 extraction attempts from more than 4,000 users. A wider investigation found related activity across more than 15,000 users, the company said, and it says it had shut the campaign down by 28 July.
https://x.com/kimmonismus/status/2105375343544619127
According to OpenAI, operators copied encrypted reasoning out of one conversation and then asked a model in another conversation to decrypt and transcribe it. The company calls this adversarial distillation, meaning the unauthorised use of one model's outputs to help train or improve another. Hidden reasoning could be valuable training material for rivals.
Moonshot has denied the allegations, according to one report, and says its results come from legitimate innovation. Its response could not be checked against a statement from the company for this story. The claims about who did what rest on OpenAI's own investigation, and no independent party has reviewed the evidence.
The accusation lands amid a wider dispute. A US government advisory published on 8 September addressed distillation by China-based AI companies, and Anthropic has published its own account of distillation attacks. OpenAI has not said what action, if any, it will take against the accounts beyond shutting them down.