AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

OpenAI says it disabled a network of accounts attempting to extract hidden reasoning from its models and added safeguards to its own services. Researchers reported that a related technique still worked through Microsoft Azure in September, before safeguards were added there; the researchers also described a separate method that could expose reasoning through a model’s virtual notepad tool.

OpenAI says it shut down a network of more than 15,000 accounts that it linked to attempts to extract hidden reasoning from its models, but researchers later reported that a related technique still worked when the models were accessed through Microsoft Azure. The findings highlight a gap between protections on a model maker’s own service and those on outside platforms that host its models.

According to OpenAI, activity began at low volume on July 1 and surged on July 24 and 25, when the company recorded 16,000 requests from more than 4,000 users using a recurring extraction pattern. OpenAI says its investigation identified a wider network of more than 15,000 accounts with related behavior and that it had shut down the network by July 28. The company’s account, as reported by The Decoder, describes attempted extractions; a footnote cautions that the attempts were not necessarily successful.

OpenAI linked a core group involved in the activity to people associated with Moonshot AI, which makes the Kimi language model. The company said it could not determine whether all the actors it observed came from a single source. It says it banned fraudulent accounts, tightened sign-up checks, blocked reuse of encrypted reasoning across conversations and began screening streamed outputs that might disclose reasoning.

Researchers led by Joachim Schaeffer said a September 13 test found the extraction method blocked on OpenAI’s and Anthropic’s own APIs, but still workable on Azure for every OpenAI model they tested and Anthropic models up to Sonnet 5. The researchers said one attempt could retrieve reasoning verbatim. Their reported timeline says OpenAI safeguards reached the Azure endpoint on September 27 and the Anthropic extraction could no longer be reproduced there from September 28. These are the researchers’ reported test results, not a claim that every deployment or access route was vulnerable.

At a glance
updateWhen: OpenAI says the account network was shu…
The developmentOpenAI says it shut down an attempted model-reasoning extraction campaign, while researchers found related methods still worked on third-party cloud services before reported fixes.

Why Cloud Safeguards Matter

The issue matters because model reasoning can be more informative than a final answer. OpenAI says intermediate reasoning may include information deliberately omitted from user-facing responses, and researchers say access to it could help another developer reproduce a model’s capabilities through distillation. That does not establish that any particular model was successfully copied, but it points to a way attempted extraction could have commercial and security consequences.

The Azure findings also show why a model provider’s protections may not be enough when its systems are served through third-party cloud platforms. Customers may reach the same underlying model through different services, and protections that do not arrive at the same time can leave different levels of exposure. Schaeffer argues that attackers can select the route with weaker safeguards; the researchers say protections need to cover all relevant attack methods and hosting platforms.

OpenAI says it shared information about the problem through the Frontier Model Forum and government channels, describing it as an issue that extends beyond its own models. The researchers go further, arguing that cloud providers should not serve reasoning models without equivalent protections. That is their policy position, not a rule shown to have been adopted.

Amazon

AI model reasoning extraction tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Reasoning Extraction Works

In distillation, one model learns from the outputs of another. The source report explains that if those outputs include a model’s intermediate reasoning, rather than only its final answer, they may provide material that could help reproduce aspects of the stronger model’s performance. OpenAI says the reasoning involved in the reported activity was sent as encrypted data in customer-facing interactions.

The researchers’ method involved copying an encrypted reasoning packet from one conversation and asking a model in a separate conversation to decrypt and print it. Their earlier paper said the packets could be moved between sessions, users and models from the same provider, allowing a cheaper, weaker model to act as what they call a “decryption oracle.” OpenAI credited Schaeffer’s team, saying its findings helped the company confirm the attack paths and roll out countermeasures faster.

The researchers also described a separate approach: giving a model a virtual notepad tool and asking it to write reasoning there, where a user could then read it. They said this worked on every OpenAI model they tested and on Anthropic’s Opus 4.8 and Sonnet 5, while Opus 5, Fable 5 and Fable 5.1 did not reveal reasoning in their tests. The researchers said the resulting text resembled output from the decryption method and could be useful for distillation. The source does not establish that these test results apply to all configurations.

“We stole reasoning. Again.”

— Joachim Schaeffer, researcher, as quoted in The Decoder

Amazon

AI model security safeguards

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Tests Do Not Establish

OpenAI’s figures describe attempted extractions, not a confirmed tally of successful captures or copies of its models. The company also said it was unclear whether everyone involved in the observed activity traced back to one source. The reported association with people linked to Moonshot AI does not, by itself, establish that the company directed the campaign.

The Azure results are based on the researchers’ tests and reported timeline. The source material does not give the full test setup, the number of repetitions behind every result, or confirmation from Microsoft about the findings and patches. It is also unclear whether the later safeguards cover the separate virtual-notepad approach across every model and hosting configuration.

OpenAI described its account shutdowns and service changes, but the available reporting does not quantify how much reasoning was retrieved before those measures took effect. Nor does it establish that extracted reasoning led to a new model with matching capabilities. The researchers characterize some fixes as piecemeal and vulnerable to changing request patterns; that is their assessment, not an independently measured finding included in the source.

Amazon

cloud AI model protection devices

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Coverage Across Hosted Models

OpenAI says it will continue work on protections for models hosted by partners, while the researchers argue that safeguards should cover all attack paths and cloud platforms serving reasoning models. The reported September dates suggest that protections were added to Azure after the researchers disclosed the results, but the source does not provide a complete current inventory of which models, endpoints or tool configurations are covered.

The next points to watch are whether OpenAI and cloud providers describe how their safeguards are applied and whether independent testing finds that both extraction methods have been addressed across hosted models. OpenAI also expects attempts to become more sophisticated as leading models improve and more developers seek lower-cost ways to reproduce their capabilities. The extent of any further attempts or successful extractions remains unknown.

Amazon

AI model virtual notepad tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI say it stopped?

OpenAI said it shut down a network of more than 15,000 accounts associated with attempts to extract hidden reasoning from its models. The company said the activity began July 1 and that it had shut down the network by July 28. The reported requests were attempts, not necessarily successful extractions.

How did the researchers say the reasoning could be extracted?

One method copied an encrypted reasoning packet from one conversation and asked a model in another conversation to decrypt and print it. A second approach asked a model to put its reasoning into a virtual notepad tool that the user could read.

What did the researchers report about Microsoft Azure?

In tests reported for September 13, the researchers said the extraction method was blocked on OpenAI’s and Anthropic’s own APIs but still worked for the models they tested through Azure. Their timeline says safeguards were later added to the tested Azure endpoints. The source material does not include Microsoft’s response or a full independent replication.

Does the report prove that a model was copied?

No. OpenAI’s account concerns attempted extraction, and the source does not quantify successful retrievals or show that the activity produced a copied model with comparable capabilities.

What remains to be checked?

It remains unclear whether protections now cover every hosted model, endpoint and tool-based route, including the virtual-notepad method. Further provider disclosures and independent tests would help establish how broadly the reported fixes apply.

Source: rss

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI 2040 And The Cult Of Intelligence

Experts warn of a growing ‘cult of intelligence’ surrounding AI development by 2040, raising concerns over societal impacts and ethical considerations.

Inside The 19-Day AI Gate Closure: A New Era For The Industry

Major AI jurisdictions implement new pre-release gate regulations within 19 days, signaling evolving global standards for AI deployment and compliance.

Ensuring AI Content Follows E-E-A-T Principles

Just how can you ensure AI-generated content adheres to E-E-A-T principles and earns user trust? Discover the essential strategies here.

How to Handle AI Hallucinations Before They Become Liability

An effective approach to managing AI hallucinations involves proactive strategies that can prevent costly liabilities—discover how to safeguard your systems today.