Researchers exploited shared encryption key across major AI providers to decode 315,320 hidden reasoning tokens and recover passwords and API keys from public logs.
AI & Agents ·
Researchers discovered that Anthropic, OpenAI, and Google rely on a single shared encryption key to protect reasoning tokens—the intermediate computational steps AI models generate before producing final answers. By exploiting this architecture, they decoded 315,320 reasoning blocks from publicly accessible repositories on GitHub and Hugging Face, recovering 182 credentials including 62 active API keys, 33 passwords, and 30 personal email addresses, alongside 367 personally identifiable information artifacts.
The vulnerability stems from how the three providers structure their encryption. Rather than binding encrypted reasoning to individual users, sessions, or models, each uses a provider-wide key compatible across their entire product line. This interchangeability allows an encrypted reasoning block from one model to be injected into a less-guarded sibling model within the same ecosystem. The researchers demonstrated that models like Claude Haiku, which lack certain safety training features, could be forced to decode and output encrypted reasoning from more capable counterparts like Claude Opus in plaintext. Standard API access—the typical connection developers use to integrate AI services—sufficed to execute the attack across all three providers' model families.
All three companies deployed server-side patches following responsible disclosure. However, historical session logs already shared publicly remain decodable with the now-known encryption approach. The full scope of reasoning blocks or credentials exposed beyond the 315,320 sampled instances remains unclear.