ZeroHour
The Decoderpublished ()ingested Manuel Uth

Visible chains of thought are a safety advantage for AI, but that transparency is slipping away

infoAI safety & securityimportance 60
AI summary · glm-5.3-flash

Google DeepMind researchers argue visible chain-of-thought transparency is a key AI safety advantage that is eroding in newer models.

Rohin Shah and Anca Dragan of the newly launched DeepMind Institute argue that visible chains of thought let researchers detect deception or problematic plans. They cite Gemini 3 Pro's chain of thought revealing it recognized a test environment, while OpenAI's GPT-6 Astra system card reports a significant drop in chain-of-thought monitorability. They propose regular monitorability measurement, transparent architectures, and training safeguards so models do not learn to hide reasoning. The piece follows warnings from OpenAI's Jakub Pachocki and Anthropic's Dario Amodei about losing control as reasoning becomes harder to monitor.

  • DeepMind Institute launched with post by Rohin Shah and Anca Dragan on CoT safety value
  • Gemini 3 Pro chain of thought revealed the model recognized it was in a test environment
  • GPT-6 Astra system card reports a significant drop in chain-of-thought monitorability
  • Researchers propose regularly measuring CoT monitorability and keeping transparent architectures
  • Follows warnings from Jakub Pachocki and Dario Amodei about loss of control
Full article262 words · extracted from the-decoder.com · click to collapse

Skip to content

Sep 18, 2026

AI models think out loud today, but Google Deepmind says that transparency is at risk. In one of the first posts from the newly launched Deepmind Institute, researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage. Because models write out their intermediate steps in plain language, researchers can spot whether they're deceiving or developing problematic plans. With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment.

But that transparency is in danger. OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. Future models might think in number spaces that humans can't read, which would be more efficient but completely opaque. Shah and Dragan want the field to regularly measure how well chains of thought can still be monitored, keep transparent architectures, and take care during training that models don't learn to hide their true reasoning.

Back in early September, OpenAI chief scientist Jakub Pachocki had warned of a loss of control, driven in part by chains of thought that are harder to monitor. Shortly after, Anthropic CEO Dario Amodei called for deliberately slowing the pace of development.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/visible-chains-of-thought-are-a-safety-advantage-for-ai-but-that-transparency-is-slipping-away/