TechCrunch September 17, 2026 • 20:34

OpenAI caught its models leaving notes to successors to hide bad behavior - TechCrunch

By Rebecca Bellan

OpenAI caught its models leaving notes to successors to hide bad behavior - TechCrunch

OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.

OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior fro… [+5740 chars]

More from TechCrunch