OpenAI’s most capable model to date reached general availability on July 9, 2026, following an unusually public delay. GPT-5.6 Sol had originally been previewed on June 26 and was set for a broader rollout shortly after, but the US government requested a pause to allow for a fuller safety review of the model’s cybersecurity capabilities.
Reuters reported that OpenAI received US approval for the broad rollout on July 8, with the model going live to developers and ChatGPT users the following day alongside two companion models — GPT-5.6 Terra and GPT-5.6 Luna — rounding out the full 5.6 model family.
The delay itself was notable. It marked one of the first documented cases of a frontier model launch being paused at a government’s request rather than for internal safety reasons, signalling how seriously the US administration is now engaging with the output of its top AI labs.
For a deep-dive walkthrough of what GPT-5.6 Sol actually delivers at the benchmark level, this overview covers the key numbers clearly: https://www.youtube.com/watch?v=LPs-SNuOqho
Benchmarks and What They Mean

GPT-5.6 Sol’s headline numbers are strong. On BrowseComp — OpenAI’s benchmark for autonomous web navigation — it scores 92.2 percent, a new state-of-the-art result. On OSWorld 2.0, which tests computer-use agent performance, it reaches 62.6 percent, surpassing Claude Opus 4.8 while using 85 percent fewer tokens to achieve comparable task completion.
On Terminal-Bench 2.1, the coding-agent benchmark that has become the industry’s de facto standard for autonomous code execution, GPT-5.6 Sol reaches 91.9 percent in Ultra mode and 88.8 percent in standard mode, placing it at or near the top of every major leaderboard in the AI coding agent category.
The safety profile has also been substantially upgraded relative to GPT-5.5. OpenAI says GPT-5.6 Sol’s cyber safeguards block roughly ten times more potentially harmful activity than its predecessor, and prompt injection failures — one of the most exploited attack vectors on deployed AI systems — dropped by sixfold compared to GPT-5.5 in internal testing.
This improvement was largely enabled by GPT-Red, OpenAI’s new automated red-teaming system announced on July 15, which uses self-play to find prompt injection vulnerabilities at a scale and speed that human red teams cannot match.
A Model Built for Agentic Work
GPT-5.6 Sol was designed explicitly for agentic and multi-step workflows rather than single-turn conversations. It powers OpenAI’s Codex coding agent across desktop, CLI, and cloud surfaces, and it is the engine behind operator-defined autonomous task pipelines that run continuously without user input.
The model is available through the ChatGPT interface on Pro, Team, and Enterprise plans, as well as through the OpenAI API. Terra and Luna, the two companion models in the 5.6 family, are positioned as faster and more cost-efficient alternatives for lower-complexity tasks within the same architectural generation, giving developers a full-family pricing structure similar to Anthropic’s Sonnet and Haiku tier approach.
The government clearance episode has already sparked wider conversation in Washington about whether formal review processes for frontier model launches should be codified, with several lawmakers citing the GPT-5.6 delay as a template worth institutionalising.
Quick Links:

Comments
Be the first to leave a comment.