OpenAI previewed the GPT-5.6 family on Friday: Sol at the top of the stack, Terra in the middle at roughly half the price of GPT-5.5, and Luna at the bottom for high-volume work. Sol sets a new top score on Terminal-Bench 2.1 at 88.8 percent. The Ultra configuration, which spins up subagents to parallelize hard problems, gets to 91.9 percent. Claude Fable 5 sits at 83.4 on the same benchmark. On ExploitBench, OpenAI’s internal cybersecurity eval, Sol is competitive with Anthropic’s Mythos Preview at roughly one third the output tokens. These are real numbers, and the headline they would normally produce is “OpenAI back on top after six months of Anthropic eating its lunch.”

That is not the headline. The headline is that you cannot use it.

GPT-5.6 launched as a limited preview to roughly twenty trusted partners, with broader rollout “in the coming weeks.” That phrasing is doing a lot of work. According to the company, the limit is at the request of the Trump administration, which raised national-security concerns about Sol’s coding and exploit-research capabilities and asked OpenAI to share the participant list with the government before release. Commerce Secretary Howard Lutnick and Sam Altman reportedly negotiated the framework. OpenAI’s own blog post, in language that reads like it was written by a lawyer who has lost several meetings, says the company “does not believe restrictions of this kind should be the norm.”

Translation: we did this once, please do not make us do this again.

The substance of the worry is not unreasonable. Sol’s ExploitBench numbers describe a model that is meaningfully better than the current public frontier at the actual task of finding and writing software exploits. If you are the federal government and you have spent the last eighteen months watching every other frontier lab ship that capability into a worldwide API with a click-through terms-of-service, you reach the part of the job where you ask the question Lutnick apparently asked. The question is whether asking it once becomes asking it every time, which is the regulatory pattern every other dual-use industry has eventually settled into.

The thing to watch is the next launch. If Anthropic ships the next Mythos with a similar gated preview list, then this is the new shape of frontier model releases and the “20 partners and the federal government” line will read like a policy regime instead of a one-off. If Anthropic ships unrestricted and OpenAI’s next release follows, this was a one-time concession around a specific capability profile. The model that wins the next coding-benchmark cycle gets to set the precedent. The federal government just learned it can ask.

openaigpt-5-6solterralunamodel-launchterminal-benchexploit-benchtrump-administrationlutnicksam-altmanai-policy