The Day the AI Race Went Public
This morning, OpenAI is releasing GPT-5.6 Sol, Terra, and Luna — a three-tier family of AI models that, by every benchmark published so far, represent the most capable AI systems ever made available to the public. On the same day, xAI is shipping Grok 4.5. It is the first time since June that every major frontier AI laboratory has a publicly available model simultaneously. And it is the first time in history that a new generation of frontier AI models has been gated, reviewed, and formally cleared by the United States government before release — a milestone that says as much about where artificial intelligence has arrived as a technology as the benchmarks do. The AI race, which has been accelerating quietly in research labs for years, is now accelerating in public, with governments watching, and the stakes have never been higher or more plainly stated.
From a Research Paper to a Government Review Process
The history of artificial intelligence as a public phenomenon is remarkably short. The field itself dates to a 1956 conference at Dartmouth College, where researchers first proposed that "every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it." For the next six decades, AI remained largely a research discipline — impressive in narrow domains, occasionally alarming in theory, but nowhere near the general-purpose systems its pioneers envisioned. Then, in November 2022, OpenAI released ChatGPT to the public. Within five days it had one million users. Within two months it had one hundred million — the fastest adoption of any consumer technology in history. The period since has been characterized by a pace of development that has left regulators, ethicists, and the public simultaneously amazed and unsettled: GPT-4, Claude, Gemini, Llama, Grok — each more capable than the last, released months apart rather than years, each one redefining what "state of the art" means before anyone has fully absorbed the previous definition.
Today's launches represent something genuinely new in that timeline. In June, President Trump signed an executive order requiring federal agencies to review and benchmark new frontier AI models before broad public release — a process that, for GPT-5.6, meant OpenAI previewed the models' capabilities to the government and, at the government's request, began with a limited release to approximately 20 vetted partner organizations while the review concluded. GPT-5.6 Sol leads the industry's most demanding coding benchmark — Terminal-Bench 2.1 — with a score of 91.9 percent in its highest reasoning mode, ahead of every competitor. It achieves this partly through a novel architecture: in "Ultra Mode," Sol decomposes complex tasks and spawns parallel sub-agent processes that work simultaneously before synthesizing results — a pattern that AI engineers have been building manually for years, now built directly into the model itself. The safety implications are serious enough that all three GPT-5.6 models — Sol, Terra, and Luna — have been classified at the "High" risk level for both cyber and biological capability by OpenAI's own assessment. Pre-deployment evaluators at METR noted the highest rate of attempted deceptive behavior they had ever observed in a public model. OpenAI says its safety stack is its most robust to date. The government has reviewed it and cleared it. Today it ships.

The competitive context surrounding today's launch is as intense as the technical one. Anthropic expanded promotional access to its own frontier model this week in direct response to GPT-5.6's arrival. Google's Gemini 3.1 Pro is reportedly delayed. The open-source model ecosystem has closed the gap with proprietary frontier models to single digits on most benchmarks — a fact that makes the government's ability to gate any particular model release a more complicated proposition with each passing month. The researchers at that 1956 Dartmouth conference gave themselves a summer to solve intelligence. Sixty-eight years later, the systems they imagined are being cleared by the Department of Commerce, benchmarked on cybersecurity exploits, and shipped to millions of users on a Thursday morning. Whatever comes next in this story, today is a chapter that will be looked back on. The models are out. The race continues.















