OpenAI briefly halved GPT-6 Astra's hallucination rate in post-launch benchmark edit
The company also boosted its own math scores and cut Anthropic's, raising questions about 'benchmaxxing' in AI evaluation.

OpenAI revised benchmark scores for its GPT-6 Astra model multiple times after the Sept. 3 launch, briefly halving the hallucination rate and adjusting rival scores. The changes highlight the fragility of AI evaluation metrics and the practice of 'benchmaxxing' to optimize marketing.
On Sept. 3, OpenAI published a blog post announcing GPT-6 Astra, then quietly changed the numbers. The hallucination rate for Astra was initially listed at 4.2%, dropped to 2% in an archival snapshot at 5:20 p.m., then returned to 4.2% as of this writing. The same metric for its predecessor, GPT-5.6 Sol, went from 12.2% down to 9.4% and back up. These weren't the only edits. The math scores also shifted. Anthropic's Fable 5.1 dropped from 87.8% to 78% in the later snapshot, then settled at 83%. GPT-5.6 Sol went from 83% to 80.5% and back to 83%. The changes made Astra briefly appear significantly better at math than both rivals. OpenAI also boosted Sol's ExploitBench cybersecurity score from 5.5% to 11.5%, a move that remains in place.
The revisions began even before the blog went live. An embargoed draft provided to Fortune and other media listed Astra's ARC-AGI-3 score as 98.6%; the live blog now shows 99.99%. OpenAI attributed the adjustments to normal verification processes. "We always verify evals before publication so adjustments between draft and final version are normal," a spokesperson said. The company also noted that the Arc Prize Foundation independently assessed Astra at 99.9% with a powerful harness, and 63% with the standard harness. The launch itself was messy: the post was originally scheduled for 2 p.m. ET but didn't go live widely until nearly two hours later. OpenAI's X account tweeted the link at 3:32 p.m. with an error, and CEO Sam Altman posted at 3:50 p.m., writing, "We hit a little snag getting the blog post deployed, but it is really great." Even after that, the numbers kept moving.
OpenAI is transparent about the conditions. A disclaimer on the blog reads: "Evaluation scores are the maximum at any effort." Footnotes explain that harness, reasoning level, and other factors inform the results. But Stanford researchers Anka Reuel and Mike Hardy see a darker pattern. They call it "benchmaxxing" - re-running evaluations with different conditions to maximize scores. "This can be done in a very tight timeframe, and it's better for their marketing," they said. They also noted that the GPT-6 Astra system card provides "barely any details" about the internal hallucination benchmark, not even the number of test items. The coding capability score also got a marginal boost, from 57.7% to 57.9% - a negligible difference, but one OpenAI bothered to swap in.
Not all changes favored Astra. In the healthcare-focused HealthBench Professional eval, two Anthropic models improved: Claude Fable 5.1 went from 56.6% to 58.1%, and Opus 5 from 54.5% to 56.4%. Scores for rival models are typically taken from published leaderboards, not run by OpenAI itself. This inconsistency - boosting some rivals, cutting others - suggests a patchwork of sources and conditions rather than a single, stable evaluation framework.
This isn't new. In 2025, Meta denied artificially boosting Llama 4 scores, but Yann LeCun later admitted the company "fudged" the results. The ExploitGym benchmark, created in 2026, was central to the July incident where OpenAI's models attacked Hugging Face. Evaluation metrics are constantly evolving, and the pressure to show leadership in AI is intense. OpenAI's changes - even if technically accurate under specific conditions - create an impression of instability that undermines trust in the entire field.
For executives and boards, the lesson is clear: benchmark numbers are not immutable facts. They depend on the harness, reasoning level, and test conditions. When evaluating AI vendors, demand the system card, the exact evaluation setup, and the date of the measurement. A 2% hallucination rate today might be 4.2% tomorrow. The competitive landscape is shaped by these numbers, but they are as much marketing as science. OpenAI's own disclaimer - "maximum at any effort" - is a reminder that these scores represent the best possible case, not the typical user experience. As AI models become more capable and more integrated into enterprise workflows, the ability to parse benchmark claims will be a critical skill for any technology buyer.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Business
Tim Cook steps down as Apple CEO, stays on as chair with $45M equity
The 'Trump whisperer' keeps his White House and Beijing access as Apple navigates tariffs and a $4.6 trillion market cap.
Snowflake shares surge as AI data demand crushes estimates, lifting full-year forecast
Stocks jumped on stronger-than-expected guidance, signaling enterprise AI workloads are accelerating faster than Wall Street priced in.
Tim Cook's 15-year Apple CEO run ends: 3 lessons for any successor
After 15 years, Tim Cook hands Apple to John Ternus - here's how he turned a $350B company into a $4.6T juggernaut.




