OpenAI announced the release of a new artificial intelligence model, GPT-6 Astra, calling it a revolutionary system rather than an ordinary update; the company’s president, Greg Brockman, even hinted at the arrival of the era of strong artificial intelligence (AGI), but independent testing did not show any outstanding performance.

Image credit: Brecht Corbeel / unsplash.com
Based on the company’s performance Manual analysis The OpenAI GPT-6 Astra test series performs on par with its predecessor, worse than the leading Anthropic model, and even worse than the recently released Meta✴ Spark of Muse 1.3. Manual analysis tests the model using a single open method, without developer involvement, and generates ratings based on these tests: Intelligence Index includes tasks in mathematics, natural sciences, programming, long document analysis, and factual knowledge; Coding Agent Index score reflects the model’s performance in the corresponding agent system Codex, Claude Code, or Muse Code. Record token consumption, task costs, and work speed.
In the Intelligence Index, the OpenAI GPT-6 Astra model scored 61 points – the same as the previous GPT-5.6 Sol. For comparison, Anthropic Claude Fable 5.1 scored 66 points, Meta✴ Muse Spark 1.3 – 62 points. In the Coding Agent Index, the new GPT-6 Astra in the Codex environment scored 67 points, roughly on par with Claude Opus 5 and Fable 5 in Claude Code and Muse Spark 1.3 in Muse Code, with the lead held by Fable 5.1 in Claude Code’s 70 points. Compared with GPT-5.6 Sol, the new GPT-6 Astra price has increased by 2.5 times: from $4 to $10 for 1 million input tokens and from $20 to $50 for 1 million output tokens; the 90% discount for reading from the cache and the 25% premium for writing to the cache remain unchanged.
Image source: artificialanalysis.ai
The price increase compensates for a significant increase in token efficiency: GPT-6 Astra generates approximately one-third the volume of GPT-5.6 Sol and approximately one-fifth the volume of Claude Opus 5. When using the code, this means that, at best, it costs the same to use as its predecessor, while the new model is rated two points higher and by the same metrics it becomes half the price of the Claude Fable 5. In the Intelligence Index, the results are somewhat different: Astra saves only 10% of the output tokens compared to its predecessor, which only partially compensates for the price difference; in mission terms, it is 75% more expensive than Claude Fable 5.
AA-Omniscience testing showed that the frequency of hallucinations decreased from 92% to 51%. In the AA-Briefcase test, the GPT-6 Astra model started to cope better with large projects involving thousands of source files and gained about 80 Elo points, but it still fell behind GPT-5.6 Sol in terms of working demonstration quality. But in the GDPval-AA v2 test, which focuses on real-world questions from 44 professions, she lost 80 Elo points. The model scored a 6 on Humanity’s final exam, which focuses on math, coding, and humanities; it also showed a drop of two to three points on the τ³-Banking (customer support), SciCode (Python tasks in the natural sciences), and AA-LCR (large file analysis) tests.
As a result, the new OpenAI GPT-6 Astra has become more profitable in programming tasks, but has not yet been able to beat Anthropic’s competitors in the field of pure intelligence.
If you find an error, select it with your mouse and press CTRL+ENTER.
