OpenAI launches GPT-6.1 Sol – almost similar to flagship GPT-6 Astra in complex tasks, but five times cheaper

OpenAI launches GPT-6.1 Sol – almost similar to flagship GPT-6 Astra in complex tasks, but five times cheaper

OpenAI announces the launch of the GPT-6.1 Sol AI model, an updated version of GPT-6 Sol that excels in agent programming, computer management, and professional applications.

    Image source: OpenAI

Image source: OpenAI

This update brings the Sol series closer to GPT-6 Astra in terms of solving complex problems, but the difference lies in more affordable tariffs – the new model’s standard price for input and output tokens is five times lower than that of Astra: $2 for input data and $10 for output. Cache input costs only $0.10 per 1 million tokens, which is 95% lower than standard input prices. This is particularly useful for proxy-based tasks that require context reuse across multiple requests.

OpenAI notes that GPT-6.1 Sol significantly outperforms GPT-6 Sol in complex professional tasks, from writing and debugging code to understanding documents and executing multi-step business processes.

For example, on the DeepSWE v1.1 benchmark, which evaluates solving complex software development problems in real-world code bases, GPT-6.1 Sol outperformed the best GPT-6 Sol by 6.4 percentage points with less inference effort and cost.

The GDP.pdf test, which measures the accuracy of models answering professional questions on complex PDF files used by professionals working in finance, healthcare, law and seven other professional fields, showed that GPT-6.1 Sol outperformed Opus 5.5 at all levels of inference work tested, with a fallback model that was more than half the cost.

In the AutomationBench test, which tests whether an agent correctly executes a multi-step business process, GPT-6.1 Sol’s average inference performance is 2.2 percentage points higher than Opus 5.5, at about one-third the cost, and 4.8 percentage points higher than GPT-6 Sol.

AutomationBench 1.0.6 Tests AI agents in end-to-end processes using 47 tools from sales, marketing, operations, support, finance, and HR.

On tasks that require interaction with computer applications (OSWorld 2.0 benchmark), GPT-6.1 Sol is 7 percentage points better than GPT-6 Sol on the maximum inference workload, while the cost is only more than half of GPT-6 Sol, only 2.1 percentage points behind Astra, and the cost is about seven times lower.

In the Terminal-Bench Science 0.1 test, which evaluates scientific workflows including data analysis, modeling, and theorem proving, GPT-6.1 Sol outperformed GPT-6 Sol, delivering more than twice the results under maximum inference effort for less than half the cost. In this test, the agent uses code and terminal tools to perform scientific research tasks: analyze data, run simulations, and select model parameters.

GPT-6.1 Sol also outperforms its predecessor in terms of practical accuracy in responses to complex queries. In very high-level inference work, its factual error rate was 4.1%, compared with 4.5% for GPT-6 Sol and close to Astra’s 4.0% under the same settings. At the same time, implementation costs are approximately 83% lower than Astra.

In tests consistent with human intentions and values, GPT-6.1 Sol showed significant improvements over GPT-6 Sol. In complex tests, this model is less likely than GPT-6 Sol to fail: it hides defective hunting tools, violates explicit restrictions, or performs unauthorized operations while performing agent tasks. There is also no record of the model trying to bypass automated security check systems.

GPT-6.1 Sol is now open to all Plus, Pro, Business, Enterprise and Edu users through ChatGPT Work and Codex. But in the regular ChatGPT chat, the new model is not available yet, but it may appear soon.

If you find an error, select it with your mouse and press CTRL+ENTER.

Exit mobile version