It turns out that this small AI model is 11 times cheaper than GPT-5.6 Luna and almost as good as GPT-5.6 Luna on logical thinking tests

It turns out that this small AI model is 11 times cheaper than GPT-5.6 Luna and almost as good as GPT-5.6 Luna on logical thinking tests

It Turns Out That This Small AI Model Is 11

Pathway Artificial Intelligence Laboratory, which specializes in non-transformer architectures, published the test results of the inference model BDH-CQ. Its parameter size is 150 million and it scores 29.5% on the ARC-AGI-1 benchmark. The computational cost of the test is $0.0007, which is 11 times cheaper than maintaining a small OpenAI GPT-5.6 Luna model.

    Image credit: Jackson Sophat/unsplash.com

Image credit: Jackson Sophat/unsplash.com

Pathway said the high cost of running artificial intelligence systems today is primarily due to the architecture of the model rather than its functionality, so the lab decided to demonstrate a switch to an alternative architecture “It opens up a whole new space in terms of wisdom and cost”. The well-known ARC-AGI-1 benchmark is designed to test a model’s reasoning capabilities and its ability to infer rules from a limited number of examples and apply them appropriately to new input material. The youngest model in the series, OpenAI GPT-5.6 Luna, showed a result of 34.2% in the same test, but already costs $0.008 per task – considering that OpenAI has reduced its price by 80%.

In comparison, the Anthropic Claude Opus 5 and Google Gemini 3.1 Pro performed 97-98% in the same test and cost $0.5-0.6 per task, which means they cost almost a thousand times more. Alibaba Qwen3 235B is more than three times more expensive than BDH-CQ, but has poorer performance. The gap is due to the reasoning mechanism. Most modern artificial intelligence models that support logical thinking produce intermediate text, which also consumes tokens, and only after that do they produce the final answer, while BDH-CQ “silently” solves the problem in memory.

Pathway’s first experiments with this model show that it adheres to standard scaling laws, similar to the mechanisms of Transformer models ranging in size from 1 billion to 600 billion parameters. The developers intend to increase the complexity of the tasks available for the model, including mathematical reasoning, ARC-AGI-2 testing, and ultimately a comprehensive evaluation of ARC-AGI-3. On larger, more complex tasks, efficiency advantages may still exist, too.

If you find an error, select it with your mouse and press CTRL+ENTER.

Exit mobile version