MWS Cloud, part of MTS Web Services, analyzed the computing infrastructure requirements for running the world’s leading Russian large language model.
A professional review of the popular open models released that spring and summer. In 2025, Kimi K2, gpt-oss-120b, Gemma 3 27B, GLM-4.5V, Qwen3 235B and Russian Cotype Pro 2 are considered. 2026 – Kimi K3, Qwen3.5 397B, DeepSeek-V4-Pro, DeepSeek-V4-Flash, GLM-5.2, Gemma 4 and Cotype Pro 3. The study shows that the average cost of running a model on the lowest NVIDIA accelerator group increased by 2.8 times from 13.7 million rubles. In 2025 it will reach 38.8 million rubles. 2026.
New AI models are becoming more and more powerful: they are better at handling complex problems and can process more information in a single request. However, to do this, they require more memory and faster GPUs. At the same time, the cost of the graphics card itself is also increasing, so models are more expensive to deploy, the findings in MWS Cloud explain.
Source: MWS Cloud/mws.ru
MWS Cloud experts said that Russian GPU prices largely depend on the global market: The global competition in the field of artificial intelligence increases the demand for computing power and supports the growth of accelerator costs. Therefore, it becomes more expensive to deploy large models on your own infrastructure, but there is no single price limit, after which the market will lose interest in artificial intelligence: if buying the equipment no longer pays off, companies will choose more compact models, optimize calculations, or move to the cloud, where they can only pay for the capacity they actually use.
source:










