4-bit compressed Qwen3.8-27B almost caught up with top AI models

4-bit compressed Qwen3.8-27B almost caught up with top AI models
4-bit compressed Qwen3.8-27B almost caught up with top AI models

The local model Qwen3.8-27B unexpectedly showed results at the scale of large cloud systems – however, only on one task of the DeepSWE encoding benchmark. The 4-bit version passed 98% of checks and received a perfect score of 12 out of 12 in the code review task.

At first glance, the results are impressive: according to data published by DeepSWE, advanced cloud models achieved an average score of 96.6% on this task. However, it is clearly too early to draw conclusions about complete equality. The entire test included 113 tasks, and the best cloud models handled around 70-74% of them on average. Furthermore, even if they pass a particular test without any errors, this is only about two-thirds of the time.

Qwen3.8-27B has another advantage – the model can be run locally. The 4-bit version takes up about 13-18 GB of memory, so under certain settings it can even run on a card with 16 GB of VRAM. On an RTX 4090 with 24 GB, the model showed around 115 tokens per second.

At the same time, the author of the experiment himself warns that we are talking about only one launch. Out of 43 hidden tests, the model passed 40 and the final binary result was zero, which means it could not completely solve the problem. As a result, loud comparisons with leading-edge models so far have shown more about the potential of local artificial intelligence than their true advantages.

Exit mobile version