DeepSeek proposes an experimental multi-modal AI model – and it’s not inferior to Claude Opus 4.8

DeepSeek proposes an experimental multi-modal AI model – and it’s not inferior to Claude Opus 4.8

Chinese artificial intelligence system developer DeepSeek demonstrated an experimental multi-modal version of its V4-Flash model. It differs from the basic version in that it is able to process not only text but also images as queries.

    Photo credit: Solen Feyissa / unsplash.com

Photo credit: Solen Feyissa / unsplash.com

The experimental model DeepSeek V4-Flash-Vision-Exp can analyze images, screenshots, and other visual queries as well as text, while maintaining existing functionality, including the ability to process text, logical reasoning skills, apply knowledge of the surrounding world, and manage autonomous artificial intelligence agents.

Image source: x.com/deepseek_ai

“In multimodal drug testing, V4-Flash-Vision-Exp and [базовой] V4-Flash, improves the performance of multi-mode agents to a level close to Opus-4.8”,— point out in the company. In addition, the DeepSeek Harness 0.1.1 tool has been released, which supports the new model out of the box.

The updated model is now available through the DeepSeek API. When billed, images are converted to tokens (up to 384 tokens per token) and charged at the price of the base DeepSeek V4-Flash model.

If you find an error, select it with your mouse and press CTRL+ENTER.

Exit mobile version