DeepSeek unveils an experimental multimodal model to compete with Anthropic.
DeepSeek unveiled an experimental multimodal model on Friday, claiming its agent performance is nearly equivalent to Anthropic’s Opus-4.8. According to DeepSeek’s published data, it outperforms that model on three out of eleven benchmarks.
The DeepSeek-V4-Flash-Vision-Exp is available on the company's API platform. This model enhances the text-only V4-Flash by incorporating the ability to interpret images and screenshots, allowing it to take actions based on visual input. The Hangzhou-based company also released version 0.1.1 of its agent harness, which includes built-in support.
“This experimental multimodal model matches DeepSeek-V4-Flash in text capabilities, covering agents, reasoning, and world knowledge,” DeepSeek stated on X. It further noted that, in terms of multimodal agent benchmarks, it “makes a significant advancement over V4-Flash, bringing multimodal agent performance close to Opus-4.8.”
Regarding Opus, the context matters less than it seems. Opus-4.8 was launched in May, and Anthropic announced Claude Opus 5 on July 24. TNW previously reported that Opus 5 was able to match Fable in coding while costing half as much.
Opus-4.8 remains active on Anthropic’s model deprecation page, labeled as fully supported and recommended for use, with no retirement scheduled before May 2027. Five Opus models share this active status simultaneously.
Thus, DeepSeek chose a current and supported competitor instead of an outdated one. Notably, the table does not include a column for Opus 5, and no other sources have published a comparison. How V4-Flash-Vision-Exp performs against Anthropic's latest Opus remains uncertain, and the release does not claim to clarify this.
Bloomberg reported on Friday, referring to Opus 4.8 as an advanced model from Anthropic.
Interpreting the actual data, DeepSeek published results from eleven benchmarks, outperforming Opus-4.8 in three cases.
It leads in DeepSWE by 1.3 points, Agents’ Last Exam by 1.6 points, and ZeroBench by 1.0 point. However, it falls short in the other eight benchmarks, with two showing a significant gap. For instance, NL2Repo records DeepSeek at 57.7 compared to 69.7 for Opus-4.8, a difference of 12 points. Meanwhile, DSBench-Hard has scores of 63.6 against 71.7.
The close results are genuinely competitive. Toolathlon-Verified shows 75.9 to 76.2, Chartography shows 64.3 to 65.0, and Terminal Bench 2.1 has scores of 83.9 against 85.0.
One score is noteworthy for its implications about the field rather than the competition itself. In AutomationBench, all three models score in the mid-twenties: 25.7, 25.1, and 27.2. This suggests that whatever agents excel at now, that benchmark is not reflecting it.
The claim of a leap in multimodal performance comes with a caveat, which DeepSeek specified. The reported jump in multimodal agent performance compared to V4-Flash shows 36.5 against 26.2 on ApexBench, and 27.3 against 25.2 on Agents’ Last Exam.
DeepSeek’s footnote illustrates part of this gap. It clarifies that in these two evaluations, the text-based V4-Flash “ignores multimodal elements contained therein,” meaning the older model is evaluated on tests that involve images it cannot perceive.
This doesn’t invalidate the new model’s scores. Nevertheless, the leap partly reflects what occurs when a model lacking vision is tested with visual elements. DeepSeek is transparent about this in their table, unlike many labs.
DeepSeek is downplaying one aspect of its performance. The company states that the vision model “matches” V4-Flash in text capabilities, though its own data indicates it actually surpasses that.
Across the seven text benchmarks, the vision variant outperforms the text-only model in six of them. Toolathlon-Verified shows an improvement of 5.6 points, DeepSWE by 4.9 points, and DSBench-Hard by 4.0 points. The only exception is Cybergym, where the vision model scores 75.3 compared to 76.7, indicating the addition of visual capabilities reduced its score by 1.4 points on a security benchmark.
All figures come from DeepSeek itself, which states it evaluated its models using its own harness in minimal mode, with temperature set to 1.0 and top_p at 0.95. While vendor benchmarks are standard and the settings are disclosed, they remain the vendor’s data.
Other articles
DeepSeek unveils an experimental multimodal model to compete with Anthropic.
DeepSeek's experimental multimodal model has secured victories in three out of eleven benchmarks when compared to Anthropic's Opus-4.8, according to the table published by DeepSeek itself.
