DeepSeek unveils an experimental multimodal model to compete with Anthropic.

DeepSeek unveils an experimental multimodal model to compete with Anthropic.

      DeepSeek unveiled an experimental multimodal model on Friday, claiming its agent performance is nearly equivalent to Anthropic’s Opus-4.8. According to DeepSeek’s published data, it outperforms that model on three out of eleven benchmarks.

      The DeepSeek-V4-Flash-Vision-Exp is available on the company's API platform. This model enhances the text-only V4-Flash by incorporating the ability to interpret images and screenshots, allowing it to take actions based on visual input. The Hangzhou-based company also released version 0.1.1 of its agent harness, which includes built-in support.

      “This experimental multimodal model matches DeepSeek-V4-Flash in text capabilities, covering agents, reasoning, and world knowledge,” DeepSeek stated on X. It further noted that, in terms of multimodal agent benchmarks, it “makes a significant advancement over V4-Flash, bringing multimodal agent performance close to Opus-4.8.”

      Regarding Opus, the context matters less than it seems. Opus-4.8 was launched in May, and Anthropic announced Claude Opus 5 on July 24. TNW previously reported that Opus 5 was able to match Fable in coding while costing half as much.

      Opus-4.8 remains active on Anthropic’s model deprecation page, labeled as fully supported and recommended for use, with no retirement scheduled before May 2027. Five Opus models share this active status simultaneously.

      Thus, DeepSeek chose a current and supported competitor instead of an outdated one. Notably, the table does not include a column for Opus 5, and no other sources have published a comparison. How V4-Flash-Vision-Exp performs against Anthropic's latest Opus remains uncertain, and the release does not claim to clarify this.

      Bloomberg reported on Friday, referring to Opus 4.8 as an advanced model from Anthropic.

      Interpreting the actual data, DeepSeek published results from eleven benchmarks, outperforming Opus-4.8 in three cases.

      It leads in DeepSWE by 1.3 points, Agents’ Last Exam by 1.6 points, and ZeroBench by 1.0 point. However, it falls short in the other eight benchmarks, with two showing a significant gap. For instance, NL2Repo records DeepSeek at 57.7 compared to 69.7 for Opus-4.8, a difference of 12 points. Meanwhile, DSBench-Hard has scores of 63.6 against 71.7.

      The close results are genuinely competitive. Toolathlon-Verified shows 75.9 to 76.2, Chartography shows 64.3 to 65.0, and Terminal Bench 2.1 has scores of 83.9 against 85.0.

      One score is noteworthy for its implications about the field rather than the competition itself. In AutomationBench, all three models score in the mid-twenties: 25.7, 25.1, and 27.2. This suggests that whatever agents excel at now, that benchmark is not reflecting it.

      The claim of a leap in multimodal performance comes with a caveat, which DeepSeek specified. The reported jump in multimodal agent performance compared to V4-Flash shows 36.5 against 26.2 on ApexBench, and 27.3 against 25.2 on Agents’ Last Exam.

      DeepSeek’s footnote illustrates part of this gap. It clarifies that in these two evaluations, the text-based V4-Flash “ignores multimodal elements contained therein,” meaning the older model is evaluated on tests that involve images it cannot perceive.

      This doesn’t invalidate the new model’s scores. Nevertheless, the leap partly reflects what occurs when a model lacking vision is tested with visual elements. DeepSeek is transparent about this in their table, unlike many labs.

      DeepSeek is downplaying one aspect of its performance. The company states that the vision model “matches” V4-Flash in text capabilities, though its own data indicates it actually surpasses that.

      Across the seven text benchmarks, the vision variant outperforms the text-only model in six of them. Toolathlon-Verified shows an improvement of 5.6 points, DeepSWE by 4.9 points, and DSBench-Hard by 4.0 points. The only exception is Cybergym, where the vision model scores 75.3 compared to 76.7, indicating the addition of visual capabilities reduced its score by 1.4 points on a security benchmark.

      All figures come from DeepSeek itself, which states it evaluated its models using its own harness in minimal mode, with temperature set to 1.0 and top_p at 0.95. While vendor benchmarks are standard and the settings are disclosed, they remain the vendor’s data.

Other articles

I used the Magic Capture feature on my cat with the Pixel 11 Pro, and I adore every funny shot. I used the Magic Capture feature on my cat with the Pixel 11 Pro, and I adore every funny shot. Magic Capture automatically extracted images from videos of my rather uncooperative cat, and even this initial rough test demonstrated the effectiveness of Google's new Pixel camera concept. Wave Browser Connects Daily Surfing to Ocean Cleanup Efforts Wave Browser Connects Daily Surfing to Ocean Cleanup Efforts The majority of individuals select a web browser based on practical factors such as speed, familiarity, compatibility, and convenience. The environmental impact is often not a significant part of this choice. However, Wave Browser is aiming to make it a consideration for users. Created by Eightpoint, Wave is a Chromium-based browser that merges typical browsing capabilities with features designed to assist […] Mark Zuckerberg purchased a castle, while the rest of us settle for a sandcastle once a year. Mark Zuckerberg purchased a castle, while the rest of us settle for a sandcastle once a year. Mark Zuckerberg purchased a 440-acre estate in Ireland during the week Meta faced trial regarding issues related to children. However, the castle itself is not the issue at hand. ChatGPT can now access your Apple Messages conversations on Mac. ChatGPT can now access your Apple Messages conversations on Mac. The recent Mac update for ChatGPT incorporates integration with Apple Messages, enabling users to import discussions from Apple's messaging application into the AI assistant. RayNeo's latest iO Smart Glasses introduce a compact display within your field of vision. RayNeo's latest iO Smart Glasses introduce a compact display within your field of vision. RayNeo's iOS Smart Glasses integrate a compact MicroLED display into your view, providing features such as real-time translation, dictation in 112 languages, and an AI memory log, all priced at $479. Nothing’s new CMF earbuds provide noise cancellation at a price point similar to that of Apple’s wired EarPods. Nothing’s new CMF earbuds provide noise cancellation at a price point similar to that of Apple’s wired EarPods. CMF has introduced its most affordable earbuds to date, featuring ANC, spatial audio, and extended battery life, but U.S. consumers currently do not have access to them.

DeepSeek unveils an experimental multimodal model to compete with Anthropic.

DeepSeek's experimental multimodal model has secured victories in three out of eleven benchmarks when compared to Anthropic's Opus-4.8, according to the table published by DeepSeek itself.