The AI revolution is prioritizing the incorrect aspects.
TL;DR: Frontier AI models are improving in coding and autonomous tasks but declining in general writing skills. Maz Ahmadi, founder of Wizard Labs, reports noticeable regressions on client-specific writing benchmarks with each model upgrade. While 44% of organizations are implementing AI on a large scale, only 37% observe a positive EBIT impact (McKinsey), highlighting a growing disparity between deployment and value. His recommendation is to create custom evaluation frameworks before development starts and to avoid choosing models based solely on public benchmarks.
Currently, most existing large language models are performing worse in writing, yet few are tracking this decline. The industry is focused on which model is the most intelligent, but a more pertinent question is: intelligent for what purpose? As major tech companies seek lucrative enterprise deals, they are enhancing frontier models for coding, logic, and autonomous workflows, leaving general writing as a secondary priority. My team's observations on client-specific benchmarks show a troubling regression in writing performance amid model upgrades, which should concern anyone utilizing AI for communication.
This issue might seem niche for engineers and content teams, but it is significant since AI is integrated into the daily tasks of millions. Almost 90% of companies use AI regularly in various business functions like report drafting, customer service, legal document preparation, coding, research analysis, and decision-making. If the models are being fine-tuned for different performance metrics, the consequences will impact everyone.
The ongoing discussion about AI text watermarking adds another layer to this issue. Recently, Anthropic revealed plans for future Claude models to watermark generated text in compliance with the EU AI Act. This technique employs statistical patterns in token selection to identify text likely generated by Claude. Anthropic claims this method won’t affect quality, creativity, or readability, citing supporting research.
However, enterprise users should consider the implications of optimizing a model for regulatory compliance, agent performance, coding, and reasoning simultaneously. In our own experience, we have seen improved model versions yield worse results for certain writing tasks. While watermarking may be harmless on its own, the broader trend warrants examination. If frontier labs prioritize benchmarks that stimulate adoption and revenue, a model can improve objectively while deteriorating for specific business needs.
This is the paradox organizations are beginning to face. Currently, 44% of organizations report that AI is scaling enterprise-wide, an increase from 38% last year. Nevertheless, only 37% note any impact on EBIT at the enterprise level, although 80% claim that AI has enhanced individual productivity. The technology is advancing faster than the value it generates. A common mistake is for companies to inquire, “Which model is best?” They often choose the one that leads public benchmarks and build their systems around it, expecting the project to function like conventional software development, but AI operates differently.
The leading generalized model might perform poorly for a specific company’s workflow. The only relevant assessment is the task itself. This necessitates the creation of custom evaluation suites before development begins, aligning the models with the company’s actual needs, and continuously testing these outcomes as models evolve.
A new discipline is required for enterprise AI, beginning with the business problem itself, rigorously testing technology against real-world performance, and continuously refining it until the economics and workflow align well, ensuring the technology enhances how the business operates before scaling it organization-wide.
A brighter future could emerge where companies build AI systems around their proprietary data and expertise, providing employees with decision-support systems that disseminate specialized knowledge across the organization. The risk lies in relying on generic systems for critical decisions, equating benchmark scores with competence, and realizing too late that the technology was optimized according to someone else’s success criteria.
I believe the future will favor companies that challenge standard assumptions. This may also entail reconsidering the belief that every enterprise must utilize the largest proprietary model. Open-weight models are increasingly capable and can be operated within a company’s chosen infrastructure. Enterprises often raise concerns about the safety of models developed in China within a geopolitical context, but the more pressing technical question is where the model is executed, who controls the infrastructure, what data is transmitted outside the environment, and what security measures are in place.
A competitive advantage will stem from making these choices deliberately. The nature of the AI revolution has evolved from a quest for sheer intelligence to a focus on engineering fit. Future market leaders won’t succeed merely by implementing the latest foundational models. Instead, those who clearly define operational needs, enforce robust benchmarks, and strategically replace hyped models with less recognized alternatives that yield better results will thrive.
Executives should refrain from asking their AI vendors which model is superior. They should demand proof of which model best suits their business. This change could transform AI from a technology procurement process into what it truly is: a long-term challenge for competitive survival.
Other articles
The AI revolution is prioritizing the incorrect aspects.
As frontier labs focus on developing coding and agents, the overall quality of general prose is diminishing. Maz Ahmadi, the founder of Wizard Labs, contends that companies should shift their focus from determining which model is the best to demonstrating which model works best for their business.
