The AI revolution is focusing on the incorrect priorities.
TL;DR: Frontier AI models are improving in coding and agent-based tasks but declining in general writing skills. Maz Ahmadi, founder of Wizard Labs, highlights a measurable drop in performance on client-specific writing benchmarks with each model upgrade. Despite 44% of organizations scaling AI across their operations, only 37% have seen a positive impact on EBIT (McKinsey), widening the gap between AI implementation and actual value. His advice includes creating custom evaluation frameworks before development and avoiding model selection based solely on public benchmarks.
Currently, every large language model is worsening in writing abilities, and few are monitoring this decline. The industry is focused on identifying the smartest model, but the more pressing question is: smartest for what purpose? As major tech firms pursue profitable enterprise contracts, they are fine-tuning frontier models for coding, logic, and autonomous agent tasks, causing general prose capabilities to be neglected. My team's observations revealed a concerning regression in our client-based benchmarks, indicating that the writing performance of upgraded models is deteriorating.
Although this might seem like a niche issue for engineers and content creators, it's significant. AI has integrated into the daily operations of millions. Almost 90% of companies are using AI regularly for various business functions, such as drafting reports, customer service, legal document preparation, code writing, research analysis, and decision-making. When the foundational models are optimized for different performance metrics, everyone faces the repercussions.
The ongoing conversation surrounding AI text watermarking adds another layer to this topic. Recently, Anthropic announced plans for future Claude models to include watermarks on generated text to comply with the EU AI Act. This method relies on statistical patterns in token selection to identify text likely produced by Claude. Anthropic claims this approach won't negatively affect quality, creativity, or readability, citing relevant research.
Nevertheless, enterprise users should consider the implications of a model optimized for regulatory compliance, agent performance, coding, and reasoning. I've encountered evidence in our work showing that model upgrades can worsen results in specific writing tasks. While watermarking might be innocuous on its own, the larger trend merits attention. If frontier labs focus on benchmarks that drive adoption and revenue, a model could technically improve while becoming less suitable for a specific business application.
This is the paradox that enterprises are starting to face. 44% of organizations now report widespread AI use in their operations, an increase from 38% the previous year. However, only 37% have seen any impact on enterprise-level EBIT from AI, even though 80% believe AI has enhanced individual productivity. The technology is proliferating faster than the associated value.
I frequently observe companies making the same error. They inquire, “What is the best model?” and select the one that excels in public benchmarks, establishing a system around it and expecting it to function like traditional software development. However, AI operates differently.
A model considered the best general-purpose large language model may underperform in a company’s specific workflow. The only relevant evaluation must come from the task itself. This necessitates developing custom evaluation suites before any development starts, measuring models against the company's actual demands, and consistently testing outcomes as models evolve.
This is where enterprise AI requires a new discipline. Begin with the business challenge, rigorously evaluate technology against real-world effectiveness, and continually refine it until it aligns economically with the workflow, ensuring that the technology complements the business processes before widespread scaling.
A more promising future can emerge in which companies build AI leveraging their proprietary data and expertise, providing employees with decision-support systems that disseminate specialized knowledge throughout the organization. The risk lies in ceding critical judgment to generic systems, interpreting benchmark scores as indicators of competence, and realizing too late that the technology was designed around another's success criteria.
I believe the future will favor businesses willing to challenge existing norms. This may also involve reassessing the belief that every enterprise should depend on the largest proprietary model. Increasingly capable open-weight models can be hosted on an organization’s chosen infrastructure. The common question regarding the safety of models developed in China often revolves around geopolitical issues. However, the more crucial technical questions focus on where the model operates, who oversees the infrastructure, what data exits the environment, and the security measures in place.
Competitive advantage will stem from making these decisions deliberately. The focus of the AI revolution has shifted from a race for raw intelligence to an engineering discipline centered on fit. Future market leaders will not simply thrive by deploying the latest foundational models; rather, they will succeed by clearly defining operational needs, employing stringent benchmarks, and exercising the strategic discernment necessary to replace an acclaimed, celebrated model with a less flashy alternative that offers better results.
Executives should refrain from asking AI vendors which model is superior. They should request evidence of which model is most effective for their specific business needs. This shift could transform AI from a technology procurement endeavor into what it truly is: a protracted test of competitive survival.
Other articles
The AI revolution is focusing on the incorrect priorities.
As frontier labs enhance their focus on coding and agents, the overall quality of general prose is deteriorating. Maz Ahmadi, the founder of Wizard Labs, contends that companies ought to shift their approach from determining which model is superior to demonstrating which one is most effective for their business.
