The writer is placing their bets on more affordable AI agents through the Palmyra X6 and a streamlined harness.
Writer, the enterprise AI firm, has introduced its latest flagship model, Palmyra X6, along with an enhanced “harness” designed to achieve something the industry has hesitated to promise: reduced token expenditure.
This model is a post-training adaptation of Z.ai’s open-source GLM-5.2, and it is available to Writer’s clients starting today. The announcement comes at a tense time, as the proliferation of agentic AI is increasing the steps—and consequently the tokens—needed for every task, leading to higher enterprise costs.
Writer is banking on a reaction to these costs, and they are not alone in recognizing the potential, with Baseten securing $1.5 billion based on the idea that AI profitability relies on inexpensive inference.
According to Writer, a harness serves as the orchestration layer surrounding a model, determining how a multi-step agent effectively carries out each request. By optimizing this layer to reduce unnecessary calls and excessive context that agents often gather, token consumption can be minimized without altering the model itself.
Writer claims the potential savings are significant, estimating that the combination of the new model and harness can reduce costs by up to 50% for basic tasks, and a Writer research paper found that efficiency changes in the harness alone lowered costs by approximately 40% on average during testing.
The interesting aspect of the harness is its model-agnostic nature; it is compatible not only with Writer’s own models but also with external ones via Microsoft Azure and Amazon Bedrock. This means the efficiency benefits are not confined to a single vendor’s offerings.
“The harness is the one component whose efficiency scales across every model an organization employs, both now and in the future,” write Writer’s researchers, suggesting that businesses have overlooked the potential of orchestration over raw model quality.
This point is notable, as Writer asserts that major labs have little motivation to help clients reduce spending, since their revenue increases with every token used. This distrust is clearly reflected in the company’s messaging.
CEO May Habib expressed it plainly: “The enterprise is absolutely sick of chasing the next benchmark,” she stated. “They want to flatten costs.” This represents a significant shift from the conventional sales narrative, where each new model is touted as faster, smarter, and inevitably, more resource-intensive.
A broader trend towards frugality, in which buyers are seeking more affordable Chinese models, has already begun to impact the valuations of potential IPOs from OpenAI and Anthropic, indicating that sensitivity to price is no longer a minor issue.
Writer’s choice to build its flagship on GLM-5.2 is significant. Instead of creating a leading-edge model from scratch, Writer opted for a robust open-source foundation and refined it through post-training, a strategy that keeps costs manageable and aligns well with the economical narrative it is promoting.
For European clients, this positioning should resonate. Doubts about the motivations of American tech giants and a preference for solutions that are not tied to a single provider have long been common in this region, and Writer is appealing directly to that sentiment.
However, it is essential to remain cautious. Efficiency claims made by a vendor about its own product should always be regarded with skepticism, and “up to” 50% reduction does not guarantee a 50% cut. Nonetheless, the overall trend is difficult to contest.
The more profound transition is what Writer terms harness engineering: the approach of minimizing token usage per task rather than simply striving to achieve a higher ranking on a leaderboard. If this mindset takes root, as incentives increasingly suggest it will, the age of the perpetually resource-hungry model may finally confront its financial limits.
Otros artículos
The writer is placing their bets on more affordable AI agents through the Palmyra X6 and a streamlined harness.
Writer has introduced the Palmyra X6 along with an enhanced harness that it claims can reduce AI agent expenses by as much as 50%, anticipating that enterprises are reacting to rising costs.
