Thrift-maxxing jeopardizes the IPOs of OpenAI and Anthropic.
Approximately 750,000 words of AI output costs $50 when purchased from Anthropic’s Fable. In contrast, the same amount from DeepSeek’s V4-Pro is priced around 87 cents, according to Fortune. Z.AI’s GLM-5.2 charges $4.40, while Moonshot’s Kimi K3, considered expensive by Chinese standards, costs $15.
Corporate America has taken notice. Companies spent a year competing to maximize token usage, but now they seek out the most economical models that can get the job done, as reported by the Wall Street Journal. Previously, the mark of achievement was token-maxxing; now it has shifted to thrift-maxxing.
The initial response to rising costs was to ration, as companies limited employee spending. The next step is substitution, presenting a much larger challenge for OpenAI and Anthropic.
“It’s akin to driving a Lamborghini just to fetch milk, while that car is meant for racing,” remarked Mike Saeks, a field chief technology officer at Cursor who advises companies on their AI investments.
He has documented evidence. When estimating the cost of developing a web browser from scratch, running the entire operation on OpenAI’s GPT-5.5 amounted to just over $10,000. In contrast, utilizing Cursor’s own Composer model alongside Anthropic’s Opus 4.8 cost only $1,339.
The focus is shifting from a single powerhouse to a collaborative approach. This situation does not indicate that companies are abandoning American labs; rather, they are assigning them tasks that validate the cost.
Telnyx, which creates infrastructure for AI agents, operated 1,000 of them using a top model from Anthropic with a $200-per-employee subscription, but Anthropic withdrew third-party operating system support, deeming it a violation of terms. Paying based on usage would have resulted in costs “approaching $100,000 per day,” stated chief executive David Casem.
Consequently, Telnyx restructured its system. Z.AI models now power its 1,400 agents, Anthropic’s Fable orchestrates and plans the tasks, open-weight models handle the execution, and OpenAI’s Sol reviews the output.
The legal AI startup Harvey trained GLM-5.2 independently, creating a mechanism to call Fable 5 for particularly challenging tasks. “We collaborate with all of them,” noted president Gabe Pereyra. Zoom’s chief technology officer Xuedong Huang referenced an old Chinese story where three ordinary individuals combine their intellect to rival a single genius, suggesting this is the essential strategy.
“It genuinely feels like a cutthroat environment,” stated Marty Kausas, chief executive of the customer support platform Pylon. Pylon has received months of unlimited free usage from providers. Kausas estimated that it has acquired roughly $1.6 million in free tokens from one vendor this year, with another contributing $65,000 and a third, $10,000.
Both labs claim to be evolving. An OpenAI representative mentioned that GPT-5.6 Sol was specifically trained for enhanced token efficiency. Anthropic recently introduced a powerful, lower-cost model. An executive indicated that clients can select between intelligence and price within their ecosystem. Both labs endorse open-weight models.
The reason for the significant price disparity is not altruistic. Electricity prices are lower in China, and new data centers face less local opposition there. Chinese companies are also willing to accept slimmer profit margins to secure market share and establish themselves as the default option.
Export restrictions may have played a role as well. Deprived of top-tier Nvidia chips, Chinese labs had to optimize performance from inferior hardware. “For what a Chinese AI firm would invest in an Nvidia chip, they could acquire 10 local chips from Huawei or other local manufacturers,” stated George Chen of the Asia Group.
Well-known companies are joining this trend too. Coinbase's chief executive Brian Armstrong mentioned in June that the exchange had reduced its AI expenditure by half by directing staff towards Kimi and Z.AI’s GLM models. DoorDash delegates what its chief technology officer Andy Fang calls “lower-level work” to Kimi for “better quality at a lower cost.” Airbnb has implemented Alibaba’s Qwen for customer service, while Cursor has developed its own Composer 2 coding model based on Kimi's foundations.
The change is evident in usage patterns. In one week in July, Chinese models accounted for 57% of the tokens used by U.S. firms on OpenRouter. During one period mid-month, all five of the top models in the marketplace were Chinese.
This trend extends beyond developers. IDC surveyed 260 decision-makers in U.S. companies with over 1,000 employees; 47% indicated they utilized a China-based model for at least one application, as reported by The Daily Upside. One in five cited extensive usage.
This situation carries weight beyond a typical price war
Other articles
Thrift-maxxing jeopardizes the IPOs of OpenAI and Anthropic.
A million AI words are priced at $50 from Anthropic and 87 cents from DeepSeek. American companies are incorporating Chinese models, and the IPO valuations of the labs are under scrutiny.
