Thrift-maxxing jeopardizes the IPOs of OpenAI and Anthropic.

Thrift-maxxing jeopardizes the IPOs of OpenAI and Anthropic.

      Approximately 750,000 words of AI output costs $50 when purchased from Anthropic’s Fable. In contrast, the same amount from DeepSeek’s V4-Pro is priced around 87 cents, according to Fortune. Z.AI’s GLM-5.2 charges $4.40, while Moonshot’s Kimi K3, considered expensive by Chinese standards, costs $15.

      Corporate America has taken notice. Companies spent a year competing to maximize token usage, but now they seek out the most economical models that can get the job done, as reported by the Wall Street Journal. Previously, the mark of achievement was token-maxxing; now it has shifted to thrift-maxxing.

      The initial response to rising costs was to ration, as companies limited employee spending. The next step is substitution, presenting a much larger challenge for OpenAI and Anthropic.

      “It’s akin to driving a Lamborghini just to fetch milk, while that car is meant for racing,” remarked Mike Saeks, a field chief technology officer at Cursor who advises companies on their AI investments.

      He has documented evidence. When estimating the cost of developing a web browser from scratch, running the entire operation on OpenAI’s GPT-5.5 amounted to just over $10,000. In contrast, utilizing Cursor’s own Composer model alongside Anthropic’s Opus 4.8 cost only $1,339.

      The focus is shifting from a single powerhouse to a collaborative approach. This situation does not indicate that companies are abandoning American labs; rather, they are assigning them tasks that validate the cost.

      Telnyx, which creates infrastructure for AI agents, operated 1,000 of them using a top model from Anthropic with a $200-per-employee subscription, but Anthropic withdrew third-party operating system support, deeming it a violation of terms. Paying based on usage would have resulted in costs “approaching $100,000 per day,” stated chief executive David Casem.

      Consequently, Telnyx restructured its system. Z.AI models now power its 1,400 agents, Anthropic’s Fable orchestrates and plans the tasks, open-weight models handle the execution, and OpenAI’s Sol reviews the output.

      The legal AI startup Harvey trained GLM-5.2 independently, creating a mechanism to call Fable 5 for particularly challenging tasks. “We collaborate with all of them,” noted president Gabe Pereyra. Zoom’s chief technology officer Xuedong Huang referenced an old Chinese story where three ordinary individuals combine their intellect to rival a single genius, suggesting this is the essential strategy.

      “It genuinely feels like a cutthroat environment,” stated Marty Kausas, chief executive of the customer support platform Pylon. Pylon has received months of unlimited free usage from providers. Kausas estimated that it has acquired roughly $1.6 million in free tokens from one vendor this year, with another contributing $65,000 and a third, $10,000.

      Both labs claim to be evolving. An OpenAI representative mentioned that GPT-5.6 Sol was specifically trained for enhanced token efficiency. Anthropic recently introduced a powerful, lower-cost model. An executive indicated that clients can select between intelligence and price within their ecosystem. Both labs endorse open-weight models.

      The reason for the significant price disparity is not altruistic. Electricity prices are lower in China, and new data centers face less local opposition there. Chinese companies are also willing to accept slimmer profit margins to secure market share and establish themselves as the default option.

      Export restrictions may have played a role as well. Deprived of top-tier Nvidia chips, Chinese labs had to optimize performance from inferior hardware. “For what a Chinese AI firm would invest in an Nvidia chip, they could acquire 10 local chips from Huawei or other local manufacturers,” stated George Chen of the Asia Group.

      Well-known companies are joining this trend too. Coinbase's chief executive Brian Armstrong mentioned in June that the exchange had reduced its AI expenditure by half by directing staff towards Kimi and Z.AI’s GLM models. DoorDash delegates what its chief technology officer Andy Fang calls “lower-level work” to Kimi for “better quality at a lower cost.” Airbnb has implemented Alibaba’s Qwen for customer service, while Cursor has developed its own Composer 2 coding model based on Kimi's foundations.

      The change is evident in usage patterns. In one week in July, Chinese models accounted for 57% of the tokens used by U.S. firms on OpenRouter. During one period mid-month, all five of the top models in the marketplace were Chinese.

      This trend extends beyond developers. IDC surveyed 260 decision-makers in U.S. companies with over 1,000 employees; 47% indicated they utilized a China-based model for at least one application, as reported by The Daily Upside. One in five cited extensive usage.

      This situation carries weight beyond a typical price war

Other articles

Apple's smartwatches are expected to retain their current design for at least the next year or two. Apple's smartwatches are expected to retain their current design for at least the next year or two. According to Bloomberg, Apple's upcoming Watch lineup will retain the same appearance as previous models, as the company emphasizes internal updates rather than design changes in its 2026 refresh. Lenovo's upcoming ThinkCenter desktop is set to be released, but the price may cause you to do a double-take. Lenovo's upcoming ThinkCenter desktop is set to be released, but the price may cause you to do a double-take. Lenovo's ThinkCentre X desktop is set to launch in the US offering up to 256GB of RAM, dual RTX 5060 Ti graphics cards, with a top-tier configuration priced at almost $14,000. Lenovo’s $99 280Hz gaming monitor is designed for those on a tight budget. Lenovo’s $99 280Hz gaming monitor is designed for those on a tight budget. Lenovo has introduced two budget-friendly 280Hz gaming monitors aimed at competitive gamers, spearheaded by a 23.8-inch version priced locally at approximately $99. Xbox is addressing one of the largest frustrations with download speeds in the PC app. Xbox is addressing one of the largest frustrations with download speeds in the PC app. Microsoft is trialing a Smart Download Client for Xbox that automatically connects to the fastest download server, which promises to speed up game installations on both PC and consoles. Multiverse Computing aims for a $570 million Series C funding round at a valuation of $1.7 billion to reduce costs associated with AI. Multiverse Computing aims for a $570 million Series C funding round at a valuation of $1.7 billion to reduce costs associated with AI. Spain's Multiverse Computing is seeking to raise as much as $570 million at an estimated valuation of $1.7 billion, focusing on compressing AI models instead of scaling them. Nadella establishes a stipulation regarding the AI surge on CNN. Nadella establishes a stipulation regarding the AI surge on CNN. Nadella may not refer to it as a bubble, but he states that AI needs to enhance GDP; otherwise, the outcome will be negative. In the meantime, Copilot is receiving chips ahead of Azure clients.

Thrift-maxxing jeopardizes the IPOs of OpenAI and Anthropic.

A million AI words are priced at $50 from Anthropic and 87 cents from DeepSeek. American companies are incorporating Chinese models, and the IPO valuations of the labs are under scrutiny.