Alibaba's new compact model operates on a laptop and achieves performance similar to that of a cloud-based model.

Alibaba's new compact model operates on a laptop and achieves performance similar to that of a cloud-based model.

      Alibaba has introduced a compact AI model, Qwen3.8-27B, which operates on personal computers and performs comparably to cloud-based versions. According to The Information, the model has had over one million downloads within just a few days, marking it as one of the company's fastest-growing models.

      This model is open source and available for free download. It comprises 27 billion parameters and can comprehend text, image, and video prompts, as reported by The Information. Alibaba made the weights available on Hugging Face on Friday under an Apache 2.0 license, as noted by VentureBeat. This license allows companies to examine, modify, and host the model independently.

      The growing interest reflects a broader trend as Alibaba pursues the market for smaller models that operate on user hardware, steering clear of paid cloud services. The 27B model is part of the same family as Qwen3.8-Max, the company's leading large model, and it is comparable to models for which Alibaba recently began charging high-usage clients.

      Its performance is remarkable given its size. Qwen3.8-27B has achieved results comparable to much larger competitors while running on standard hardware, according to the South China Morning Post. Xinmei Shen reported that it performed similarly to OpenAI’s GPT-5.6 Luna, which was described as the most cost-effective model in OpenAI's latest flagship lineup.

      The model also closely matched two substantial Chinese open-weight models, nearly equaling DeepSeek’s V4-Pro, which was released the previous week with 1.7 trillion parameters, and getting close to Zhipu’s GLM-5.2 from June, which has 753 billion parameters, as stated by SCMP. These figures were provided by the benchmarking firm Artificial Analysis.

      Independent evaluations were published on Monday, with Artificial Analysis giving Qwen3.8-27B a score of 52 on its Intelligence Index, which aggregates results from nine different tests covering coding, science, and reasoning. This score matches that assigned to GPT-5.6 Luna at its highest reasoning setting. The index places the model at the top of its class among 135 tested models.

      Alibaba initially set the stage with its launch metrics, reporting scores of 61.7 on SWE-bench Pro and 90.3 on a coding test called LiveCodeBench, according to VentureBeat. In Alibaba's presented data, the 27B model even surpassed results from Claude Opus 4.6 on those assessments. However, VentureBeat pointed out that some of these evaluations were conducted internally, and the testing setups did not match exactly, making the results less reliable for determining a definitive winner.

      Developers took notice of the comparisons. The open-source coding tool Cline remarked on X that this was the first instance of a local model achieving frontier model capabilities. In a distinct agentic test by the firm, the model scored 51, surpassing Claude Opus 4.8 at maximum reasoning, which was released by Anthropic less than three months ago.

      The key focus is on the hardware rather than just leaderboard standings. The model requires about 56GB of GPU memory to run at full precision, as noted by VentureBeat. A compressed 4-bit version reduces the file size to roughly 17GB, making it feasible for high-end gaming desktops or well-equipped laptops.

      One developer, Simon Willison, tested this by running the roughly 17GB version on both an Apple laptop and an Nvidia desktop. He noted it wrote code, interpreted images, and executed a coding-agent loop, stating that it was remarkable for a 17GB file to perform these functions on personal machines.

      Usage surged quickly, with Qwen3.8-27B surpassing 3 million downloads on Hugging Face within its first three days, Cybernews reported, a figure greater than the one million mentioned by The Information for a similar timeframe.

      However, its high-quality performance comes with a trade-off regarding speed. Qwen3.8-27B generates significantly more reasoning text than its competitors, according to VentureBeat. Artificial Analysis reported that the model produced 160 million output tokens during testing, while comparable open-weight models averaged only 43 million.

      Willison encountered similar delays; a simple image request took 21 minutes and over 22,000 reasoning tokens, as the model defaults to its highest reasoning level. He suggests lowering this setting for typical local use. New inference software might help reduce the speed disparity, as Willison observed about a 72 percent performance increase on his Nvidia machine when activating a feature known as Multi-Token Prediction.

      Investor Tomasz Tunguz found a comparable trade-off during his own testing. With reasoning enabled, Qwen showed slightly better quality but exhibited substantially slower performance compared to a cloud model. He cautioned that his limited nine-task sample was insufficient for a conclusive assessment, arguing that the appropriate measure should focus on the time taken to reach a complete answer, rather

Other articles

Prevalent AI secures $22 million to address the data issues contributing to unsuccessful AI initiatives. Prevalent AI secures $22 million to address the data issues contributing to unsuccessful AI initiatives. London's Prevalent AI has secured $22 million from Integrity Growth Partners, marking its first major funding in nine years, to facilitate its expansion into the US market. The Top Ergonomic Office Chairs of 2026 The Top Ergonomic Office Chairs of 2026 Extended hours at your workstation require more than just a standard office chair. If you're enhancing your home office or substituting an old chair, these ergonomic options are crafted to ensure your comfort and support during the entire workday. Googlebook has officially announced a launch date, and it’s approaching quicker than anticipated. Googlebook has officially announced a launch date, and it’s approaching quicker than anticipated. Google is set to present Googlebook in New York on September 15, highlighting its software powered by Gemini, a native Android codebase, and laptops from its hardware partners. A recent report indicates that Meta compensates influencers to advocate against social media restrictions on teenage accounts globally. A recent report indicates that Meta compensates influencers to advocate against social media restrictions on teenage accounts globally. As nations seek to prohibit teenagers from using social media, Meta is compensating actors, psychologists, and parenting influencers in over a dozen countries to advocate against a universal ban on teen accounts. The $220 PlayStation speakers from Sony come with one feature that I really wish every desktop configuration included. The $220 PlayStation speakers from Sony come with one feature that I really wish every desktop configuration included. Sony has announced the price and release date for its first PlayStation wireless speakers, which feature planar magnetic drivers, wireless gaming audio, and the capability for voice chat without a headset, priced at $219.99. The Top Gaming Chairs to Look Out for in 2026 The Top Gaming Chairs to Look Out for in 2026 A gaming chair may appear stylish but can still have you adjusting your position after just an hour. The true distinction becomes evident when brief matches extend into lengthy gaming sessions. Here’s what differentiates appealing design from a chair designed for extended comfort.

Alibaba's new compact model operates on a laptop and achieves performance similar to that of a cloud-based model.

Alibaba's Qwen3.8-27B, a compact open model that operates on a laptop, achieved parity with leading cloud models on benchmarks and surpassed one million downloads within a few days.