Thomson Reuters developed its own model using a Chinese open-source foundation.
Thomson Reuters introduced its first proprietary large language model on Monday, naming it Thomson. The company reported spending $40 million on talent and computing resources to train the model, which is based on what it describes as “a strong open-source foundation.” However, the specific foundation is not mentioned in the announcement. The chief technology officer revealed it in an interview.
Thomson Reuters, which trades as TRI on the Toronto and Nasdaq exchanges, provides legal, tax, accounting, and compliance products and owns Reuters. The announcement indicates that the $40 million expenditure encompasses both talent and computing resources. This investment is compared to frontier labs, which the company claims have typically invested billions in computing and years in infrastructure to advance their projects. Thomson Reuters states that the end result is a model it fully controls, “without the heavy inference costs of typical frontier models.”
According to the release, Thomson has been trained on less than 10% of its own content so far, with Westlaw, Practical Law, Checkpoint, and Reuters identified as sources.
In an interview with Business Insider, CTO Joel Hron disclosed that Thomson is based on a model named Snowdon, which was developed by “realigning” an open-source Qwen model from Alibaba. Business Insider specified the base model as Qwen3.5 and referred to the new model as Thomson-1. The company's press release only broadly characterizes the starting point as “a strong open-source foundation” without detailing its identity.
Regarding the development of Snowdon, Hron stated that a collaborative team from Thomson Reuters and Imperial College in the UK adapted Qwen over several months, ensuring that the outcome was “ethically and politically de-biased and safe to use.” He remarked, “There’s nothing that necessarily ties us to Qwen,” adding that Alibaba recently indicated plans to charge its largest users.
Thomson's initial deployment will be within Tabular Analysis in CoCounsel Legal, the AI assistant for lawyers. This functionality focuses on “high-volume, structured document review” and will soon reach law firms and corporate legal departments in the next release. The company notes that CoCounsel Legal “remains multi-model by design” and aims to integrate Thomson across its legal and tax offerings.
On August 20, Thomson Reuters and document-management firm iManage announced an expanded partnership, enhancing the integration of CoCounsel Legal within the iManage platform, alongside HighQ, Noetica, and Legal Tracker. The companies will also implement support for Model Context Protocol, allowing approved Thomson Reuters tools to reason over iManage content while maintaining access controls, ethical boundaries, and privilege limits. Rawia Ashraf, co-head of CoCounsel Legal, highlighted the challenge of managing legal work across various platforms.
Hron mentioned that CoCounsel still primarily relies on Claude. Thomson Reuters had expanded its partnership with Anthropic in May for this specific product. He expressed the goal of making Thomson the driving model behind more of CoCounsel's functionalities over time, asserting that the new model does not supplant the company's collaborations with Anthropic and other labs.
Hron cited cost as a principal reason for developing its own model, allowing the company to build on its intellectual property without incurring expenses from external AI providers, likening it to the difference between renting and buying a house. He stated, “Renting a house, you still have a roof over your head, and somebody’s taking care of it, and it’s great. But you’re not building any equity that compounds into something valuable for you long term.” River AI raised $1.1 billion recently to enable companies to train and retain their own models.
Anthropic has alleged that Chinese labs have unlawfully distilled the outputs of its models to develop their own, naming Alibaba in June as part of what it described as the largest distillation campaign against Claude. The firm has called for the US government to impose restrictions on such practices. Senator Tom Cotton has expressed concerns regarding US companies utilizing Chinese open-source models, citing potential security risks, including backdoors. Neither Anthropic nor Alibaba has responded to Business Insider's inquiries.
Thomson Reuters reported involving legal and AI academics with the model prior to its launch, quoting two individuals in its announcement. Jonathan H. Choi from Washington University School of Law tested Thomson against ChatGPT and Claude using questions from his Corporate Tax class, finding that while all three models answered correctly, he preferred Thomson’s responses for their connections to treatises. Professor Samuel Dahan, who leads the Queen’s Conflict Analytics Lab and the Cornell Legal AI Lab, concluded that Thomson’s “citation quality generally competes with leading frontier models,” including in Canadian employment law queries.
Thomson Reuters is releasing a “small” version of Thomson on Hugging Face as an open-weight model for academic and non-commercial purposes, alongside a technical report detailing the foundation model's development.
In a related move, Harvey introduced Tenet on Sunday as its first proprietary legal model, which was
Other articles
Thomson Reuters developed its own model using a Chinese open-source foundation.
Thomson Reuters invested $40 million in Thomson, marking its first internal model. The announcement mentions an open-source foundation. The Chief Technology Officer referred to Alibaba's Qwen.
