Thomson Reuters built its own frontier model for $40 million
4 min read
- Instruction following
- 0.914
- Harvey Legal Agent
- 0.857
- Stanford LegalBench
- 0.823
- Long context
- 0.753
- Reasoning
- 0.684
- Coding
- 0.399
- PrBench Legal Hard
- 0.352

Thomson Reuters launched Thomson, its first proprietary large language model, on 24 August 2026, and the claim attached to it is the one every company sitting on a large private corpus wants to hear: domain-specific training matching frontier labs at a fraction of their cost. SiliconANGLE puts the investment at roughly $40 million over two years, with the final training run costing about $450,000. The model did not start from scratch - Thomson Reuters says it began from "a strong open-source foundation" and added proprietary content and training methods on top.
What the money bought
The corpus is the asset. Thomson Reuters says Thomson was trained on content from Westlaw, Practical Law, Checkpoint and Reuters, shaped by what it describes as hundreds of subject matter experts. The Globe and Mail reports the work was done over two years with Safe Sign Technologies, a British startup Thomson Reuters acquired last year - a detail the company's own press release does not mention.
How much of the archive has been used is reported three ways: Thomson Reuters says less than 10%, SiliconANGLE says about 10%, and The Globe and Mail says 8%. Consistent within rounding, and all three carry the same implication the company is selling - most of the corpus is still unused, so there is headroom. SiliconANGLE adds that customer data was not used for training.
The benchmark numbers, and their footnotes
LawNext published the only detailed look at the scores, testing Thomson-1-Large against Gemini 3.1 Pro, Claude Opus 4.8 and GPT-5.5 across seven categories. Thomson takes first place in instruction following, long context and PrBench Legal Hard, second on the Harvey Legal Agent benchmark, third on reasoning - and last on coding.
Then the footnotes. LawNext found the runs were not matched: Thomson benefited from test-time scaling while Gemini and Opus ran in reasoning mode, and GPT-5.5 was tested only in non-reasoning mode, which LawNext notes "is typically less capable when handling complex, advanced or multi-step challenges." Thomson Reuters said the scaling advantage was minor, moving its average from 0.787 to 0.789, and that a reasoning-mode GPT-5.5 evaluation "took substantially longer" and was not ready to publish.
The second evaluation is weaker still. It compared Thomson connected to Westlaw and Practical Law against frontier models with web access, which LawNext summarises as "less a test of the models than of the retrieval sources behind them." Its bottom line is the sentence to keep: the results are entirely self-reported, with no independent third-party verification. Andrew Bean, a senior research scientist at the company, describes the method as assessing "both the completeness of answers and whether citations supported their claims."
Where it ships, and what is coming for developers
Thomson is not replacing anything wholesale. It ships first in Tabular Analysis, a high-volume document-review feature inside the CoCounsel Legal assistant, where SiliconANGLE says it becomes the default while administrators can still pick alternatives. CoCounsel stays multi-model by design, using Thomson where the company believes domain training wins and third-party frontier models elsewhere. The Globe and Mail reports the internal test is literally framed as running "a horse race" between Thomson and the leading frontier models, plus future licensing to large law firms for on-premises deployment.
Two developer-facing plans matter more than the launch itself. SiliconANGLE reports Thomson Reuters intends to release a smaller open-weight version on Hugging Face under a noncommercial academic licence, and is building an API portal for outside developers.
What this means for developers
The transferable lesson is the cost structure, not the leaderboard. A $450,000 final training run on top of an open-weight base is within reach of a mid-sized company, and the moat was the corpus plus expert labelling, not compute. If your employer sits on decades of proprietary structured data, the build-versus-buy calculation just moved, and the honest first step is auditing whether your data is actually clean and labelled enough to train on - that is where these projects die, not at the GPU bill.
Read the evaluation as a template for your own. LawNext's critique is a checklist worth stealing: were competing models run in their strongest mode, was retrieval held constant, and did anyone outside the vendor verify the numbers? A domain model that wins because it alone had Westlaw is telling you about retrieval, not about the model. Run the same test with your retrieval pipeline wired to both and the answer often changes.
The coding score is the useful caution against over-reading domain training. Thomson finishes last there, which is exactly what a legal-corpus model should do - it says nothing bad about the model and everything about scope. A model trained on your domain will be narrow in precisely the way your domain is narrow, so keep a general model in the stack for everything else, which is what Thomson Reuters itself has done.
Sources
- Thomson Reuters Leverages its World-Class Data Assets to Launch Its Own Frontier Model - Thomson Reuters
- Thomson Reuters launches proprietary AI model for legal work - SiliconANGLE
- Thomson Reuters Says Its Homegrown AI Model Now Rivals the Frontier Labs - I Take A Closer Look At the Benchmarks - LawNext
- Thomson Reuters to roll out in-house AI model as quarterly revenue rises - The Globe and Mail
Related articles

Z.ai and IBM ship open-weight reasoning models a day apart
Z.ai released GLM-5.3-Flash under MIT with 320B parameters. IBM released Granite 4.2 under Apache 2.0 at 3B to 30B. They suit very different hardware.

Researchers document a near-autonomous AI-agent attack on Taiwan
Israeli firm Dream says AI agents built on open-source frameworks Hermes and OpenClaw ran a four-day intrusion on Taiwan's government, compromising 85 accounts with little human input.

DeepSeek warns developers a significant API price rise is coming
DeepSeek warned developers on August 6 that API prices will rise 'significantly' soon, with no figures or dates - while V4-Flash is the cheapest well-known model to run.
The developer AI briefing
3–5 stories a day, what they mean for developers. Free, no spam.