Thomson Reuters built its own AI model that now ranks among the best

Cynddl1 pts0 comments

Thomson Reuters Built Its Own AI Model That Now Ranks Among the World's Best - Thomson Reuters Institute

Skip to content

Jul 31, 2026 |<br>AI and product innovation

Thomson Reuters Built Its Own AI Model That Now Ranks Among the World’s Best

Joel Hron Chief Technology Officer, Thomson Reuters

Jonathan Schwarz Head of AI Research, Thomson Reuters

The most capable AI models no longer come only from frontier AI labs. One now comes from Thomson Reuters.

Today, we are sharing early benchmarking results for Thomson, a first of its kind AI model. Across a range of benchmarks assessing legal and general capabilities, Thomson performed competitively with the strongest frontier models on the market, including Claude Opus 4.8, and ahead of GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro.

Why? Because it knows the work.

Launching later this summer, Thomson is the newest layer of the Thomson Reuters AI strategy, and a demonstration of what becomes possible when authoritative content, expert judgment, professional tools, and model development come together.

In 2024, Thomson Reuters acquired Safe Sign Technologies, an AI research company. At the time, the market was betting that access to increasingly powerful general-purpose models would be enough.

We made a different bet. We believed the future of professional AI would require more than general-purpose intelligence. It would require models built specifically for the domains, standards, and consequences of professional work. We believed that the distinct advantages of Thomson Reuters decades of world-class content and expertise could be best expressed in a model that we ourselves crafted.

Thomson is the result of that bet. And it is why Thomson Reuters will continue to set the standard for Fiduciary-Grade AI™.

Meet Thomson

Thomson starts from a strong open-source foundation, so it performs general-purpose work just as effectively as the frontier models. It then goes further: trained using state of the art mid-training and post-training techniques on decades of authoritative content from Westlaw, Practical Law, Checkpoint, and Reuters, content professionals have staked their reputations on for generations.

That training was shaped by hundreds of subject matter experts who evaluated outputs, identified failure modes, and validated that the model reasons the way legal professionals actually work. The same professional standard governs how Thomson is deployed. Customer data is never used to train the model.

The result is a model that thinks and reasons like a lawyer while outperforming models multiple times larger on the work that matters.

Thomson Matches the Best. And Beats the Rest.

We evaluated Thomson against the leading general-purpose models on the market for general professional work and categories spanning:

Legal<br>Coding

Tax<br>Math

Accounting<br>Multilingualism

Journalism<br>Agentic tasks

Safety<br>Long context

Reasoning<br>Following instruction

Thomson is competitive with the world’s leading frontier models despite being a fraction of their size and cost to train and operate. Thomson Reuters has achieved that performance  by combining exceptional AI talent with authoritative proprietary content and deep domain expertise. And with less than 10% of Thomson Reuters content used in its training so far, there remains significant opportunity to expand its capabilities.

Instruction Following is a composite average of the IFEval and FollowBench benchmarks.  Reasoning is a composite average of the GPQA Diamond, HLE, and MMLU-Pro benchmarks.   Coding is a composite average of the SWE-Bench Pro and Terminal-Bench 2.1.  Long Context is a composite average of the Infinity Bench as well as some internal benchmarks developed by Thomson Reuters.

These evaluations show Thomson’s competitiveness with industry recognized benchmarks. Our internal evaluation and training cover a wide range of scenarios, including carefully designed agentic use cases optimized for real-world professional work, tens of thousands of real-world queries written by experts, end-to-end deep research training with human-calibrated judges, and data-centric mid-training on a large amount of our content. To increase safety and robustness, we conduct training and evaluations consistent with Thomson Reuters values, and stress-test the models through human and automated red-teaming. Thomson is still early in its development. To date, less than 10% of Thomson Reuters content has been used in its training, leaving significant opportunity to expand its domain knowledge and capabilities through additional training, rigorous evaluation, and expert validation.

Thomson also showcases its strength when a native integration with Thomson Reuters content is added. When compared with leading frontier models given unrestricted access to the web, Thomson’s access to proprietary data sources such as Westlaw, Practical Law, and Reuters news ensures both superior completeness and factuality (i.e. the ability to back up...

thomson reuters content models training model

Related Articles