NewsAI ModelsEnterprise AIAlibaba Qwen

Thomson Reuters Invests $40 Million in Proprietary AI Model – Breaking Free from OpenAI Dependency

The legal information giant launches 'Thomson,' a custom language model built on Alibaba's Qwen. The strategy: independence over reliance on external AI vendors.

$40 million

Thomson Reuters Invests $40 Million in Proprietary AI Model – Breaking Free from OpenAI Dependency

Thomson Reuters has developed its own language model to break free from expensive dependence on OpenAI and Anthropic. The model "Thomson" is built on Alibaba's Qwen and was specifically trained for legal applications – with an investment of $40 million over two years for personnel and compute resources.

Quick Facts

  • Thomson Reuters develops "Thomson" language model based on Alibaba's Qwen (reportedly Qwen3.5-397B most recently)
  • Investment: $40 million over two years (of which only $450,000 for the final training run)
  • Benchmarks show top performance only with access to proprietary corporate content like Westlaw and Practical Law
  • First use case: document review; a smaller version will be released non-commercially

Model Development and Training

Thomson Reuters uses Alibaba's open Qwen model as its foundation. Working with the Imperial College, the Chinese model was first trained on safety, ethics, and political neutrality – this intermediate stage is called "Snowdon," named after a mountain in Wales. This was followed by pre-training on proprietary content, post-training with domain experts, and agentic reinforcement learning within the company's own tool environments.

Notably, less than ten percent of available content has been used in training so far. CTO Joel Hron notes that the company has changed its starting point roughly half a dozen times. The real achievement is not the individual model but the reusable "model factory," says Chief Research Officer Jonathan Schwartz.

Benchmarks with Caveats

Thomson Reuters markets "Thomson" as one of the world's best models – but its own numbers tell a more nuanced story:

Benchmark Thomson Competition Note
Stanford LegalBench 0.823 Gemini 3.1 Pro, GPT-5.5 (higher) Thomson trails competitors
Harvey Legal Agent just behind Opus 4.8
Deep-Research (web only) 0.53 GPT 5.4: 0.65 Thomson significantly weaker
Deep-Research (with corporate content) 0.83 GPT 5.4: 0.82 Thomson marginally ahead

Methodological comparability is questionable: Thomson uses test-time scaling, while GPT-5.5 runs without reasoning mode. More importantly, the advantage stems primarily from training on proprietary tools and content – access external providers don't have. Evaluation lead Andrew Bean concedes that Thomson with web-only access is "certainly not leading."

"What matters is not intelligence itself, but knowing which intelligence you need to own."

This statement from CTO Hron encapsulates the strategy: it's not about the best universal model, but specialization for your own business processes.

What This Means for German Companies

Thomson Reuters' approach highlights a dilemma many large enterprises face: dependence on US-based providers like OpenAI and Anthropic is expensive and strategically risky. The $40 million investment is substantial but manageable for a company of this scale – and it pays off when the model can access proprietary data. German enterprises with similar data assets (financial services, insurance, media) might consider similar moves. However, Thomson Reuters also demonstrates that proprietary development alone isn't enough – you need the strategic data advantage. Building a model without specialized training data and internal tools won't be competitive.

Sources

Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.

Share
← All articles

All analyses are based on i6eal's own measurements or on clearly labelled sources. Figures are snapshots and may change; corrections are disclosed transparently.