Urgent.News

What's breaking now, across thousands of outlets.

AI

Thomson Reuters launches proprietary AI model for legal work

Global content powerhouse Thomson Reuters Corp. today launched Thomson, its first proprietary large language model, combining the company’s trove of legal knowledge with LLMs from outside providers to provide legal advice. The company said Thomson will first be deployed in Tabular Analysis, a high-volume document review capability in its CoCounsel Legal AI assistant. CoCounsel will […] The post…

Thomson Reuters launches proprietary AI model for legal work

Thomson Reuters has unveiled its first proprietary large language model, named Thomson, designed specifically for legal work. The model amalgamates the company's vast legal knowledge with LLMs from external providers to deliver legal advice. Initially, Thomson will power Tabular Analysis, a high-volume document review feature in Thomson Reuters' CoCounsel Legal AI assistant.

The assistant remains multi-model, utilizing Thomson for tasks where its domain-specific expertise is advantageous, while using third-party frontier models for other tasks. Thomson Reuters invested approximately $40 million over two years on personnel and computing resources for the project, but the final training cost was reduced to around $450,000.

Rather than constructing a foundation model from the ground up, Thomson Reuters began with an open-weight model, added its proprietary content, training methods, and professional expertise, which streamlined both training and inference costs compared to general-purpose frontier models.

Thomson's development process involved aligning the base model with Thomson Reuters' values, pretraining on the company's content, post-training adjustments guided by professionals, and reinforcement learning to teach the model to collaborate with tools like Westlaw and Practical Law. Over 40,000 databases and more than 150 years of legal publishing and editorial curation comprise Thomson Reuters' flagship Westlaw platform.

The model's development involved hundreds of subject-matter experts who defined training objectives, created legal question examples, and evaluated responses in blind comparisons.

Specialization can deteriorate a model's broader abilities if not managed properly, said Jonathan Schwartz, head of foundational research at Thomson Reuters. To counter this, the team focused on continual learning, adding domain skills without losing existing capabilities. Thomson Reuters internal tests revealed Thomson's broad competitiveness with leading models when all models had only web access.

When connected to Thomson Reuters' content, the model's performance either equaled or slightly surpassed leading models, assessing both answer completeness and citation support. The company plans to share Thomson with legal experts and academic institutions for testing and eventually release a smaller open-weight version on Hugging Face under a noncommercial academic license.

The next step involves adding more material and transforming the most useful content and product activity into more potent training signals. Thomson Reuters also owns customer data, which is not used for training the model, providing the company with greater control over deployment, governance, and future development.

Written by urgent.news from SiliconANGLE's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at siliconangle.com →

More in AI

More from Monday 24 August →