Artificial intelligence development and the use of copyrighted data: Lessons from a dispute

Insights
Artificial intelligence development and the use of copyrighted data: Lessons from a dispute
Posted on: 22/08/2025

    In the era of rapidly developing artificial intelligence (AI) technology, the use of data to train AI models has become a core factor for creating competitive products. However, unauthorized use of copyrighted data can lead to serious legal disputes, causing financial and reputational losses. Thomson Reuters Enterprise Centre GmbH v. West Publishing Corp. v. Ross Intelligence Inc. (2025)[1] in the District Court of Delaware, USA is a good example, clearly illustrating the risks of copyright infringement during AI development.

     

     

    Thomson Reuters sues Ross Intelligence

    Thomson Reuters, the owner of legal research platform Westlaw, is one of the leading providers of legal data, including case law, laws, regulations, legal journals, and editorial notes organized under the Key Number System. These headnotes, although based on copyright-free case law, are considered creative works because of the editing and compilation process, and are therefore protected by copyright law. Ross Intelligence Inc., (Ross) a startup, developed an AI-based legal search engine to compete with Westlaw. To train the AI, Ross needed a large amount of legal data, but was denied permission by Thomson Reuters to use Westlaw's content.

    Ross then partnered with LegalEase, a third party, to create "Bulk Memos" – collections of legal questions and answers based on Westlaw's headnotes. Although LegalEase was instructed not to directly copy the headnotes, the questions in the Bulk Memos bear a significant resemblance to the headnotes, leading to allegations of copyright infringement from Thomson Reuters. In 2023, Judge Stephanos Bibas initially rejected most of Thomson Reuters' requests for summary judgments, but after further research, he revised his opinion in 2025. The court concluded that Ross had infringed copyright on 2,243 headnotes, rejecting Ross's arguments, including fair use, unintentional infringement, copyright abuse, and other legal doctrines. The lawsuit emphasizes that the use of copyrighted data to train AI, even at an intermediate step, can lead to serious legal consequences.

    Legal perspective from the case

    Allegations of direct copyright infringement

    Thomson Reuters accused Ross of direct copyright infringement of Westlaw's headnotes and Key Number System. The Court determined that:

    • Validity of copyright. Headnotes and Key Number System meet the threshold of originality according to Feist Publications, Inc. v. Rural Tel. Serv. Co. (1991).[2] Headnotes are the product of creative editing, similar to a sculpture carved out of stone (case law without copyright). The Key Number System, although based on popular legal topics, is still an independent organizational system of Thomson Reuters.
    • Replicate reality and significant similarity. Ross's expert report acknowledged that the 2,243 questions in the Bulk Memos were very similar to Westlaw's headnotes, but different from the original case law. The judge compared each pair of headlines, questions, and case law, concluding that this similarity was clear evidence of violation.

    Arguments for fair use

    Ross cited fair use under Chapter 17 of the U.S. Copyright Act[3], but was rejected by the court based on an analysis of four factors:

    1. Purpose and nature of use.

    Ross's use of headnotes is commercial and not transformative, as Ross's AI tool competes directly with Westlaw. Intermediate copying is not considered fair use, as it does not involve computer code as in the cases of Google v. Oracle (2021)[4] or Sony v. Connectix (2000).

    1. The nature of the copyrighted work.

    Headnotes and the Key Number System are creative, but not as high as fiction or art, so this element is in favor of Ross, but is less important.

    1. The number and importance of the part used.

    Ross's final product (which is case law) does not show headnotes, so this factor is in favor of Ross.

    1. Impact on the market.

    Ross downplayed the current market, which is Thomson Reuters' legal research market and AI training data potential, in favor of Thomson Reuters. This was the most important factor, leading to the rejection of Ross's fair use of defense.

    Other arguments

    Ross offered defenses such as unintentional infringement, copyright abuse, merger doctrine, and scenes à faire, all of which were dismissed. In particular, the court emphasized that headnotes have copyright notices, eliminating the possibility of reducing damages due to unintentional infringement, and that Thomson Reuters does not abuse copyright to stifle competition.

     

     

    Lessons for Vietnamese tech companies

    This lawsuit provides many important lessons for Vietnam, where the AI technology industry is developing rapidly but intellectual property law still has many challenges in applying to new technologies. Here are the specific lessons:

    First, respect intellectual property rights and apply for permission to use data

    The lawsuit shows that unauthorized use of copyrighted data, even through a third party, can lead to copyright infringement. In Vietnam, the Intellectual Property Law 2005, which was amended in 2009, 2022 protects creative works, including data aggregation or compilation. Technology companies, such as FPT, VinAI, or AI startups, need to: (i) Check the validity of copyrights. Before using data from sources such as books, newspapers, or legal databases, the Law Library, Vietnam Law or a law firm's website needs to verify whether the data is public data such as the original law issued by the State or the data was formed and developed and protected by an individual.  organization or not. (ii) Licensing negotiations. Where it is determined that the data is owned by a third party, it is necessary to sign an explicit license agreement or at least a written permission to use the data from the data owner. (iii) Employee training. Companies need to organize courses on intellectual property law and copyright to ensure that employees and implementing parties understand their legal responsibilities when using data created and owned by others.

    Second, distinguish between public data and copyrighted data

    The court determined that Westlaw's headnotes, even though based on public precedent, were protected for editorial creativity. In Vietnam, companies should:

    Understand originality. Materials such as legal commentaries, lectures, or organizational databases (such as library book catalogs) can be protected if there is an element of creativity. For example, legal analyses from the Supreme People's Court or lectures from Hanoi Law University may be copyrighted.

    Use public resources carefully. When using data from the Government Portal or the Official Gazette, it is necessary to ensure that copyrighted annotations or analyses are not copied.

    Create independent data. Companies should invest in aggregating data from publicly available sources, such as legal documents or case law, to avoid relying on copyrighted data.

    Third, carefully evaluate the appropriate legal basis

    Ross's fair use argument in the case fails because the use of data is commercial and competes directly with Westlaw. In Vietnam, Article 25 of the Intellectual Property Law allows fair use in cases such as non-profit research or fair citations, but businesses should also understand the fair use limits. The use of copyrighted data to create commercial products, especially if competing with owners, is hardly considered reasonable. For example, a Vietnamese AI company that uses data from textbooks to create a learning app that competes with a publisher could be sued. In addition to self-assessing the reasonableness of using data, technology businesses should consult legal expertise from law firms, especially companies that may be strong in intellectual property to assess whether the use of data falls within the scope of fair use. At the same time, businesses should also focus on the variability and differentiation of the data they create. AI products need to have a purpose or function that is markedly different from the original data to be considered fair use.

    Fourth, build a transparent data management process

    Ross did not thoroughly check the provenance of the data from LegalEase, which led to the breach. This experience shows that Vietnamese technology companies when developing AI-based data need to:

    Check the origin of the data. Establish a process to check the legality of data from the vendor, including a request for proof of ownership or license.

    Clear contracts. When working with third parties, it is necessary to have a contract that clearly states the liability for the data and the rights involved. For example, if a Vietnamese AI company hires a third party to collect legal data, it is necessary to require a commitment that the data does not infringe copyright, does not violate regulations on personal data collection,...

    Record keeping. It is essential for businesses to record details of the source of the data and how it will be used as evidence in the event of a potential dispute over the use of this data that may occur between stakeholders in the future.

    Fifth, businesses should invest in proprietary data and research and development

    The court noted that Ross was able to generate similar headnotes data on his own without infringing copyright. Vietnamese companies should be careful in creating proprietary data. The investment in the collection and analysis of data from publicly available sources, such as legal documents from the Official Gazette or case law from the Supreme People's Court. In addition, increasing R&D, investing in research and development to build its own database, and reducing dependence on third-party data is also an effective way to avoid unnecessary disputes in this field.

    The Thomson Reuters v. Ross Intelligence case  is a wake-up call for tech companies about the legal risks of using copyrighted data in AI development. For Vietnam, where the AI industry is thriving, the lessons from this lawsuit underscore the importance of complying with intellectual property laws.  building proprietary data, and managing legal risks. By applying these lessons, Vietnamese technology companies can not only avoid costly disputes but also build AI products that are sustainable, transparent, and competitive in the global market.