Close Menu
    Facebook X (Twitter) Instagram
    Trending
    • ChatGPT Hit by Major Outage as OpenAI Reports Errors and Delays
    • CeX Website Confirms Plans for Retro-Only Games Store in Birmingham
    • Artificial Intelligence Could Trigger Global Economic Downturn, Bank of England Governor Warns
    • Rising carbon dioxide is accelerating grass growth across African savannas, study finds
    • Fortnite cheat codes: Full list of Lobby Hack codes and island cheats explained
    • RMT announces fresh East Midlands Railway strikes over Aurora train safety concerns
    • Boardmasters Festival 2026: Emerging South West Artists Ready to Take Centre Stage
    • Xbox Free Play Days: Five Games Available to Try for Free This Weekend
    Mediarun Search
    • Home
    • Top News
    • World
    • Economy
    • Science
    • Technology
    • Sport
    • Entertainment
    • Contact Us
    Mediarun Search
    Home»Economy»Risk: OpenAI destroyed databases containing more than 100,000 books used to train ChatGPT
    Economy

    Risk: OpenAI destroyed databases containing more than 100,000 books used to train ChatGPT

    Charlotte WhitmoreBy Charlotte WhitmoreMay 11, 2024No Comments2 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Risk: OpenAI destroyed databases containing more than 100,000 books used to train ChatGPT
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Photo: Getty Images

    The conflict between the authors' union and OpenAIowner of ChatGPT, has just begun a new chapter, with documents proving that the startup used thousands of books to train its algorithms.

    The consortium is suing the startup, claiming that OpenAI infringed the copyright of published works for AI training.

    New evidence suggests that the startup deleted two databases, known as books1 and books2, which contained more than 100,000 published works. According to Business Insider, OpenAI has been reluctant to acknowledge the existence of these files. More recent documents, dated 2020 and now released, reveal that the books1 and books2 databases account for 16% of the total training used to create GPT-3, totaling 50 billion words. OpenAI's lawyers claim that the textbook training was retired at the end of 2021 and the databases were deleted the following year, and that none of the current ChatGPT models were created using these files. Furthermore, those responsible for creating the files are no longer in the company. Using published books is crucial to training high-quality AI models, but the lack of financial compensation for copyright holders has led to legal disputes, including lawsuits brought by the Authors' Union. The startup seeks to keep the contents of databases and the identity of employees confidential.

    Charlotte Whitmore

    Charlotte Whitmore is a contributor at Mediarunsearch.co.uk, covering a broad range of topics including news, politics, business, technology, sport, entertainment, and lifestyle. She focuses on delivering clear, balanced reporting and practical information that helps readers stay informed about current events and emerging developments. Her work highlights stories that matter to everyday audiences, with an emphasis on accuracy, relevance, and accessible journalism that keeps readers connected to the issues shaping the UK and beyond.

    See also  Don't do this while your cell phone is charging or it will call for help
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

    Related Posts

    South East Water Ordered to Fund £30.5 Million Improvement Programme Following Major Supply Failures

    July 14, 2026

    UK Green Economy Surpasses £100bn as Net Zero Sector Drives Jobs and Investment

    June 3, 2026

    BYD to cooperate with Senate to deregulate electric vehicles

    October 28, 2025
    Leave A Reply Cancel Reply

    Navigate
    • Home
    • Top News
    • World
    • Economy
    • Science
    • Technology
    • Sport
    • Entertainment
    • Contact Us
    Pages
    • About Us
    • Contact Us
    • DMCA
    • Editorial Policy
    • Privacy Policy
    © 2026 Media Run Search. All Rights Reserved. Designed by Media Run Search.

    Type above and press Enter to search. Press Esc to cancel.