Tech Giants Clash Over AI Training as Executives Call Data Scraping a Labor Theft

Newly unsealed court documents in The New York Times’ 2023 lawsuit against OpenAI and Microsoft show executives from both companies feared AI training on news content could threaten journalism, with remarks likening mass scraping and paywall circumvention to the 'largest theft of labor in history.' The materials highlight how training data was assembled by bypassing paywalls and erasing copyright notices, fueling a broader debate over whether such use qualifies as fair use. Microsoft, while distancing itself from the statements, argues transformative AI uses align with copyright law, while OpenAI executives described AI products as largely substitutive to journalism; the case remains pivotal in defining data sourcing for large language models.
- Microsoft Executive Called OpenAI's Web Scraping The 'Largest Theft Of Labor In Human History' engadget.com
- Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History The New York Times
- Tech Companies’ Staff Knew Their AI Tools Posed ‘Existential Threat’ to Publishers WSJ
- Did OpenAI Pull Off the Biggest IP Heist of All-Time? The Hollywood Reporter
- OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web The Verge
Want the full story? Read the original reporting
Read on engadget.com