Tag

Ai Training Data

All articles tagged with #ai training data

UMG and Sony Hit Suno with Second AI Training Lawsuit Over Unlicensed Recordings
business20 days ago

UMG and Sony Hit Suno with Second AI Training Lawsuit Over Unlicensed Recordings

Universal Music Group and Sony Music Entertainment filed a second lawsuit against Suno in Boston, alleging Suno copied about 60,000 copyrighted recordings to train its AI models, including the v6 suite launched Sept. 9, and seeking damages potentially in the billions plus an injunction and a jury trial. The complaint argues v6 is not a fresh start but the product of training on outputs of infringing models, while Suno contends it trained v6 from scratch; the case follows their earlier suit and underscores concerns about market harm from AI-generated music.

US DOJ Sides with OpenAI in High-Stakes Copyright Fight Over AI Training Data
technology1 month ago

US DOJ Sides with OpenAI in High-Stakes Copyright Fight Over AI Training Data

The U.S. Department of Justice urged a New York court to let OpenAI broadly access copyrighted material for training its AI, arguing that such use is fair, transformative, and essential to maintaining U.S. leadership in AI. The Intercept and other media plaintiffs contend this would amount to an uncompensated transfer of IP rights to tech companies. The case, which involves open DMCA claims against OpenAI and has been consolidated with other media lawsuits (including The New York Times, Tribune, and Ziff Davis), has seen mixed rulings and continues as arguments over fair use and the impact on press freedom unfold.

Publishers Claim Claude Was Trained on Pirated Lyrics, Demand Up to $150K per Song
technology1 month ago

Publishers Claim Claude Was Trained on Pirated Lyrics, Demand Up to $150K per Song

Sony Music Publishing and Warner Chappell sued Anthropic in a Northern California court, alleging Claude was trained by scraping thousands of copyrighted songs and lyrics (including hits like Eye of the Tiger and All I Want for Christmas Is You), leading to AI-generated lyrics that resemble originals. The publishers seek a jury trial and statutory damages of up to $150,000 per composition; Anthropic denies the allegations.

Lawsuit Accuses Grok of Training on CSAM, Expanding Musk AI Controversy
technology1 month ago

Lawsuit Accuses Grok of Training on CSAM, Expanding Musk AI Controversy

A new class-action lawsuit alleges Elon Musk’s Grok AI was trained on child sexual abuse material (CSAM), including depictions of the plaintiff, Jane Doe. The filing claims CSAM was part of training data used by xAI to develop Grok, with hash-listed material from the National Center for Missing & Exploited Children and notifications from the Canadian Centre for Child Protection, and argues that public posts on X feed into Grok’s training pipeline. The suit seeks class-action status, destruction of any Grok-generated CSAM, and measures to prevent further generation. It notes prior reporting on related CSAM datasets but does not confirm those datasets were used, and it frames the case as a landmark challenge to xAI’s training practices.

Publishers sue Anthropic over AI training data, seeking billions
ai1 month ago

Publishers sue Anthropic over AI training data, seeking billions

Sony Music Publishing and Warner Chappell have filed a federal lawsuit against Anthropic in California, seeking potentially billions in damages for allegedly using tens of thousands of copyrighted works to train Anthropic’s Claude AI, with damages up to $150,000 per work and $25,000 for each instance where copyright data was stripped. The complaint also names co-founders Dario Amodei and Benjamin Mann, accusing Mann of using BitTorrent to download millions of pirated books and alleging that employees scraped lyrics from licensed sites; songs cited include Ain’t No Mountain High Enough, Livin’ On a Prayer, September, Hallelujah, and Paper Rings.

Grok Accused of Being Trained on CSAM in New Class-Action Lawsuit
artificial-intelligence1 month ago

Grok Accused of Being Trained on CSAM in New Class-Action Lawsuit

A new proposed class-action accuses Elon Musk’s xAI and Grok of training its AI chatbot on child sexual abuse material (CSAM), alleging CSAM posted via Grok fed back into training to generate more abuse content; the plaintiff seeks damages and destruction of all Grok-generated CSAM, arguing there is no AI exception to federal child-protection laws. This is the first suit to claim Grok was trained on CSAM, following prior legal actions over safeguards and user abuse of the tool.

Lawsuit alleges xAI trained Grok on CSAM, including AI-generated images
technology1 month ago

Lawsuit alleges xAI trained Grok on CSAM, including AI-generated images

A plaintiff accuses Elon Musk’s xAI of training its Grok image/video generator on child sexual abuse material (CSAM), including AI-generated CSAM, using hashed survivor images; the suit seeks to halt Grok’s harmful outputs, destroy generated CSAM, and argues violations of federal CSAM laws and Masha’s Law, naming the survivor as Doe. The case is the first to target xAI for CSAM training, and xAI did not comment.

Handshake AI Offers Up to $30K for User-Owned Written Documents
technology1 month ago

Handshake AI Offers Up to $30K for User-Owned Written Documents

Handshake AI is hiring a Professional Document Contributor, offering up to $30,000 for owner-owned, high-quality written documents (Word or PDF) submitted at $6 per page for up to 50 documents, each up to 100 pages; acceptance criteria aren’t defined and slide decks are generally ineligible, raising questions about data ownership, confidentiality, and how submitted documents will train AI models.

Anthropic's $1.5B Settlement Marks Major Win in Copyright Case Over AI Training Data
technology2 months ago

Anthropic's $1.5B Settlement Marks Major Win in Copyright Case Over AI Training Data

A San Francisco federal judge approved a $1.5 billion settlement between Anthropic and authors who accused the company of using pirated books to train its AI. Described as the largest copyright class-action settlement, the deal pays roughly $3,000 per book for claims covering around 440,000 titles; Anthropic must delete the pirated files. The accord resolves liability for past data acquisition, not for future AI outputs or new claims.

Hack-cached data ties Suno’s AI training to YouTube and Deezer in new legal spotlight
technology2 months ago

Hack-cached data ties Suno’s AI training to YouTube and Deezer in new legal spotlight

A hacked code dump reportedly shows Suno scraped YouTube Music, Deezer, and Genius (among others) to train its AI models, a detail that could strengthen Universal Music Group and Sony Music Entertainment’s copyright-infringement case. Suno has said its training relied on publicly available data from the open internet and fair use, while the breach also exposed customer emails/phone numbers and Stripe payments, adding factual context to the ongoing litigation and the scale of potential damages.

Rivian bets on a separate Mind Robotics to pioneer humanoids
technology3 months ago

Rivian bets on a separate Mind Robotics to pioneer humanoids

Rivian CEO RJ Scaringe has launched Mind Robotics as a standalone company with over $1 billion in funding to develop humanoid robots, with Rivian as a major shareholder and launch customer. Mind aims to reveal its first product within a year, while Scaringe says the unit will stay separate from Rivian. He envisions robots working alongside human manufacturing teams, using Rivian data to train AI, and notes a long‑term automation shift—though Musk’s Tesla/AI strategy remains a contrasting approach.

technology4 months ago

US Court Orders Global Action Against Anna's Archive in $19.5M Judgment

A New York federal judge issued a $19.5 million default judgment and a permanent injunction targeting Anna's Archive, a shadow library tied to Sci-Hub and LibGen, naming more than 20 intermediaries to shut down its domains. Operators remain anonymous, complicating enforcement, and while U.S. courts can bind some entities (like Cloudflare), many registrars and country-level registries may resist; the case underscores the conflict between piracy, AI training data needs, and global enforcement.

Wikipedia signs paid-data deals with AI firms to fund its infrastructure
technology8 months ago

Wikipedia signs paid-data deals with AI firms to fund its infrastructure

The Wikimedia Foundation has begun paid data-access deals with AI firms including Amazon, Meta, Microsoft, Mistral AI, and Perplexity to monetize Wikipedia’s data and help cover rising infrastructure costs from automated scraping, signaling a shift from donation-based funding to enterprise partnerships; the foundation also envisions AI tools to assist editors and a conversational search experience that cites verified text.

Wikipedia opens paid licensing for AI training with major tech firms
technology8 months ago

Wikipedia opens paid licensing for AI training with major tech firms

The Wikimedia Foundation expanded Wikimedia Enterprise to offer paid licensing of Wikipedia content to major tech firms—Microsoft, Meta, Amazon, Perplexity, and Mistral AI—letting them use Wikipedia’s 65 million articles to train AI models, joining Google's existing deal. The move monetizes API access to offset rising infrastructure costs as AI scrapes increase, while editors push back on AI-generated content and bot traffic remains a concern.