Tag

Ai Training Data

All articles tagged with #ai training data

Handshake AI Offers Up to $30K for User-Owned Written Documents
technology7 days ago

Handshake AI Offers Up to $30K for User-Owned Written Documents

Handshake AI is hiring a Professional Document Contributor, offering up to $30,000 for owner-owned, high-quality written documents (Word or PDF) submitted at $6 per page for up to 50 documents, each up to 100 pages; acceptance criteria aren’t defined and slide decks are generally ineligible, raising questions about data ownership, confidentiality, and how submitted documents will train AI models.

Anthropic's $1.5B Settlement Marks Major Win in Copyright Case Over AI Training Data
technology1 month ago

Anthropic's $1.5B Settlement Marks Major Win in Copyright Case Over AI Training Data

A San Francisco federal judge approved a $1.5 billion settlement between Anthropic and authors who accused the company of using pirated books to train its AI. Described as the largest copyright class-action settlement, the deal pays roughly $3,000 per book for claims covering around 440,000 titles; Anthropic must delete the pirated files. The accord resolves liability for past data acquisition, not for future AI outputs or new claims.

Hack-cached data ties Suno’s AI training to YouTube and Deezer in new legal spotlight
technology1 month ago

Hack-cached data ties Suno’s AI training to YouTube and Deezer in new legal spotlight

A hacked code dump reportedly shows Suno scraped YouTube Music, Deezer, and Genius (among others) to train its AI models, a detail that could strengthen Universal Music Group and Sony Music Entertainment’s copyright-infringement case. Suno has said its training relied on publicly available data from the open internet and fair use, while the breach also exposed customer emails/phone numbers and Stripe payments, adding factual context to the ongoing litigation and the scale of potential damages.

Rivian bets on a separate Mind Robotics to pioneer humanoids
technology2 months ago

Rivian bets on a separate Mind Robotics to pioneer humanoids

Rivian CEO RJ Scaringe has launched Mind Robotics as a standalone company with over $1 billion in funding to develop humanoid robots, with Rivian as a major shareholder and launch customer. Mind aims to reveal its first product within a year, while Scaringe says the unit will stay separate from Rivian. He envisions robots working alongside human manufacturing teams, using Rivian data to train AI, and notes a long‑term automation shift—though Musk’s Tesla/AI strategy remains a contrasting approach.

technology3 months ago

US Court Orders Global Action Against Anna's Archive in $19.5M Judgment

A New York federal judge issued a $19.5 million default judgment and a permanent injunction targeting Anna's Archive, a shadow library tied to Sci-Hub and LibGen, naming more than 20 intermediaries to shut down its domains. Operators remain anonymous, complicating enforcement, and while U.S. courts can bind some entities (like Cloudflare), many registrars and country-level registries may resist; the case underscores the conflict between piracy, AI training data needs, and global enforcement.

Wikipedia signs paid-data deals with AI firms to fund its infrastructure
technology7 months ago

Wikipedia signs paid-data deals with AI firms to fund its infrastructure

The Wikimedia Foundation has begun paid data-access deals with AI firms including Amazon, Meta, Microsoft, Mistral AI, and Perplexity to monetize Wikipedia’s data and help cover rising infrastructure costs from automated scraping, signaling a shift from donation-based funding to enterprise partnerships; the foundation also envisions AI tools to assist editors and a conversational search experience that cites verified text.

Wikipedia opens paid licensing for AI training with major tech firms
technology7 months ago

Wikipedia opens paid licensing for AI training with major tech firms

The Wikimedia Foundation expanded Wikimedia Enterprise to offer paid licensing of Wikipedia content to major tech firms—Microsoft, Meta, Amazon, Perplexity, and Mistral AI—letting them use Wikipedia’s 65 million articles to train AI models, joining Google's existing deal. The move monetizes API access to offset rising infrastructure costs as AI scrapes increase, while editors push back on AI-generated content and bot traffic remains a concern.

ElevenLabs Debuts AI Music Generator in Industry Collaboration
technology1 year ago

ElevenLabs Debuts AI Music Generator in Industry Collaboration

ElevenLabs has launched an AI music generator that is claimed to be cleared for commercial use, expanding beyond its traditional text-to-speech tools. The company has shared samples of AI-generated music and announced partnerships with major music publishers to use their material for training, amid ongoing legal concerns about copyright infringement in AI music development.

US Authors Prepare Class-Action Lawsuit Against Anthropic Over AI Book Piracy
legal1 year ago

US Authors Prepare Class-Action Lawsuit Against Anthropic Over AI Book Piracy

A California federal judge has allowed a class-action lawsuit against Anthropic, alleging the AI company downloaded up to seven million copyrighted books from pirated sources to train its chatbot Claude, violating the Copyright Act. The lawsuit, filed by authors including Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, joins a broader trend of legal actions against AI firms over copyright issues, with some cases focusing on unauthorized data use and others on licensing disputes.

"Report: Tumblr and WordPress Strike Deals to Sell User Data for AI Training"
technology2 years ago

"Report: Tumblr and WordPress Strike Deals to Sell User Data for AI Training"

Automattic, the owner of Tumblr and WordPress.com, is reportedly in talks with AI companies Midjourney and OpenAI to provide training data scraped from users' posts, potentially creating a new revenue stream for the site. The company plans to launch a new setting allowing users to opt out of data sharing with third parties, including AI companies. However, it's unclear what data has been sent to the AI companies and how it has been used. This move reflects a trend of companies striking deals with AI tool makers for training data, but it has also sparked concerns from the creative community about their work being used for training. Automattic has struggled to monetize Tumblr since acquiring it in 2019 and is seeking new avenues for revenue.