Harvard Unveils 1 Million Public-Domain Books for AI Training

TL;DR Summary
Harvard University has launched a dataset of nearly one million public domain books for AI model training, funded by Microsoft and OpenAI. This initiative aims to provide legal data sources for AI development amidst ongoing legal challenges faced by companies like OpenAI for using copyrighted material without permission. The dataset includes classics and diverse texts, but AI companies still seek exclusive, modern data to enhance their models. The project highlights the tension between data accessibility and copyright protection in AI training.
- Harvard Makes 1 Million Books Available to Train AI Models Gizmodo
- Harvard Is Releasing a Massive Free AI Training Dataset Funded by OpenAI and Microsoft WIRED
- Supporting New Open Data Initiatives: Institutional Data Initiative and CORE Microsoft
- Harvard and Google to release 1 million public-domain books as AI training dataset TechCrunch
- Harvard’s Library Innovation Lab launches initiative to use public domain data to train artificial intelligence Harvard Law School
Reading Insights
Total Reads
0
Unique Readers
5
Time Saved
3 min
vs 4 min read
Condensed
88%
663 → 81 words
Want the full story? Read the original article
Read on Gizmodo