Harvard Unveils 1 Million Public-Domain Books for AI Training

1 min read
Source: Gizmodo
Harvard Unveils 1 Million Public-Domain Books for AI Training
Photo: Gizmodo
TL;DR Summary

Harvard University has launched a dataset of nearly one million public domain books for AI model training, funded by Microsoft and OpenAI. This initiative aims to provide legal data sources for AI development amidst ongoing legal challenges faced by companies like OpenAI for using copyrighted material without permission. The dataset includes classics and diverse texts, but AI companies still seek exclusive, modern data to enhance their models. The project highlights the tension between data accessibility and copyright protection in AI training.

Share this article

Reading Insights

Total Reads

0

Unique Readers

5

Time Saved

3 min

vs 4 min read

Condensed

88%

66381 words

Want the full story? Read the original article

Read on Gizmodo