zlacker

Anthropic apparently did it both ways. After realizing that pirating mass quantities of books for training wasn't a great legal look, it hired someone previously responsible for Google Books, who in turn contacted publishers about mass licensing their content for training use.

However, that option was ultimately not pursued as instead...

>> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books). Anthropic created its own catalog of bibliographic metadata for the books it was acquiring. It acquired copies of millions of books, including of all works at issue for all Authors.

(from the ruling)

replies(1): >>bonobo+Q2

>>ethbr1+(OP)
Yes. And from the article "That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for theft but it may affect the extent of statutory damages." Sounds reasonable (except for the "stole" and "theft" language -- a copyright violation is a copyright violation, not theft, not stealing).

If the actual model was trained from the unauthorized copies, and then they post-hoc bought the books, that doesn't retroactively cancel the initial copyright violation. As I understand they did not retrain the model using the OCR'd scans