> And no, no AI company has ever come to us and asked to run training on all of our scanned copies
The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.
There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.
Agreed, in addition LLMs are trained on a lot of useless data, such as the whole Reddit dataset which has to be at least 90% Reddit garbage, and that doesn’t seem to be a problem. I don’t see why ai labs wouldn’t want to also train on older data
The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.
There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.