No need for a data dump, just list all URLs or whatever else of their training data sources. Afaik that's how the LAION training dataset was published.
providing a large list of bitrotted URLs and titles of books which the user should OCR themselves before attempting to reproduce the model doesn't seem very useful.