undefined | Better HN

0 pointsneurostimulant2y ago0 comments

There is a rumor that OpenAI might've used libgen in their training data.

0 comments

3 comments · 1 top-level

mcculley2y ago· 2 in thread

Someone will. The potential gains are too high to ignore it.

nojvek2y ago

We are talking about trillions of tokens.

I’m sure the big players like Google, Meta, OpenAI have used anything and everything they can get their hands on.

Libgen is a wonder of the internet. I’m glad it exists.

mcculley2y ago

I am also glad that libgen exists. Liberating human knowledge from copyright will improve humanity overall.

But I don’t understand how you can be sure that the big players are using it as a training corpus. Such an effort of questionable legality would be a significant investment of resources. Certainly as the computronium gets cheaper and techniques evolve, bringing it into reach of entities that don’t answer to shareholders and investors, it will happen. What makes you sure that publicly owned companies or OpenAI are training on libgen?

j / k navigate · click thread line to collapse