It is just as legal as when Uber and AirBNB were running illegal taxis and hotels during their growth phase. I'm just waiting for some corporate IP law firm to learn about Huggingface.
bandwidth and storage are literally free when compared to the cost of GPU clusters. HF gets rewarded heavily on capital market for being in AI without actually doing much AI stuff, that is a huge win when compared to costs they are paying for bandwidth and storage.
To be precise, Amazon Cloudfront is the CDN. Maybe they got some startup deal?
Amazon does now also have flat rate plans that are a lot cheaper.
Presumably they already know. The issue is that IP law firms are tiny compared to the trillions of capital pouring into "AI". And if you believe the USA is a capitalist country where the side with deeper pockets win, you know you're not going to win against the trillionaires.
Probably as a base to use by people buying NVIDIA hardware to train their own.
Open-source data coverage: The released datasets cover an estimated 8–10T tokens
(~40–50% of the internal 25T blend). Missing categories include code (~14% of blend),
nemotron-cc-code (~2%), crawl++ (~2%), and academic text (~2%). Users should
supplement with their own data for these categories and adjust train_iters
accordingly.
Nemotron is the strongest model (on most benchmarks) that has its full training pipeline and most of the data open. Olmo 3 from AllenAI, and K2 Think V2 from Mohamed bin Zayed University of Artificial Intelligence are both fully open, but not as capable as the Nemotron family. Granite has much of the training pipeline and data open, but is missing some of each.