undefined | Better HN

0 pointsretetr1y ago0 comments

Unrelated, but is this a case of the Pareto Principle? (Admittedly the first time I'm hearing of it) Wherein 80% of the effect is caused by 20% of the input. Or is this more a case of diminishing returns? Where the initial results were incredible, but each succeeding iteration seems to be more disappointing?

0 comments

1 comments · 1 top-level

klabb31y ago

Pareto is about diminishing returns.

> but each succeeding iteration seems to be more disappointing

This is because the scaling hypothesis (more data and more compute = gains) is plateauing, because all text data is used and compute is reaching diminishing returns for some reason I’m not smart enough to say why, but it is.

So now we're seeing incremental core model advancements, variations and tuning in pre- and post training stages and a ton of applications (agents).

This is good imo. But obviously it’s not good for delusional valuations based exponential growth.

1 more reply

j / k navigate · click thread line to collapse