DeepSeekMath 7B achieved 51.7% on MATH benchmark (opens in new tab)

(github.com)

114 pointsmdp2y ago39 comments

39 comments

22 comments · 4 top-level

deepseekfake2y ago· 10 in thread

I have spoken to team members, and they all say the results of this and coder are very, very much leakage (no suprisse given the result!!)

godelski2y ago

That's good to know, and better to admit. Earns a lot of respect, at least for me. Recall is still a pretty useful task. I just wish more would be less afraid to admit spoilage.

pclmulqdq2y ago

There's a good chance that's also true for GPT-4 given how they train. Without known completely new evals, it's hard to say that any LLM benchmark results aren't leakage.

godelski2y ago

This is most certainly true. If you look back to my comment and the discussion from the main thread I have two quotes from the GPT 4 technical paper

> We measure cross-contamination between our evaluation dataset and the pre-training data using substring match. Both evaluation and training data are processed by removing all spaces and symbols keeping only characters (including numbers). For each evaluation example, we randomly select three substrings of 50 characters (or use the entire example if it’s less than 50 characters). A match is identified if any of the three sampled evaluation substrings is a substring of the processed training example. This yields a list of contaminated examples. We discard these and rerun to get uncontaminated scores.

> The RLHF post-training dataset is vastly smaller than the pretraining set and unlikely to have any particular question contaminated. However we did not check explicitly.

These are not great at building confidence that OpenAI does not have spoilage. Given what we know about the dedupe process (even from early 2023) this is not enough to purge contamination. Exact string matching has been the de facto method for quite some time and for quite some time we've known that this has issues. Just that 5 years ago these issues weren't as critical as they are today because performance was much lower back then.

1 more reply

CuriouslyC2y ago

If you're trying to prove the model has reasoning abilities, ask it the question in a language other than English, even better give it multiple sentences in different languages and tell it to answer the question without first translating the sentences.

1 more reply

rgbrgb2y ago

yea, that's my first thought seeing the result too. we need a reputable proprietary eval.

godelski2y ago

> reputable proprietary eval

I think this is self-conflicting. If the evaluation is proprietary then it is most certainly not reputable. We'd want open metrics where we can analyze the limitations. Of course, we'd need open data too, but that's exceptionally rare these days. Plus, a metric isn't going to really tell us if we have have spoilage or not. You can get some evidence for spoilage through a trained model, but it is less direct, fuzzier, and more tells us about what information it was able to memorize rather than if the data was spoiled.

1 more reply

riku_iki2y ago

It's actually interesting results in a sense we see the limitation of LLM to memorize complicated information correctly. Gemini ultra also reported around 50% accuracy

Davidzheng2y ago

I think the SOTA is GPT4+tool use? I heard near 80%

1 more reply

WiSaGaN2y ago

So you created this account just to make this comment.

gowld2y ago

Why haven't they updated the github page?

godelski2y ago· 4 in thread

Does anyone know how much spoilage are in these datasets? Common crawl has a lot of websites in it, including Reddit and Stack*. I'm certain there are lots of questions in those datasets and we want to differentiate recall from problem solving (often confused). I have a deep distrust when using large datasets like this given a common one with 60 authors assumed writing leet code style programs by hand would mean they wouldn't appear in the training data (github) and didn't even bother to check. It's really hard to sanitize datasets of this size and deduplication is a much harder task than many realize.

https://arxiv.org/abs/2107.03374

https://arxiv.org/abs/2303.09540

Imnimo2y ago

The paper claims:

>To avoid benchmark contamination, we follow Guo et al. (2024) to filter out web pages containing questions or answers from English mathematical benchmarks such as GSM8K (Cobbe et al., 2021) and MATH (Hendrycks et al., 2021) and Chinese benchmarks such as CMATH (Wei et al., 2023) and AGIEval (Zhong et al., 2023). The filtering criteria are as follows: any text segment containing a 10-gram string that matches exactly with any sub-string from the evaluation benchmarks is removed from our math training corpus. For benchmark texts that are shorter than 10 grams but have at least 3 grams, we employ exact matching to filter out contaminated web pages.

However, benchmark contamination is difficult, and ngram matching is often insufficient. See https://arxiv.org/pdf/2311.04850.pdf for some examples of how this approach can fail.

In general, if a benchmark is available online before a model's dataset is collected, I put very little stock into that model's performance on that benchmark. It's just too hard to know what's a true improvement and what's contamination. It's especially true for a paper like this that specifically hunts down MATH-like data.

godelski2y ago

So I think we're in agreement and I find very little discussion about this within the community (being a researcher myself). This wouldn't particularly bug me if we were saying that the measurements do not distinguish the ability to recall with generalization, but I find that the discussion is always about generalization and AGI, leading to a very confused public.

Unfortunately I'm just not aware of any metric that can adequately quantify meaningful similarity between data. Curse of dimensionality I suppose. Personally I try not to lean too hard on benchmark results not only because the aforementioned spoilage, but due to metric limitations as well. Personally I think our progress has out paced our ability to properly measure and it feels like we've only become more reliant upon them rather than more nuanced in our evaluations (am I alone in this?). I am wondering if this will create a stall or plateau (or even reversal) in practical performance as our measurements become less meaningful as our quality increases. I'm in vision, so a good example is how it is common to think that the L2 distance between a norm layer of a classification network (even if better than InceptionNet) is an accurate measurement of visual fidelity. Or to even think we have such metrics in even special cases (I guess PSNR or SSID are closest but that's more accurately described as reconstruction quality).

Btw, I think you might like the second paper I linked. It's a META/Stanford paper and mostly deals with vision (LAION) but a bit with C4. The short of it is that they can prune about 40% of LAION and still get good "Zeroshot" ImageNet accuracy. I actually found the results for random pruning quite enlightening, especially around all the toy datasets (Fig A4).

Zeroshot in quotes because it's pretty dubious to call ImageNet out of distribution (same with COCO) when a model is trained on LAION considering all the classes (at least an abstracted version of the class since LAION is more specific. i.e. ImageNet _distribution_ ⊂ LION _distribution_).

Another pet peeve of mine is arxiv links direct to PDF ;)

ianbutler2y ago

Some have a lot and those models are often ignored (except by lay people or hobbyists which is a different problem), but many serious submissions from serious groups for benchmarks like this check for contamination to specifically avoid the problem you’re suggesting. Process for decontamination has been outlined by many groups so you can often check it out.

godelski2y ago

So I am a ML researcher. Note that part of my comment is specifying how difficult it actually is to ensure lack of spoilage. The second paper I link is actually a pretty good proof of this. Though I'll say that I wish they had been a bit more explicit about how a random pruning significantly improves results. Because that is quite the result in of itself, given that the datasets they look at are already filtered. Dedupe is fucking hard. So I'm not looking for a handwavy "trust me" I'm looking for the explicit vetting processes applied to these specific datasets. It's incredibly important to know the limits of your tools and that includes datasets (as well as metrics).

2 more replies

rgbrgb2y ago· 4 in thread

Supports commercial use!

Interesting what's unsupported:

- In any way that violates any applicable national or international law or regulation or infringes upon the lawful rights and interests of any third party;

- For military use in any way;

- For the purpose of exploiting, harming or attempting to exploit or harm minors in any way;

- To generate or disseminate verifiably false information and/or content with the purpose of harming others;

- To generate or disseminate inappropriate content subject to applicable regulatory requirements;

- To generate or disseminate personal identifiable information without due authorization or for unreasonable use;

- To defame, disparage or otherwise harass others;

- For fully automated decision making that adversely impacts an individual’s legal rights or otherwise creates or modifies a binding, enforceable obligation;

- For any use intended to or which has the effect of discriminating against or harming individuals or groups based on online or offline social behavior or known or predicted personal or personality characteristics;

- To exploit any of the vulnerabilities of a specific group of persons based on their age, social, physical or mental characteristics, in order to materially distort the behavior of a person pertaining to that group in a manner that causes or is likely to cause that person or another person physical or psychological harm;

- For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories.

ronsor2y ago

The irony is that anyone who was going to do those things isn't going to care about a license anyway.

BossingAround2y ago

True, but at least the author wouldn't be liable.

2 more replies

austin-cheney2y ago

Who would qualify as a user that seeks to violate applicable laws and yet is somehow identified as an official part of some legally recognized military? Furthermore, how would anybody know?

As a dumb Army guy if I were doing military research I would just keep it on my private military internet that does not exist for non-military users.

wand3r2y ago

Its virtue signaling. I know its over used, but seriously, who is intentionally harming minors BUT unwilling to break a ToS contract?

4 more replies

mdpOP2y ago

Related paper - https://arxiv.org/pdf/2402.03300.pdf

j / k navigate · click thread line to collapse

39 comments

22 comments · 4 top-level

deepseekfake2y ago· 10 in thread

I have spoken to team members, and they all say the results of this and coder are very, very much leakage (no suprisse given the result!!)

godelski2y ago

That's good to know, and better to admit. Earns a lot of respect, at least for me. Recall is still a pretty useful task. I just wish more would be less afraid to admit spoilage.

pclmulqdq2y ago

There's a good chance that's also true for GPT-4 given how they train. Without known completely new evals, it's hard to say that any LLM benchmark results aren't leakage.

godelski2y ago

This is most certainly true. If you look back to my comment and the discussion from the main thread I have two quotes from the GPT 4 technical paper

> The RLHF post-training dataset is vastly smaller than the pretraining set and unlikely to have any particular question contaminated. However we did not check explicitly.

1 more reply

CuriouslyC2y ago

1 more reply

rgbrgb2y ago

yea, that's my first thought seeing the result too. we need a reputable proprietary eval.

godelski2y ago

> reputable proprietary eval

1 more reply

riku_iki2y ago

It's actually interesting results in a sense we see the limitation of LLM to memorize complicated information correctly. Gemini ultra also reported around 50% accuracy

Davidzheng2y ago

I think the SOTA is GPT4+tool use? I heard near 80%

1 more reply

WiSaGaN2y ago

So you created this account just to make this comment.

gowld2y ago

Why haven't they updated the github page?

godelski2y ago· 4 in thread

https://arxiv.org/abs/2107.03374

https://arxiv.org/abs/2303.09540

Imnimo2y ago

The paper claims:

However, benchmark contamination is difficult, and ngram matching is often insufficient. See https://arxiv.org/pdf/2311.04850.pdf for some examples of how this approach can fail.

godelski2y ago

Another pet peeve of mine is arxiv links direct to PDF ;)

ianbutler2y ago

godelski2y ago

2 more replies

rgbrgb2y ago· 4 in thread

Supports commercial use!

Interesting what's unsupported:

- In any way that violates any applicable national or international law or regulation or infringes upon the lawful rights and interests of any third party;

- For military use in any way;

- For the purpose of exploiting, harming or attempting to exploit or harm minors in any way;

- To generate or disseminate verifiably false information and/or content with the purpose of harming others;

- To generate or disseminate inappropriate content subject to applicable regulatory requirements;

- To generate or disseminate personal identifiable information without due authorization or for unreasonable use;

- To defame, disparage or otherwise harass others;

- For fully automated decision making that adversely impacts an individual’s legal rights or otherwise creates or modifies a binding, enforceable obligation;

- For any use intended to or which has the effect of discriminating against individuals or groups based on legally protected characteristics or categories.

ronsor2y ago

The irony is that anyone who was going to do those things isn't going to care about a license anyway.

BossingAround2y ago

True, but at least the author wouldn't be liable.

2 more replies

austin-cheney2y ago

Who would qualify as a user that seeks to violate applicable laws and yet is somehow identified as an official part of some legally recognized military? Furthermore, how would anybody know?

As a dumb Army guy if I were doing military research I would just keep it on my private military internet that does not exist for non-military users.

wand3r2y ago

Its virtue signaling. I know its over used, but seriously, who is intentionally harming minors BUT unwilling to break a ToS contract?

4 more replies

mdpOP2y ago

Related paper - https://arxiv.org/pdf/2402.03300.pdf

j / k navigate · click thread line to collapse