A recent experience with ChatGPT 5.5 Pro (opens in new tab)

(gowers.wordpress.com)

728 points_alternator_1mo ago535 comments

https://twitter.com/wtgowers/status/2052830948685676605

https://xcancel.com/wtgowers/status/2052830948685676605

535 comments

275 comments · 58 top-level

ziotom781mo ago· 54 in thread

I am a physics professor and often use Gemini to check my papers. It is a formidable tool: it was able to find a clerical error (a missing imaginary unit in a complex mathematical expression) I was not able to find for days, and it often underlines connections between concepts and ideas that I overlooked.

However, it often makes conceptual errors that I can spot only because I have good knowledge of the topic I am discussing. For instance, in 3D Clifford algebras it repeatedly confuses exponential of bivectors and of pseudoscalars.

Good to know that ChatGPT 5.5 Pro can produce a publishable paper, but from what I have seen so far with Gemini, it seems to me that it is better to consider LLMs as very efficient students who can read papers and books in no time but still need a lot of mentoring.

nopinsight1mo ago

I assume you're using the "regular" Pro version of Gemini 3.1 for the above, rather than the Deep Think mode, which is more comparable to GPT-5.5 Pro. To my knowledge, regular 3.1 Pro is a tier below and often makes mistakes.

Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon.

You might want to track the progress of these models on the CritPt benchmark, which is built on *unpublished, research-level* physics problems:

https://critpt.com/

Frontier models are still nowhere near solving it, but progress has been rapid.

* o3 (high) <1.5 years ago was at 1.4%

* GPT 5.4 (xhigh), 23.4%

* GPT-5.5 (xhigh), 27.1%

* GPT-5.5 Pro (xhigh) 30.6%.

https://artificialanalysis.ai/evaluations/critpt.

FrojoS1mo ago

> there's no reason to believe the progress of LLMs [...] will stop anytime soon

Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

14 more replies

civvv1mo ago

There are many indications that model progress is slowing down, so that is not entirely accurate.

3 more replies

Davidzheng1mo ago

Deep think still makes many many many more mistakes than gpt 5.5 pro on math

maximamas1mo ago

LLMs are at their best when you have an expectation for their output. I generally know the shape of the correct response and that allows me to evaluate it's output on it's "vibes", rather than line by line. If there's no expectation then I have to take everything at face value and now I'm at the mercy of the machine.

jillesvangurp1mo ago

Exactly, if I generate a large chunk software, I'm going to have expectations about what it will do, how it will do it, etc. You don't just accept the statement that "it's done" for fact but you start looking for evidence.

A scientific approach here is to look to falsify the statement. You start asking questions, running tests, experiments, etc. to prove the notion that it is done wrong. And at some point you run out of such tests and it's probably done for some useful notion of done-ness.

I've built some larger components and things with AI. It's never a one shot kind of deal. But the good news is that you can use more AI to do a lot of the evaluation work. And if you align your agents right, the process kind of runs itself, almost. Mostly I just nudge it along. "Did you think about X? What about Y? Let's test Z"

2 more replies

ziotom781mo ago

I agree, but I would add that they can be very useful even if you do not have clear expectations but have some solid ways to verify their claims. Often in doing this verification I came up with new ideas.

tags2k1mo ago

I'm no physics professor but this aligns with the way I use the tools in my "senior engineer" space. I bring the fundamentals to sanity-check the trigger-happy agent and try to imbue other humans with those fundamentals so they can move towards doing the same. It feels like the only way this whole thing will work (besides eventually moving to local models that do less but companies can afford).

illiac7861mo ago

Using the word “Mentoring” is anthropomorphic and subconsciously makes you think it will learn. It does not, and it is for the human brain a formidable task to remember that something as smart as an LLM does not learn. I keep catching myself making the same mistake.

It’s also because it is so annoying to have to manage the memory of the LLM with custom prompts/instructions manually.

I have not yet played with the long term memory feature, but I fear it will be even less reliable than prompts, simply because in one year or two years so much will have changed again that this “memory” will have to be redone multiple times by then.

timschmidt1mo ago

They can form new associations between concepts via their input prompts and thinking text. That is a form of learning. Just not very durable. I liken it to https://en.wikipedia.org/wiki/Anterograde_amnesia

illiac7861mo ago

yeah, I should have been more specific: I meant the type of learning that mentoring fosters, the long term learning.

2 more replies

kybernetikos1mo ago

Current LLM architecture doesn't learn - and you're right this is a huge piece that normal folks fail to understand, since in many ways, it's the opposite of what years of AI research has been trying to create.

However, I think it's important to remember that LLMs are embedded in larger systems, and those larger systems do learn.

2 more replies

freedomben1mo ago

I mostly agree, though after a mentoring session you can ask it to write skill or a memory and it can be reasonably durable. For Claude at least, the memories work pretty well (though I am still at a small scale with them. As they grow it might start to break somewhat. Doesn't always work, but has often enough that I thought it worth a mention.

stingraycharles1mo ago

> Using the word “Mentoring” is anthropomorphic and subconsciously makes you think it will learn.

I think this is a bit pedantic. Obviously the parent you’re replying to is referring to the concept of “in-context learning”, which is the actual industry / academic term for this. So you feed it a paper, and then it can use that info, and it needs steering / “mentoring” to be guided into the right direction.

Heck the whole name of “machine learning” suggests these things can actually learn. “reasoning” suggests that these things can reason, instead of being fancy, directed autocomplete. Etc.

In other news: data hydration doesn’t actually make your data wet. People use / misuse words all the time, and that causes their meaning to evolve.

kasey_junk1mo ago

I agree it’s pedantic and personally don’t get bent out of shape with people anthropomorphizing the llms. But I do think you get better results if keep the text prediction machine mental model in your head as you work with them.

And that can be very hard to do given the ui we most interact with them in is a chat session.

1 more reply

DoctorOetker1mo ago

> ... that something as smart as an LLM does not learn.

what? training is learning, as long as weights are available continual learning is perfectly feasible: just keep training the LLM with the user corpus alternated with a frozen version to prevent catastrophic drift / collapse.

it's not because model providers don't want to provide user specific continual learning, that we don't know how to do it.

it would be a lot more expensive to host user-specific model weights, and would prevent amortizing the weights over many requests in batches...

_the_inflator1mo ago

I agree and put it this way: LLMs sound so convincing presenting you the work it does rose colored and promising to give you more if you keep going.

There is a 50/50 chance that it turns out to be right or letting you jump of the cliff.

Only the trip stays the same beautiful 5 star plus travel.

Also, spotting an error and telling LLM makes it in most cases worse, because the LLM wants to please you and goes on to apologize and change course.

The moment I find myself in such a situation I save or cancel the session and start from scratch in most cases or pivot with drastic measures.

Gemini to me is the most unpredictable LLM while GPT works best overall for me.

Gemini lately gave me two different answers to the same question. This was an intentional test because I was bored and wanted to see what happens if you simply open a new chat and paste the same prompt everything else being the same.

Reasoning doesn’t help much in the Coding domain for me because it is very high level and formally right what the LLM comes up with as an explanation.

I google more due to LLMs than before, because essentially what I witnessed is someone producing something that I gotta control first before I hit the button that it comes with. However, you only find out shortly afterwards whether the polished button started working or gave you a warm welcome to hell.

MattPalmer10861mo ago

Reusing the same prompt several times is something I've started doing too. The contrast is often illuminating.

In one case, it made a thoroughly convincing argument that an approach was justified. The second time it made exactly the opposite argument, which was equally compelling.

I now see LLMs as persuasion machines.

scotty791mo ago

Before AI happened I watched youtube. Occasionally I encountered there very convincing arguments. Same person often made very convincing arguments on many subjects.

But noticed that the closer the domain they were talking about was to my area of competence the less convincing their arguments were. There were more holes, errors and wrong conclusions.

I recalibrated my bs meter thanks to that.

Since AI came I successfully used this strategy of being extremely cautious towards convincing arguments to not become mislead by AI.

However this year I'm working with AI more in the domain of software development. Where I can see the competence. And I see the competence. This had opposite effect on me. I tend to trust AI outside my domain of expertise much more after I saw what can it do in software.

One caveat though is that there are a lot of areas of human culture where there's very little actual knowledge, but a lot of opinions, like politics, economy, diet, business, health. I still don't trust AI in those domains. But then again, I don't trust humans there either.

For me basically AI achieved the threshold of useful reliability for any domain that humans are reliable at.

I don't really care about sycophancy. I might have a slight advantage that I don't talk to AI in my native language. So its responses don't have a direct line to my emotions.

eitally1mo ago

One thing I've been doing lately -- and I'm in a business function, not a technical one, although I have an engineering background -- is pitting LLMs against each other. For example, if I'm structuring a proposal or a contract with the assistance of Claude, I'll begin my 360 feedback review first by asking Claude how it would react if it were the counter-party receiving the proposal. After some iterative changes, mostly manual, I will then run the same output document past Gemini and ask it to adopt personas from both sides and provide reactive feedback. The result of this is almost always a stronger proposal that I can also accompany with proactive objection handling and a solid FAQ, as well as clear points of negotiation that will likely be acceptable to both parties.

For this sort of thing, using multiple LLMs is extremely helpful.

taneq1mo ago

Ever since they started getting really sycophantic, I’ve been presenting my ideas as “my co-worker says this is a good approach but I disagree, can you help me convince him that it’s wrong?”

pbhjpbhj1mo ago

>LLM wants to please you

I was using Copilot and asked it a question about a PDF file (a concept search). It turned out the file was images of text. I was anticipating that and had the text ready to paste in.

Instead, it started writing an OCR program in python.

I stopped it after several minutes.

Often Copilot says it can't do something (sometimes it's even correct), that's preferential to the try-hard behaviour here.

freedomben1mo ago

> Gemini to me is the most unpredictable LLM while GPT works best overall for me.

This nails an important thing IMHO. I've absolutely noticed this, for better or worse. Gemini can produce surprisingly excellent things, but it's unpredictability make me go for GPT when I only want to ask it once.

miki1232111mo ago

I think that ultimately, the largest change brought on by LLMs will be due not to their intelligence, but to their tenacity.

If you had an infinite number of monkeys, each with a typewriter, one would eventually write Shakespeare. If you had an infinite number of college-educated interns, each with access to all the public records you can possibly get via FOIA, one would eventually get enough evidence to prove that a top politician is cheating on their partner, evidence which you could use to blackmail that politician.

You don't need that much intelligence to do that, you just need somebody who's willing to dedicate their life to knowing everything there is to know about that guy from Louisiana.

With humans, the amount of money you'd need to pay such a person just isn't worth the reward. With LLMs, it may very well be.

mixtureoftakes1mo ago

please, sign up for a paid plan of either chatgpt or claude. gemini is while close, still noticeably behind

you deserve opinions shaped by interactions with the best tools that are out there.

wg01mo ago

Gemini feels deep and philosophical. Especially for product management. Tell him you're a product manager and we're a team of two.

But regular reminder - All LLMs can be wrong all the time. I only work with LLMs in domains I'm expert in OR I have other sources to verify their output with utmost certainty.

wafflemaker1mo ago

Or when you don't care about results being very correct.

When I'm cooking meatballs with sauce and the recipe calls for frying them, I'll have an LLM guestimate how long and which program to use in an air fryer to mimic the frying pan, based on a picture of balls in a Pyrex. So I can just move on with the sauce, instead of spending time browsing websites and stressing about getting it perfect.

I used to hate these non-deterministic instructions, now I treat it as their own game. When I will publish my first recipe, I'll have an LLM randomize the ingredient amounts, round them up to some imprecise units and also randomize the times. Psychologists say we artists need to participate and I WILL participate.

smartmic1mo ago

> I only work with LLMs in domains I'm expert in

This. Should become a general rule for any non-trivial use of LLM in a professionel setting.

1 more reply

peyton1mo ago

Seriously, it’s not worth reaching for less intelligence. Use Extended Pro 100% of the time for things you’d spend the amount of time GP spent writing their post.

cubefox1mo ago

Gemini is certainly not behind Claude in terms of physics.

ainch1mo ago

Agreed, Gemini is clearly a capable model, but the tool use is lagging behind the other two. Ironically it regularly gets things wrong (ie. the current version of some software) because of an unwillingness to use web search.

hodgehog111mo ago

ChatGPT and Gemini are actually fairly comparable.

Claude has been utterly useless with most math problems in my experience because, much like less capable students, it tends to get overly bogged down in tedious details before it gets to the big picture. That's great for programming, not so much for frontier math. If you're giving it little lemmas, then sure it's great, but otherwise you're just burning tokens.

Quothling1mo ago

We've got a rather extensive AI setup through our equity fund and I've setup a group of agents for data architecture at scale. One is the main agent I discuss with and it's setup to know our infrastructure and has access to image generation tools, websearch, hand off agents and other things. I tend to use Opus (4-6 currently) and I find it to be rather great. As you point out it comes with the danger of making mistakes, and again, as you point out, it's not an issue for things I'm an expert on. What I rely on it for, however, is analysing how specific tools would fit into our architecture. In the past you would likely have hired a group of consultants to do this research, but now you can have an AI agent tell you what the advantages and disadvantages of Microsoft Fabric in your setup. Since I don't know the capabilities of Fabric I can't tell if the AI gives me the correct analysis of a Lakehouse and a Warehouse (fabric tools).

What I do to mitigate this is that I have fact checking agents configured to be extremely critical and non-biased on Opus, Gemini and GPT. Which are then handed the entire conversation to review it. Then it's handed off to a Opus agent which is setup to assume everything is wrong. After this, and if I'm convinced something is correct I'll hand the entire thing off to a sonnet agent, which is setup to go through the source material and give me a compiled list of exactly what I'll need to verify.

It's ridicilously effective, but I do wonder how it would work with someone who couldn't challenge to analytic agent on domain knowledge it gets wrong. Because despite knowing our architecture and needs, it'll often make conceptional errors in the "science" (I'm not sure what the English word for this is) of data architecture. Each iteration gets better though, and with the image generation tools, "drawing" the architecture for presentations from c-level to nerds is ridiclously easy.

stasomatic1mo ago

Are you using this agent hive for any repeatable tasks? What you described, superficially, seems like a one off. Genuinely curious.

1 more reply

cyanydeez1mo ago

I've been watching the automation of things like flight control systems for the past decade, and the evolution of the fallback to a real pilot in the event of a emergency is what's most concerning about where LLMs are being embedded.

Right now, we have a lot of smart people who have trained for decades to understand where these things go wrong and how to nudge them back, but the pool of people are going to slowly be replaced by less knowledgeable.

At some point, a rubicon will be crossed where these systems can't fallback to a human operator and will fail spectacularly.

pbhjpbhj1mo ago

Watching a teenager approach their homework, instead of struggling to answer questions they don't know, they ask Gemini. Unfortunately, I think the mental struggle to approach an answer is where much of the learning is. They also miss out on the reward for persistence of seeing things fall together.

It is troubling. It suggests a plateauing of human understanding.

regularfry1mo ago

It absolutely is where the learning is, that's pretty well established brain science.

1 more reply

regularfry1mo ago

What that means practically is that we've got a generation - 25 years or less - to evolve these things not to need the fallback. If such a thing is possible.

leptons1mo ago

We're on the road to Idiocracy.

wccrawford1mo ago

This doesn't surprise me since the coding agents are similar. I've previously compared them to very fast, ambitious junior programmers. I think they are probably mid-level coders now, but they continue to make mistakes that a senior programmer wouldn't. Or at least shouldn't.

recursivecaveat1mo ago

This is close to my experience with code. LLMs can pick out small mistakes from giant code changes with surprising accuracy, or slowly narrow down a weird. On the other hand I've seen them bravely shoulder on under completely incorrect conceptual models of what they're working with and churn around in circles consequently, spin up giant piles of slop to re-implement something they decided was necessary, but didn't bother to search for, or outright dismiss important error signals as just 'transient failures'. Unlimited stamina, low wisdom.

northzen1mo ago

Hi ziotom! I wonder about you work in 3D Cifford Algebras. May you share some links to the research you do? I also have interest in this topic I research on my own.

Just in case if you don't want to disclose your name my email is northzen@gmail.com

port111mo ago

Gemini’s smug and over-confident “this is the gold standard in 2026” definitely leaves little space for nuance if you don’t know the subject matter. Human students would, hopefully, know they don’t know everything.

quantummagic1mo ago

> Gemini’s smug...

Anthropomorphizing these systems is dangerous, whether coming from the bullish or bearish perspective. The output is statistically generated by a machine lacking the capability to be smug.

2 more replies

danielparsons1mo ago

I find exactly the same for legal analysis. Great at ideation and proofreading but frequently misunderstands concepts and hallucinates conclusions from faulty premises.

wslh1mo ago

I assume that once LLMs are trained with large [synthetic] information about 3D Clifford algebras it will work better.

tasuki1mo ago

> in 3D Clifford algebras it repeatedly confuses exponential of bivectors and of pseudoscalars.

I have no idea what any of those words even mean. I'm sure LLMs make similar obvious-to-professors mistakes in all the domains. Not long ago, we didn't even have chatbots capable of basic conversation...

jiggawatts1mo ago

Ironically, it's sort of the other way around! Every frontier chatbot since GPT 4 (at least) has had a pretty good understanding of even very esoteric technical concepts.

Bivectors and pseudoscalars (in a 3D context) are "just" signed areas and volumes. Easy!

Back around the GPT 3, 3.5, and 4.0 era I used to ask the bots to explain "counterfactual determinism", which is one of the most complex topics I personally understand.

Then I would lie to the bot about it, and see if it corrected me or not.

This test is useless now, the frontier models can't be fooled any longer on such "basic" concepts.

Conversely, LLMs are basically useless at anything that doesn't have enough (or no) public information for their training. Think: obscure proprietary product config files and the like, even if the concepts involved are trivial.

Similarly, Clifford Algebra is a relatively niche (even "alternative") area of mathematics and physics, with vastly less written material about it than the competing linear algebra. Hence, the AIs are bad at it.

eth0up1mo ago

Any experience with NotebookLM?

Mine has been epically bad.

DeathArrow1mo ago

I don't think the experience with Gemini will be the same when using GPT.

wood_spirit1mo ago

Chiming in to agree but clarify that the latest sota models are no better than Gemini.

I put my stuff through several sota models and round robin them in adversarial collaboration and they are all useful even though, fundamentally, they don’t “understand” anything. But they are super useful delegates as long as deciding on the problem and approach and solution all sits safely in your head so you can challenge them and steer them.

So I know the article is about one particular new model acing something and each vendor wants these stories to position their model as now good enough to replace humans and all other models, but working somewhere where I am lucky enough to be able to use all the sota models all the time, I can say that all keep making obvious mistakes and using all adversarially is way better than trusting just one.

I look forward to the day one a small open model that we can run ourselves outperforms the sum of all today’s models. That’s when enough is enough and we can let things plateau.

energy1231mo ago

Basically all Erdos problems that get solved with AI use ChatGPT 5.* Pro, not Gemini/Opus.

1 more reply

ed_balls1mo ago

intern that never sleeps

ieieaaa1mo ago

LLM’s are the most powerful tool invented to search across a huge information space in response to human input.

That’s all they are. They don’t ‘know’ anything intrinsically and do know ‘know’ what reasoning even is.

NotOscarWilde1mo ago· 27 in thread

As a TCS assistant professor from Eastern Europe, I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models.

Paying for Pro from any of my current academic budgets is completely ouf of the field of reality here -- all budgets tend to have restricted uses and software payments fit into very few categories. Effectively, I'd have to ask for a brand new grant and hope the grant rules allow for large software payments and I won't encounter an anti-AI reviewer; such a thing would take one year at least.

As a nail to the coffin, I was "denied" all Claude Opus recently as part of Microsoft's clampdown on individual (and academic) use of Copilot.

(Chagpt 5.5 Plus does not seem sufficient for any deeper investigations into new research topics, I've tried.)

Apologies for the rant.

vthallam1mo ago

@NotOscarWilde drop your email here, I will reach out and happy to get you a pro account for a few months so you can try 5.5 pro.(work at OAI)

teiferer1mo ago

While this sounds generous (and in some ways it is), it does not address the general point that GP is making. That is, the systematic disadvantage which large parts of humanity have w.r.t. to access to the tools. You could say they can't drive a Lambhorgini either, but that also doesn't solve the problem.

NotOscarWilde1mo ago

You're absolutely right (pun intended).

An aside: It was a very nice gesture and completely unexpected by me, so even if it doesn't work out, it made my day. I personally believe that kind gestures have a lot of power.

Back on topic: There is a real danger of the gap between rich and poor universities significantly widening in all fields if the rich can afford Pro level models, or even hardware that can run their own comparable models, and this being fiscally inaccessible to the rest.

One can sweep this under the rug by blaming the educational funding but this just shoots down all discussion. Even if GDP of a country goes up by a lot -- such as Poland -- it takes time before any budget benefit trickles to the education budget, and with some governments it might never do.

I believe Microsoft et al do have the most power here to boost affordable access to AI for researchers on a large scale; the fact that they cut some too expensive models (Opus, 5.5) from their academic benefits package is a grim omen. I do realize they would like universities to pay them also, and ultimately the universities should do that -- but then we are back at the institutional level of the problem.

1 more reply

Scea911mo ago

Its a problem of the individual institutions and countries. The budget required for AI tools currently is negligible compared to other university expenses. We don't need to call everything a systemic disadvantage when the disadvantaged (at the institution level) have agency here.

2 more replies

snayan1mo ago

I mean, I don't think OpenAI should be wading into the policies and practices of foreign institutions and governments. Look at all the blowback we see from the collision of Anthropic or OpenAI and the US government.

At present, the tools are available for whomever wants to buy them. Not OpenAI's fault that parent comment's government and/or institutions policies haven't been updated to allow for their purchase and use.

I'd argue that the OpenAI dude/dudettes level of generosity is appropriate given the circumstances.

NotOscarWilde1mo ago

This requires a major "dox" of myself, but I am really grateful for the offer, so these are my academic contacts:

https://pastebin.com/hNYrCjhL

I probably will erase the contents in a few days.

Even if you just drop an email and it doesn't work out, I appreciate this gesture so much. Thank you.

vthallam1mo ago

Got the contact, will reach out tomorrow, you can delete them.

thierrydamiba1mo ago

Shoutout to you-I will match it if they need other resources. (I don’t work at OAI, just think this is cool)

alsetmusic1mo ago

You know what, I'm ashamed that I didn't think of this. I'll sponsor three months. Email in my hn profile. I don't understand the math in the article, but I'd love to help you make progress in it.

1 more reply

NotOscarWilde1mo ago

I will leave the contact up for a bit longer if people want to get in touch and share their experience with the research gap of the models -- or anything, really -- but I do not think there is any need of further support. Like I said elsewhere, the offer of support made my day and the gesture is enough.

Thank you.

lelanthran1mo ago

This doesn't solve the problem, though: having the ability to finish a field of study without paying a toll to a token provider.

layer81mo ago

And what should he do after the few months?

johndough1mo ago

At my university, everyone had to pay their AI subscriptions out of their own pocket, until a communal AI service was introduced recently. It took 2 years to set up and only serves gpt-oss-120b, so everyone is still using other services. But at least some admin can scatter the word "AI" all over the university's website now and has an excuse to reject any requests for AI subscriptions because "we already have AI".

alsetmusic1mo ago

It’s a classic example of the best positioned people being in the best position to keep reaping all the rewards.

There’s the example of a poor person and a rich person buying boots. The poor person’s boots wear out and have to be replaced while the rich persons boots last for many years due to higher quality craftsmanship. Over years, the poor person’s boots wear will pay may for boots.

huijzer1mo ago

I know the example, but as a counter-argument: often more expensive boots are not more durable. It’s about spending time to learn to spot the quality.

Of course if you are really poor, then you have to take expensive shortcuts, but for most people that shouldn’t be the case. Learning to do more with less money isn’t as bad as many people think. It’s also good for the brain to be a bit more creative.

NotOscarWilde1mo ago

> Learning to do more with less money isn’t as bad as many people think.

We are wading into philosophy here, but I believe this analogy doesn't track in this case -- my suspicion from this blog post and others is that already today, the Pro level thinking models are a positive multiplier to your research output similar to how the models one level lower are a multiplier to one's programming output.

Maybe one can someday use the cheaper models similar to how you can use cheaper models than Opus/5.5 and still be nearly as productive as a programmer -- but I am trying and failing doing exactly that for research questions.

1 more reply

m_mueller1mo ago

here I think it's less about "poverty" (non-US acedemic budgets are still high, though not in the same sphere), but it's about having red tape when it comes to software. My experience doing a PhD in Japan was: Everything you can touch was basically a free for all - including $500 keyboards and $10k Mac Pros, especially if you are a valued researcher. But software, oh man, how can we prove receipt of goods to accounting...

bambax1mo ago

OpenRouter lets you pay by the token only (no subscription), has all the frontier models (including Opus 4.7, GPT-5.5) and most of the others, and if you use it sparingly it usually turns out to be quite cheap.

johndough1mo ago

API pricing for Claude is about an order of magnitude more expensive than subscriptions (numbers: https://she-llac.com/claude-limits). But it may be worth it with DeepSeek V4 Pro, which is currently on discount.

bambax1mo ago

Depends very much on usage! If you connect it to tools like Cursor, etc. then yes a subscription is probably cheaper -- although, you'd have to subscribe to each provider if you want to use them all.

But if you ask questions occasionally, (and don't resend, for example, your whole codebase with each request), then the API feels really cheap, even for the frontier models.

tasuki1mo ago

My problem with pay-by-the-token is that it discourages me using the thing ("oh the prompt will cost me $0.1"), so I pay a subscription which I'm pretty sure costs me about two-three times what I'd pay just for the api costs, but encourages me to use it more ("oh I have a subscription already, better make use of it").

nerdsniper1mo ago

I believe ChatGPT 5.5 Pro access is available for $100/month, is that an unrealistic level of expense for someone in your position and geography? Even if the university won't pay for it, it seems you'd like to use this tool for your own goals.

I'm not trying to shame here, just curious whether this is completely unattainable for most researchers in your area.

Computer01mo ago

It appears that in their country someone in their position makes about 50k usd annually. I make a similar amount in my country and cannot justify it.

ziotom781mo ago

I fully understand your rant! I pay ~20€/month for the Pro account, as my university has a deal with Microsoft and only seems to recognize Copilot, so it’s very hard to use one own’s funding for paying something else.

qq661mo ago

Paste what you want me to ask 5.5 Pro and I'll paste you the response.

bwfan1231mo ago

> I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models

I am starting to see folks saying - ok, so LLMs can do this, what value have you added ? modulo llm is becoming the norm.

iberator1mo ago

Its good. You should work hard if you are in the public sector! Using claude (cheating) and getting check from government is unethical

pmontra1mo ago· 17 in thread

It's a very long post with a mix of technical (math) and philosophical sections. Here are the most striking points to reflect upon IMHO.

> It seems to me that training beginning PhD students to do research [...] has just got harder, since one obvious way to help somebody get started is to give them a problem that looks as though it might be a relatively gentle one. If LLMs are at the point where they can solve “gentle problems”, then that is no longer an option. The lower bound for contributing to mathematics will now be to prove something that LLMs can’t prove, rather than simply to prove something that nobody has proved up to now and that at least somebody finds interesting.

Training must start from the basics though. Of course everybody's training in math starts with summing small integers, which calculators have been doing without any mistake since a long time.

The point is perhaps confirmed by another comment further down in the post

> by solving hard problems you get an insight into the problem-solving process itself, at least in your area of expertise, in a way that you simply don’t if all you do is read other people’s solutions. One consequence of this is that people who have themselves solved difficult problems are likely to be significantly better at using solving problems with the help of AI, just as very good coders are better at vibe coding than not such good coders

People pay coders to build stuff that they will use to make money and I can happily use an AI to deliver faster and keep being hired. I'm not sure if there is a similar point with math. Again from the post

> suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but the LLM did all the technical work and had the main ideas. Would we regard that as a major achievement of the mathematician? I don’t think we would.

bambax1mo ago

Yes but it's not just that if you solved a problem yourself, you're better at solving other problems; it's also that you actually understand the problem that you solved, much better than if you simply read a proof made by somebody (or something) else.

I see this happening in the enterprise. People delegate work to some LLM; work isn't always bad, sometimes it's even acceptable. But it's not their work, and as a result, the author doesn't know or understand it better than anyone else! They don't own it, they can't explain it. They literally have no value whatsoever; they're a passthrough; they're invisible.

tempaccount50501mo ago

Are you a cutting edge research scientist or something? Everyone I know works in the same domain every day. The problems are the same. People aren't solving brand new problems to humanity every day. We make budgets and look at ticket counts. Roll out patches. Replace hardware. Upgrade software packages. Make a new dashboard to track a project. I guess if every day is a completely novel thing for you, ok. I feel like the goalposts have moved to an absolutely ridiculous place. Oh no, I won't have a bunch of random error log numbers memorized anymore? Who gives a shit. I just want to afford a place to live so I can play my guitar and make something good for dinner. Maybe I'm just old, but I don't see why the average person needs to be a fuckin genius problem solver.

_vertigo1mo ago

I think that’s fine, but 1) that mentality leaves you extremely vulnerable to being disrupted by LLMs and 2) IMO, if you are solving the same problems every day it means you are not making progress on solving the root causes of those problems. What you are describing is toil, not knowledge work

doginasuit1mo ago

I don't think it matters much what kind of problem it is. If it is challenging enough to benefit from assistance and you end up playing a minor role in the solution, it seems like you are putting yourself in the worst position possible. You lose your edge for functioning within the problem space and it raises the question why you are even in the loop at all. If its job security you want, transforming your role into LLM babysitter seems like the worst way to ensure it.

2 more replies

kobe_bryant1mo ago

so how would an LLM being able to do your job help you afford a place to live

slibhb1mo ago

> I see this happening in the enterprise. People delegate work to some LLM; work isn't always bad, sometimes it's even acceptable. But it's not their work, and as a result, the author doesn't know or understand it better than anyone else! They don't own it, they can't explain it. They literally have no value whatsoever; they're a passthrough; they're invisible.

According to the blog post linked in the OP, the LLM-generated results were read, understood, and confirmed by the mathematician whose work they built on.

I notice a dichotomy here between people who care about results and people who care about process. The former group wants to use LLMs insofar as they can contribute to getting results. The latter group is wary of LLMs because they're more interested in the process and less interested in the results themselves. Needless to say, I think the former group is right, and I'm happy to see that mathematicians (or some of them) agree.

1 more reply

joe_mamba1mo ago

> They literally have no value whatsoever; they're a passthrough; they're invisible.

Then middle management also have no value, since they're also a passthrough between upper management and ICs, yet they never went extinct.

1 more reply

panarky1mo ago

I delegate writing a binary executable to the compiler and the linker.

I don't know or understand the binary executable.

I don't own the binary executable, I don't understand it, I can't explain it, it's not my work.

I'm a passthrough; I'm invisible.

I have literally no value whatsoever.

kerabatsos1mo ago

But perhaps we should regard it as a major achievement.

lmpdev1mo ago

I mean in the same way getting Wolfram Alpha to solve a really hard/ugly differential equation I suppose

aspenmartin1mo ago

Insane that we have a system capable of making innovative math proofs and people dismiss it as unimpressive

1 more reply

palata1mo ago

I feel like you slightly miss both points.

> Training must start from the basics though.

Sure, but the point is that at some point (e.g. when starting a PhD) one needs to do research, not learn the basics. And LLMs make that harder, because they solve the "easy research" part.

Take a young lion "fighting/playing" with another young lion as a way to learn how to fight, and later hunt. And suddenly they get TikTok and are not interested in playing anymore. Their first encounter with hunting will be a lot harder, won't it?

> People pay coders to build stuff that they will use to make money and I can happily use an AI to deliver faster and keep being hired.

Again, that's true but missing the point: if you never get to be a "good coder", you will always be a "bad vibe coder". Maybe you can make money out of it, but the point was about becoming good.

sdeframond1mo ago

> Would we regard that as a major achievement of the mathematician? I don’t think we would.

1. Does it matter, really? 2. Is it very different from previous computer-aided proofs, philosophically?

agnishom1mo ago

1. It matters because there are human mathematicians who pride themselves for their mathematical achievements. Mathematics is art to them.

2. Yes, it is. Because pre-LLM era computer-aided proofs were about using the computer to either solve a large number of cases or to check that each step in a proof mechanically follows from the axioms.

1 more reply

layer81mo ago

It matters because most mathematicians thrive on the recognition of their achievements. If what you do any mediocre mathematician could have done, that takes away motivation and fulfillment.

andai1mo ago

> Training must start from the basics though. Of course everybody's training in math starts with summing small integers, which calculators have been doing without any mistake since a long time.

Yeah, it's the same way with learning programming. LLMs can handle basic programming (and increasingly advanced programming) but I think it's necessary to write code by hand. As a beginner, of course, and arguably to maintain skill later too.

The alternative would be like, just asking ChatGPT to do your math homework and then "verifying" it by looking at it and saying "yeah, that looks okay." What are you going to learn?

We do stuff by hand for a reason.

coderenegade1mo ago

You only get good at the things you actually do. Our ancestors had to maintain a minimum level of fitness in order to be able to eat -- a level that most people today never reach, because the modern world has removed that need. Thinking is a skill just like any other, so what happens when people no longer have to exercise that skill to survive? It's a scary thought.

1 more reply

Jweb_Guru1mo ago· 16 in thread

This jives with what I've experienced in the brief time I had access to 5.5 Pro. It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not.

The downside (not noted in the article, but noted by others here) is cost. It uses tokens at an insane rate, the tokens cost a lot, and using it with subagent flows that you can use to have it tackle large problems with high accuracy costs even more. It is also much "slower" for large scale problems because of context limitations -- it has to constantly rediscover context for each part of the problem, and in order to make it accurate you need to wipe its context before progressing to the next small part, or launch even more agents. For mathematical proofs like these, where the required context to understand the problem and proof besides stuff that's already available in its training set is small and the problems are considered "important" enough, this might not be a problem, but for many of the tasks I would like to use it for (ensuring correctness of code that affects large codebases, or validating subtle assumptions) it definitely is one.

So I think it will be a while before the impressive capabilities of these models really percolate into our lives as programmers, unless you're one of the lucky ones given unlimited access to 5.5 Pro.

elAhmo1mo ago

> It's the very first LLM that I feel like I can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided, but it does a pretty good job of tracing its own reasoning and correcting itself in a way that the other models do not.

I swear that people have said the same thing with effectively every new model that came out in the last six months.

fluidcruft1mo ago

I think it's because people walk every model up to its limits and become very aware of a task they can't make work. They do a lot of work simplifying and understanding limitations at that boundary. Then an improved model comes out and they immediately toe that barrier and make swift progress. They will also notice that the new model is natively doing tricks they had done manually.

The reality is likely that everyone is hitting similar barriers and the solutions are somewhat generalizable and get added to training new models.

Eventually people will reach the new limits and the cycle repeats.

fasterik1mo ago

> I swear that people have said the same thing with effectively every new model

That is definitely true, and at the same time, we can measure progress by who is making that claim. When Timothy Gowers, a Fields Medalist, says that models are now capable of "producing a piece of PhD-level research in an hour or so, with no serious mathematical input from me," we can be pretty confident that we are getting into seriously interesting territory.

Jweb_Guru1mo ago

Many people may have, but I certainly haven't.

hackable_sand1mo ago

They did. The scam continues.

1 more reply

y1n01mo ago

> This jives with what I've experienced

Just as an fyi, the word you are looking for is jibes. Jive is something else entirely.

jibe1mo ago

I'm with you!

1 more reply

pfdietz1mo ago

Excuse me stewardess, I speak jive.

jfaat1mo ago

I'm going to start using malapropisms so people know I didn't use an llm to write things

boring-human1mo ago

Cut me some slack, Jack.

billfor1mo ago

Blame The Bee Gees: https://en.wikipedia.org/wiki/Jive_Talkin'

1 more reply

bicepjai1mo ago

Interesting I did not know that I would have used jives :) thanks

refulgentis1mo ago

That ship sailed looooong ago.

1 more reply

idiotsecant1mo ago

The only thing worse than complaining about this is being the guy complaining about the guy complaining about this. So congratulations on being second most annoying.

1 more reply

Forgeties791mo ago

> can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided,

I don’t know about the rest of y’all but I find “rigidly guiding” LLM’s incredibly tedious and frustrating in the same way seeing an error code throw for the 40th time while troubleshooting something on my computer for two hours is frustrating. It also feels somewhat like micromanaging a direct report. I don’t find that process fun or enjoyable in the slightest and it teaches me little in the process. It’s just trading styles of work, and I guess the response to that is “some people prefer that of work.” I just don’t like being told by the world we all have to work that way now I guess.

Jweb_Guru1mo ago

I agree. I find it endlessly frustrating and kind of hate what programming has become. But at least for me it meets the minimum bar of "it works if you push things" now. For past models, under no circumstances could I get them to semi reliably solve these kinds of problems correctly without giving them so many "hints" that they weren't actually saving me time. The kind of reasoning I'm talking about is stuff like "can you actually construct a trace from program start for this condition that looks locally reachable?" Past model simply cannot reliably answer such questions as soon as the control flow involves enough hops or requires tracing through enough function calls.

MrDrDr1mo ago· 15 in thread

> "Even though I can motivate it in retrospect, ChatGPT’s idea to use h^2-dissociated sets to control relations of order at most h feels quite ingenious. As far as I can tell, this idea is completely original."

The question that keep bothering me is can an LLM generate an idea that is truly novel? How would/could that actually happen? But then that leads to the question - what are we actually doing when we think?

Perhaps it's as simple as the ability to just make mistakes that matters, the same things that powers evolution. As long as the LLM can make mistakes, it's capable of generating something genuinely novel. And it can make more mistakes much faster than we can.

charleshn1mo ago

Yes, they can.

Some people like to parrot "next token prediction", "LLMs can only interpolate", and other nonsense, but it is obviously not true for many reasons, in particular since we introduced RL.

Humans do not have the monopoly on generating novel ideas, modern AI models using post training, RL etc can come to them in the same way we do, exploration.

See also verifier's law [0]: "The ease of training AI to solve a task is proportional to how verifiable the task is. All tasks that are possible to solve and easy to verify will be solved by AI."

This applied to chess, go, strategy games, and we can now see it applying to mathematics, algorithmic problems, etc.

It is incredibly humbling to see AI outperform humans at creative cognitive tasks, and realise that the bitter lesson [1] applies so generally, but here we are.

[0] https://www.jasonwei.net/blog/asymmetry-of-verification-and-...

[1] http://www.incompleteideas.net/IncIdeas/BitterLesson.html

vld_chk1mo ago

I genuinely start to think that we, as humanity, severely overestimate our cognitive abilities. We act so surprised “just a few years of LLM with a few RL tweaks match our PhD levels! It must be hidden inside our knowledge base!”. Em, what if no? What if our “PhD level” is just very low level comparing to upper boundaries of measurable intelligence? What if we need to learn being humble and stop treating our minds as “sacred source of creativity and intelligence”?

energy1231mo ago

RL or no RL, AI cannot escape the distribution it's trained on. It's just that the labs will put so much into the distribution that we won't be able to tell the difference that easily, nor will it matter for most tasks. The reason AI does well on ARC-AGI-2 is because the labs created synthetic training data using similar puzzles.

2 more replies

jdub1mo ago

Reinforcement learning for "reasoning" perturbs the model to generate completions in a particular chain of thought / alternative selection structure. It's three next token predictors in a trench coat.

charleshn1mo ago

> Some people like to parrot "next token prediction", "LLMs can only interpolate", and other nonsense

Thank you for illustrating my point.

eterm1mo ago

My own take, and it's veering into the Philosophy of Mathematics, but there's a debate about whether Mathematics is "Invented" or "Discovered".

If it's "invented", then it requires ingenuity.

If it's "discovered", then it was always already there, just waiting for the right connections to be made for it to be uncovered and represented in a way we can understand.

Invention requires ingenuity, but discovery does not. So if LLMs can generate truly novel mathematics, for me that settles it that mathematics is indeed discovered, as LLMs are quite capable of discovery yet I don't consider them possible of invention.

layer81mo ago

Mathematical concepts are invented, but they live in a space of possible (conceivable) mathematical concepts, and we can only invent concepts from that selection of possible concepts. This can be reframed as a process of discovery regarding which conceptions are possible.

Furthermore, the results of theorems aren’t an invention, they are a discovery of what the base assumptions (axioms) logically entail. Finding out which theorems are true and provable is a discovery process. For example, the results of Gödel’s incompleteness theorems were a discovery. They weren’t invented, in the sense that the results couldn’t have been otherwise. We merely could have failed to discover them.

This also holds for physical inventions. You discover a working way to build some functioning mechanism. It’s a process of discovery of what is possible in the physical world.

Whether you portray somethings as a discovery or as an invention is more a matter of degree, a matter of from which angle one is looking at it.

The possible states of an LLM are finitely enumerable. The same likely holds for the possible states and configurations of a human brain, in approximation. Therefore there is only a finite set of possible ideas, thoughts, and conceptualizations an LLM or a human can have, and in principle they could be exhaustively enumerated and thus “discovered”.

MrDrDr1mo ago

I like this distinction, but it would then seem the only 'invention' would be the axioms of your mathematics. There exists numbers (natural, imaginary...), there exist shapes (a point, a line...). All the work from that point on could be 'discovered'. I agree that I don't see LLMs inventing in this way. But again, it raised the question - what are our brains doing when we 'invent' something?

1 more reply

eiieue1mo ago

Mathematical objects are an invention of the mind - they are abstract objects that only an entity who can process abstractions can make sense of.

There is no ‘discovery’ here nor was it waiting to be found. The human has to sacrifice and pursue the path of exploring reality and thereby is inherently inventing.

Humans built up mathematics iteratively from smaller bases extending into large ones. Is this what LLM’s do? Of course not - They are fed with vast amounts of information from the off.

LiamPowell1mo ago

Trivially the answer is yes by the infinite monkey theorem. If we allow the sampler to pick any token then any stream of arbitrary tokens can be generated. Therefore if an original idea can be represented with written words then a LLM can generate it. That is perhaps not the most satisfying answer, but if you want a better one you'll need to provide a function that determines if an idea is original.

humanfromearth91mo ago

For my paper about ME/CFS, I let an LLM integrate lots of findings of other scientific papers. Then I ask the LLM to "creatively brainstorm", given all we know of ME/CFS and the newly integrated paper, to generate new hypotheses, treatment ideas or any other kind of insight it can think of.

This works really well.

Now, it's clear that I have no idea how much of this is something we would consider new and original, and how much is a kind of systematic, but not novel, easy of thinking.

What I couldn't do so far is get an LLM to generate a truly new maths theory, with new abstract concepts and dimensions and points of view. The kind that is not just a combination of existing theories and logic.

realtalk_sp1mo ago

Would you mind posting the outcome of this? A person I love dearly is struggling with Long Covid/CFS. I’ve been doing something similar to what you describe, but I’m always looking for more angles that could help.

jasfi1mo ago

It's about the ability to combine ideas in novel ways, without breaking the rules in relevant frameworks. Sometimes the idea may even be to contradict existing theories where they are weak.

ikari_pl1mo ago

How do you define a new idea?

To me, it's rearranging the information you had in a way that hasn't been applied or published before.

That's literally what LLMs are built for.

eiieue1mo ago

Theres a simple test for this.

Limit the knowledge an llm to some point in time at which a discovery was made. And check to see if the llm could produce the discovery.

If you think OAI hasn’t already tried this then think again - they have every incentive to do so and announce it to the world.

MinimalAction1mo ago· 13 in thread

As a graduate student, this piece made me sad. I always believed that my work speaks for itself and transcends beyond my limited time on this cosmic experience. This notion of immortality was just a small intangible bonus I hoped for when I jumped into grad school. AI is making me feel less worthy.

hodgehog111mo ago

As someone who is much further down the track, I would kindly suggest you drop that line of thought. I've seen far too many brilliant and ambitious people drop into depression because of it.

You are worthy of doing this work because you are able to do it. Do the work because you love it and because you love the mystery. Enjoy every moment that you get to do it. Find joy in the great fortune you have to do this work while others toil away on tasks that bring them no satisfaction. Sometimes it's tedious, but sometimes it's incredibly rewarding in its own right.

Don't work for the possibility of eternal glory though, it just doesn't exist anymore.

MinimalAction1mo ago

Thank you for this comment. I often fall into the why of graduate school many times. The pay is insufficient, hours are long, but at least I find it very satisfying on good days. It is just the feeling that what I do may not be unique anymore is what sucks. I didn't necessarily mean to find glory through incredible work alone, but through being unique in the problems I choose. Anyway, I digress.

1 more reply

whatever1201mo ago

You are worthy. You will hone your skills in grad school and be able to command these AIs better than somebody who hasn’t struggled with hard problems for a long time.

jlarcombe1mo ago

A depressing thought that all that work is just so you can "command AIs better"

folderquestion1mo ago

It could happen than the AI, in a near future, is not something external but just a part of your brain, so you retain the glory.

2 more replies

alexashka1mo ago

All that work to kick a ball into a net.

Nobody looks at this species and goes hm, rational and reasonable :)

helloplanets1mo ago

"If you value intelligence above all other human qualities, you're gonna have a bad time." - Ilya Sutskever, 2023

timedude1mo ago

Let me tell you, there is a ton more to learn in this reality than llms are capable of finding out on their own, especially when it comes to truth, ethics and morality. And those are the only thing that matter in the end when you leave this reality. A greater challenge does not exist.

ionwake1mo ago

I feel bravery transcends time better than the odd scientific breakthrough which are often attributed to one, but whose roots came from a "lesser" unknown

kranke1551mo ago

try meditation.

MinimalAction1mo ago

Thanks. This might help. Are you suggesting any particular form?

2 more replies

alexashka1mo ago

> I always believed that my work speaks for itself and transcends beyond my limited time on this cosmic experience

Any statement preceded by the word 'believe' is a coping mechanism.

> This notion of immortality was just a small intangible bonus I hoped for when I jumped into grad school

Any statement preceded by the word 'hope' is a coping mechanism.

> AI is making me feel less worthy

Worth comes from understanding, not achievement.

MinimalAction1mo ago

I strongly disagree on beliefs and hopes are coping mechanisms. Coping from what? Beliefs and hopes are what they are.

But I agree worth should be derived from understanding, not through achievement.

1 more reply

mxwsn1mo ago· 8 in thread

> Here’s a thought experiment: suppose that a mathematician solved a major problem by having a long exchange with an LLM in which the mathematician played a useful guiding role but the LLM did all the technical work and had the main ideas. Would we regard that as a major achievement of the mathematician? I don’t think we would.

This is a cultural choice. It makes sense that in the mathematics culture we currently have, this is alien. But already, other fields, and many individuals, would disagree and say that the human did have a major achievement here. As long as human-AI collaborations are producing the best results, there is meaningful contribution by the humans, and people that are deeper experts and skilled LLM whisperers should be able to make outsized contributions. The real shoe drops when pure AI beats humans and human-AI collaboration.

pmontra1mo ago

I replied to a comment about AI in sports and I build on that.

We praise car drivers despite most of the performance in their sport comes from the car. The driver makes the difference when two cars are close in performance. Brilliances or mistakes. Horse riders too.

In the case of math, the human can lead the LLM on the right track, point it to a problem or to another one. So it deserves some praise.

Then the team that built the car, cared about the horse, built the AI might deserve even more praise but we tend to care more about the single most visible human.

butlike1mo ago

Well, kind of. In F1 there's 2 championships: the driver's championship, and the constructor's championship. The constructor's championship is for the best engineered auto. Both the driver and the engineers win separate things because it takes both.

dmbche1mo ago

Could you win an F1 race with the latest winning car against F1 drivers?

3 more replies

djeastm1mo ago

>Would we regard that as a major achievement of the mathematician? I don’t think we would

For some reason this reminds me of AI images and a domain like comedy.

If an image makes people laugh, the person who prompted it to make the image certainly doesn't get credit for the vast majority of the work in its creation, but perhaps they do get credit for the initial prompt idea and then the "taste" to select that particular one from whatever drafts they went through or otherwise guiding it.

So if a mathematician comes up with an amazing result that an LLM "did", I think they could still get a bit of credit for prompting it to do it and being its guide.

But whereas the first person could perhaps be called a comedian and not an artist, would the mathematician still be called a mathematician or something else?

gobdovan1mo ago

I would. Even if someone found a prompt or even automated the conversation and just searched all open math problems I still would. If they produced a useful result without harm to anyone, that's a valuable human activity that should be rewarded just as well as we reward the other mathematicians, which I imagine is quite a lot, given all the billionaire mathematicians...

5424581mo ago

> given all the billionaire mathematicians

We just call those ones “quant traders”.

ComplexSystems1mo ago

Absolutely. Who cares if the LLM automates some of the grunt work? Mathematicians are artists, and they paint with ideas. The goal is map out more of the beautiful structure of how things work. The enjoyment in it derives entirely from the payoff of seeing the larger view of how things fit together. If part of their process involves bouncing things off of other people, or even LLMs, I don't think it matters much, nor does it take away from the enjoyment in getting things figured out.

bambax1mo ago

It may not be a major achievement by the mathematician (although it's debatable) but it would still be a major result.

few1mo ago· 8 in thread

>So if your aim in doing mathematics is to achieve some kind of immortality, so to speak, then you should understand that that won’t necessarily be possible for much longer — not just for you, but for anybody.

This made me a little sad

mentos1mo ago

I watched the movie '21' (2008) for free on YouTube yesterday.

The opening of the movie features the MIT campus full of students navigating its grounds and all the promise and status that higher education brings. [0]

Gave me the same sense of sadness realizing how much will fall to AI.

[0] - https://youtu.be/0lsUsWdkk0Y?si=TJl7f_b1RcWcDqF8&t=278

lexoj1mo ago

Not free in my country, didn’t know YouTube was broadcasting full movies in certain regions as you imply.

1 more reply

vessenes1mo ago

This was the most interesting line in the essay to me as well — I flashed back to quitting an academic math career instantly; the way I thought about it at 19 or 20 was that I didn’t think I could be world class at it. (Rightly). The next thought I had was “what am I good at?” And implied in that was at the very least “What could I be world class at?” Or at least very good at.

I don’t think I ever thought I was good enough to try and get (math) immortality by finding and naming some result that would live beyond me, but if I had, perhaps this bad news would have had a similar impact on me.

That said, I think I disagree with the premise at the margin, at least. I don’t care how many proof assistants or cluster compute is used - the team or person that proves the Riemann Hypothesis will be famous, or at least math famous.

jdale271mo ago

I don't know that it's that disappointing. I doubt most of the great mathematicians were actually doing it to achieve immortality. I suspect most of them were either after (possibly indirect) practical applications (via the math -> physics -> engineering pipeline) or just "for the love of the game", appreciation of the beauty of math and the intellectual joy of doing it. AI might also take over the practical application side, but the other aspects are still there for the taking.

hodgehog111mo ago

Exactly. Gowers is in the unique position to think about the "glory" of frontier mathematics, but for essentially everybody (especially those working outside of number theory), that dream died long ago. There are far too many mathematicians now.

Many mathematicians work because they love the breakthrough (a certain quote of Villani comes to mind). They love finding new results, uncovering new mysteries. From that point of view, having an AI that can build on your basic ideas and refine them into more powerful arguments is awesome, regardless of who gets the credit. There are those that treat it more like solving puzzles so the result is not of interest. From that point of view, I can see the dissatisfaction. But I have found those with that viewpoint don't tend to make it as far in academia as those with the other viewpoint.

bananaflag1mo ago

Now repeat that for every sort of human achievement

bel81mo ago

Machines are comming even after table tennis :(

https://www.youtube.com/watch?v=VVEzgYxDdrc

pmontra1mo ago

Sports are safe. Machines came after runners (motogp, formula 1) and yet we cheer the winners of the 100 m at the Olympics Games. Fully autonomous bikes and cars won't change that. AIs destroy chess players. We still cheer the world champion.

We care about sports with humans.

1 more reply

robot-wrangler1mo ago· 7 in thread

A very interesting comment from Baez, I'll just quote part of it.

> Where does the value of thinking and having deep ideas come from? We need to think about this now. If it comes primarily from their scarcity – the fact that having certain ideas is hard – then indeed this value may drop precipitously when the manufacture of ideas can be automated. But if the value comes from the utility of the ideas – the benefit that the idea brings – then the story changes: perhaps creating more good ideas is actually better, not worse. Here I’m using “utility” in a broad sense, not just in the sense of what people often call applied mathematics.

> In other words, mathematicians may need to adjust to a transformation from a scarcity economy to an abundance economy.

https://gowers.wordpress.com/2026/05/08/a-recent-experience-...

zarzavat1mo ago

There are three species of mathematicians:

The first species is the pure problem solver. Tao is the poster child for this group. Their currency is interesting problems and solutions to those problems.

The second species is the pure theory builder. The poster child for this group is Conway. Their currency is theories and ideas rather than theorems, they are most interested in expanding the territory of mathematics and discovering new mathematical lands.

The third species is the applied mathematician. They see mathematics as a means to an end, they have some problem outside of mathematics and they want to use mathematics to solve it.

It seems like the first group (the problem solvers) are the most immediately threatened by AI, although so far AI is better at solving problems than finding new conjectures.

The second group (the theory builders) are more distantly threatened by AI, since thus far AI has shown limited ability to come up with novel and interesting mathematical ideas and nobody has any clue how to train an AI to do such a thing.

The third group stands to gain the most from AI. If an AI can answer your mathematical question then you can spend less time doing mathematics and more time on whatever it is outside of mathematics that you wanted to use mathematics to help solve.

robot-wrangler1mo ago

I'm cautiously optimistic that the theory camp will benefit as much the applied guys, but later. They dream and scheme and shape fields by making vague intuitions and questions sharper, but they also need interesting patterns to work with as their raw material.

Identifying suitable problems in this sense rather than solutions is an AI use-case you don't hear about much. We don't quite have the infrastructure for this yet but by combining language models / anomaly detection / knowledge-bases we might be in a position to give a Conway 25 interesting high-quality puzzles before breakfast. Funny that it's like a kids nightmare, chatgpt giving them homework instead of solving it, but if it had good taste, people would love it for research.

Anyway, for now, dreamers will probably find more inspiration by cross-pollination with colleagues from different disciplines, or just going for a walk.

lhd11mo ago

I was expecting Grothendieck. Conway is hardly the poster child for theory building.

1 more reply

pocksuppet1mo ago

This is generally true across the economy. It is known as the diamond-water paradox. Diamonds are (for everyday life) useless, so then why are they so much more valued than water, which you need to live?

https://en.wikipedia.org/wiki/Paradox_of_value

qwrahg1mo ago

I note that it is always the same online pundits (even if they are distinguished academics) who push anything new.

Meanwhile Wiles and Perelman stayed offline and solved real problems.

robot-wrangler1mo ago

I don't necessarily think engaging in the personalities is interesting, but I'm struggling to see what is the beef here. Is it personalities? Pure vs applied math? Or AI?

DoctorOetker1mo ago

would Wiles be willing to transcribe his proof for the metamath verifier? it can be done offline indeed...

1 more reply

einrealist1mo ago· 7 in thread

"After 16 minutes and 41 seconds, it came back" ... "further 47 minutes and 39 seconds" ... "After 13 minutes and 33 seconds" ... "After 9 minutes and 12 seconds" ... "After 31 minutes and 40 seconds" ... plus other computations

Anyone spotting the issue here? What did that really cost?

I am not against compute being used for scientific or other important problems. We did that before LLMs. However, the major LLM gatekeepers want to make all industries and companies dependent on their models. And, at some point, they need to charge them the actual, unsubsidized costs for the compute. In the meantime, companies restructure in the hopes that the compute costs remain cheap.

sidkshatriya1mo ago

> "After 16 minutes and 41 seconds, it came back" ... "further 47 minutes and 39 seconds" ... "After 13 minutes and 33 seconds" ... "After 9 minutes and 12 seconds" ... "After 31 minutes and 40 seconds" ... plus other computations Anyone spotting the issue here? What did that really cost?

Whatever the Joules... (convert to $ using your preferred benchmark price) it is a fraction to what it might take a human Ph. D. weeks to feed and sustain themselves when working on the same problem. The economics on LLMs is just unbeatable (sadly) when compared to us humans.

einrealist1mo ago

Compute in science was already subsidized by public funding or by donations. Most supercomputers are financed this way. And that's a good thing. If you have a good science problem that can be computed, apply for compute time. There is nothing wrong to apply that to LLMs as well, like I wrote in my initial post. The human is still required to identity problems that are worth to be computed, to create prompts that the LLM can act on, and to verify results. But, OpenAI providing compute for basically free is still tied to a different incentive: to fuel the hype and to capture the market, while distorting/obfuscating the real costs. That's also the reason for why we cannot claim that 'economics on LLMs is just unbeatable'. It depends on the problem, the reason for a prompt.

intexpress1mo ago

Not necessarily. Humans brains use a tiny amount of power. Most of the human cost would be due to the very high cost of housing in many locations.

1 more reply

hackable_sand1mo ago

Are you guys being intentionally obtuse?

Do you understand the difference between an apple and an apple tree?

colordrops1mo ago

Still not as bad for the environment as animal agriculture, and animal agriculture is absolutely not necessary and only causes harm and suffering for taste pleasure. At least with LLMs we get many positive advancements from them. I don't see these sorts of comments every time someone posts a burger review.

einrealist1mo ago

Did I praise our animal agriculture anywhere?

colordrops1mo ago

We have a wide audience.

kang1mo ago· 4 in thread

> The lower bound for contributing to mathematics will now be to prove something that LLMs can’t prove, rather than simply to prove something that nobody has proved up to now and that at least somebody finds interesting.

5.5pro is amazing but this implication might not be true & is the core argument of this piece.

AI will prove all sort of things - interesting, boring & incorrect.

To sort it will be the task of the PhD.

layer81mo ago

The task of a proof verifier is much simpler than the task of a proof finder (it’s basically equivalent to P vs. NP), and hence the bar for the required skills is lower. Merely verifying proofs isn’t research, and doesn’t impart research skills.

kang1mo ago

Verification on its own is not research, but judgement is research.

"Hey, Prove something a machine can't", sure I can't, "Hey, Say something worth proving & judge it well", ah, now I might have a few unique observation/ideas/curiosities/problems from my having being a human.

Imo, the feeling of intelligence or the process of originality(originativity) test for ai is subjective & is coming down to 4 paths: novel relative to a reference class, valuable within a domain, counterfactually sensitive to internal state and environment, and revisable through learning.

ianm2181mo ago

Verification is generally a much lower bar than solution generation. I don’t think it’s likely sorting out the right from wrong will end up being this huge PhD level effort.

kang1mo ago

Verification & solution generation are both part of problem generation & defining the passing test - judgement.

YeGoblynQueenne1mo ago· 4 in thread

All this sounds to me like mathematicians spooking themselves with stories of how ChatGPT solved a problem, when it's mathematicians solving a problem using ChatGPT as a tool. E.g. from the twitter thread by Timothy Gowers:

>> All I did was say things like, "Yes, it would be great if you could explore that idea and see whether you can get it to work," or "Could you rewrite that argument as a LaTeX file in the style of a standard mathematical preprint?"

Yeah, so all he did was take the horse to the water and make the horse drink. The collaboration with the other two mathematicians wasn't a trivial part of the problem solving either: every time Timothy Gowers figured ChatGPT had goe somewhere with its problem-solving, he stopped, asked it to render the answer in LaTex, and sent the answer off to be verified by the other two.

The reason for that is not to be underestimated: ChatGPT can produce answers to questions you ask it for as long as you ask it to do so but it has no capability to determine whether an answer is correct or not. That's why it needs a human with domain expertise to evaluate those answers. And of course to discard wrong answers in the process, because of course the process that's described here glosses over many false starts and back-and-forths and "you're absolutely rights, here's a new version of that"'s etc. that are common experience when using LLMs for problem-solving tasks.

The existential questions that the article poses about mathematics then are easily answered by taking all of the above into account. If LLMs are a useful tool for mathematicians, then nothing changes. Mathematicians of all levels can still do their job and perhaps do it faster or better with the new tool.

If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening.

whimsicalism1mo ago

Sorry, just so I fully understand your comment - your claim is that asking it to “explore that idea further” and “write the paper in latex” constitutes “taking the horse to the water and making the horse drink”?

thank you for the morning laugh

YeGoblynQueenne1mo ago

https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

famouswaffles1mo ago

>If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening.

I mean that has happened so yeah ?

https://www.scientificamerican.com/article/amateur-armed-wit...

Actual GPT transcript. Zero such input https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...

And maybe the other guy wasn't the most polite about it but his point is very valid. Replace chatgpt with a human in both of these stories and nobody would say that timothy 'took the horse and made it drink'. The 'Horse' would be the first and likely only Author so this just sounds like denial.

That there are multiple of these stories in the last few months by the latest set of models (there are even more than these 2) should provoke this sort of consideration and discussion.

YeGoblynQueenne1mo ago

These are different cases, yes? The person in the SA article you link is described as an "amateur", but Timothy Gowers is not an amateur and he is much more capable of guiding an LLM with domain expertise than an amateur.

Then there's the kind of problem we're talking about. The "amateur" in the SA article solved one of Erdős problems and Gowers himself seems to think that, on its own, is not a cause for concern. He distinguishes his own result from that kind of earlier result at the start of his article:

>> The background is that, as has been widely reported, LLMs are now capable of solving research-level problems, and have managed to solve several of the Erdős problems listed on Thomas Bloom’s wonderful website. Initially it was possible to laugh this off: many of the “solutions” consisted in the LLM noticing that the problem had an answer sitting there in the literature already, or could be very easily deduced from known results.

So we have an "amateur" who "vibe-solved" an Erdős problem, on one hand, which may or may not already had a solutiuon lurking in the wings on the one hand; and an expert who solved a harder problem by interactive use rather than vibe-solving, on the other hand. There's no reason to believe that we can "Replace chatgpt with a human in both of these stories" as you say.

And btw there's scholarship that indicates vibe-solving is not yet ready to replace mathematicians like Timothy Gowers:

First Proof

To assess the ability of current AI systems to correctly answer research-level mathematics questions, we share a set of ten math questions which have arisen naturally in the research process of the authors. The questions had not been shared publicly until now; the answers are known to the authors of the questions but will remain encrypted for a short time.

https://arxiv.org/abs/2602.05192

See Appendix A for initial results.

1 more reply

globular-toast1mo ago· 4 in thread

I wish people would stop generating stuff they don't understand only to forward it to someone who does. Something about that really rubs me the wrong way.

hodgehog111mo ago

May I remind you that this is Timothy Gowers. He says he doesn't understand, but he most certainly has far greater capacity than most to detect complete junk from a maybe plausible argument. His colleague is even better able to judge this, hence why he sent it to him.

Also if he did send me complete junk, I would still parse it for multiple days to see what is there.

globular-toast1mo ago

Yeah, it doesn't make a difference for me. It's the generation part. Gowers should have sent his prompts to the colleagues, not the generated paper. That's all. I feel like it's creating obligations for others to help with the remaining 20% which always takes the most time, while you get to have all the fun of doing the first 80%.

I'm not criticising Gowers directly in this instance because he's exploring the possibilities, my disdain is towards the more general pattern I see emerging where people just send each other LLM outputs.

1 more reply

auggierose1mo ago

Lol. If Gowers sends you a piece of math he doesn't quite understand because he thinks that you might, that is something you celebrate.

frozenseven1mo ago

You are criticizing a Fields Medalist for consulting with another mathematician.

CharlesLau1mo ago· 4 in thread

Is the assessment system of undergraduate mathematics education no longer effective?

margalabargala1mo ago

Undergraduate? No. We've had calculators able to solve undergraduate problems for decades. AI doesn't change the need to understand how calculus works any more than calculators did. The foundations remain valuable.

Graduate? Yes.

whatever1201mo ago

How should graduate school be changed then? Specifically for mathematics

2 more replies

dyauspitr1mo ago

I don’t think it’s just mathematics. We don’t hear enough about this, but if I think back to my undergraduate years, which were less than 10 years ago, every homework assignment and every take-home exam I had would be trivial for LLMs to solve at this point I wonder what is actually happening on the ground.

crocdundae1mo ago

Well... here's something from "boots on the ground": I teach a bachelor's degree where programming is a smallish facet of a curriculum. My course is the last of a series of 3 courses which progressively introduce more concepts and try make practical implementations more feasible. I've been able to grade the course purely based on returns to take-home exercises, some of which are complex, some trivial. When ChatGPT (& Co.) came along I was still able to do that but with a major added workload to me (suddenly everyone started producing mountains of code, often nonsensical, but I still had to read it all). I always requested targeted, atomic changes to code (vs. rewrites) which served me well up to a point (I was still able to grade fairly). I requested them originally to avoid "github copies", but that worked kind of OK with ChatGPT too. However, when ClaudeCode came along it was obvious to me I'm loosing the battle. It does not particularly matter to me whether students use AI or not as long as the rows they add and alter in the assignments make sense, but the "last nail to the coffin" problem now with ClaudeCode is that in the latest batch (this spring) it is clear some students "pay themselves" a good grade (i.e. they pay for ClaudeCode, thus bypassing the need to actually learn). I cannot make assignments that are both complex enough to cause ClaudeCode tripping on something and still humane for those who do not use AI or only use free chatbot options. Essentially ClaudeCode plays havoc with the whole grading process: students not using it (whether they try to write code fully manually or ChatGPT assisted) are left with far less points that students who just push all the code I give to ClaudeCode and "let it rip" for some 15 minutes. This really irks me. So, my solution? Still working on it and hoping to find one! For sure no more points from most take-home assignments: lowest grades still achievable through them (the trivial ones), but that's it, the rest it preparation for an exam. Practically this already means anyone with ChatGPT is going to pass, no doubt about it... As for the higher grades, for autumn I'm desperately now figuring out how to even make a meaningful paper based exam for my course. I've myself completed a master's degree writing C language on paper with a pencil. I sure did not want to start doing that to others, but here we are. Besides, back in my youth the only "library" was pretty much ANSI-parts-of-C! I'm not sure what kind of a 2 inch thick stack of papers I'd have to give my students into the exam these days as reference material. One horrible aspect is that students are now far more dependent on compiler errors to spot pretty much anything and everything... I worry the first paper exam from me will be a total horror story to us all. In any case, interesting times.

1 more reply

TrackerFF1mo ago· 3 in thread

The vast, vast majority of students going into higher education this fall will not contribute much to science until 4-5 years down the road (should they do research). Realistically 6-7 when they're in full swing with their Ph.D.

If we look where these models were 5-7 years ago...the existential threat of the Ph.D. was not even on the radar back then. The people finishing up their doctorate now are the first that can truly leverage these tools.

Now, if these to-be researcher students feel defeated (enough to quit), or completely lean on AI models the work for them, we're going to have a problem. Same with the funding of those Ph.D. positions. If we move away from "funding to produce researchers" to "funding to achieve results", will money that was usually spent to fund Ph.D. students start to flow towards compute?

If we look at it a bit cynically: Some researcher will be able to pump out a lot more papers by spending money on compute, than a couple of years of training students.

Interesting times. But also so much uncertainty. I feel terrible for the students that will have to decide now what they want to do, with all this knowledge.

robot-wrangler1mo ago

> Now, if these to-be researcher students feel defeated (enough to quit), or completely lean on AI models the work for them, we're going to have a problem. [..] If we look at it a bit cynically: Some researcher will be able to pump out a lot more papers by spending money on compute, than a couple of years of training students.

Obviously this is already happening and will accelerate. Outside of grad work, you could already just buy a degree. Certainly in the softer disciplines, you can currently just buy a phd thesis and a good publication history. If you're in industry instead of academics, you can even buy a promotion. If your employer gives an AI budget to all workers then you quietly double that budget out of your own pocket for as long as it takes to get a promotion, then stop and just enjoy a bigger paycheck.

ndkap1mo ago

PhD students are already using AI models to work for them. Most of the PhD candidates I know have $200 Claude Max plan which they use to their fullest.

I see that they are able to do researches that they were not previously able to do. And although I see that using AI has certainly diminished their ability to code some stuff up, I see it the same way as someone using scikit-learn or Pytorch to code their ML models -- indeed the underlying details is abstracted away from you, and without AI, you won't be able to do much, but the research that you do is indeed happening because of you and wouldn't have happened with just the AI doing the research.

odyssey71mo ago

It’s not as if institutions have been lavishing PhD students with money up until now.

As an afterthought budget item, those funds aren’t exactly attractive targets to raid for pursuing an expensive, different process.

robot-wrangler1mo ago· 3 in thread

Despite this coming from an independent expert and not from OpenAI, we need to be honest that this is more like a marketing campaign than open science. I assume the progress really is valid, the experts are indeed impressed, and we have the accurate time that it takes to produce the results. What we don't have is detail about true cost, or a CoT trace, or anything like that.

The implication is: we're ready to let everyone go wild with this very soon. Ok, go wild with what exactly? How do we know that influential VIP users who might make very friendly blog posts aren't getting allocated exclusive access to a billion dollars worth of hardware when they ask questions? I mean literally giving a certain group of people temporary privileged access to like 90% of all available compute would be a completely reasonable business decision for OpenAI.

Would a reveal like that change how we think about the result? What if half that amount of cash/compute could enable some completely non-AI approach of numerical brute forcing that settles the question even if it didn't write the paper?

My other question is always whether the latest is purely using giant models or if we're now deeply into harnesses that use MCTS and such. Understandable to keep that a trade secret I guess. But IMHO we should at least get the CoT trace as a proxy for true cost, or else maybe we're just getting played to do the hype for corporate.

electriclove1mo ago

Marketing campaign???

energy1231mo ago

We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro, not by familiar names that OpenAI is sneaking additional compute to behind the scenes (which is too conspiratorial of an explanation for my liking regardless).

robot-wrangler1mo ago

> We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro

Where's that? The stuff I've seen is from celebrities. Were those problems as hard as this one, or the ones that Tao posts about? Regardless.. what's the argument against more transparency here to just settle this kind of thing?

> which is too conspiratorial of an explanation for my liking regardless

OpenAI is not, in fact, open. Why do they deserve the benefit of the doubt?

Regardless.. special treatment for special customers isn't conspiracy, it's SOP literally everywhere and especially if you're helping to beta test. Anyone who's ever interacted with any technical account manager has seen waived quotas, free resource allocations, etc. The quid-pro-quo is obviously that your cheap early access means you get to give talks at a conference (or make a blog post that a lot of people read and talk about).

1 more reply

quinndupont1mo ago· 3 in thread

[flagged]

vessenes1mo ago

I don’t love the tone here, but I do think you get at a key question in mathematical philosophy.

Mathematicians have engaged, vigorously, on this very philosophical question for centuries - is math discovered truth, or is it more akin to building an edifice where you first define the materials, then the structure, and see where it leads?

There are lots of strong feelings on both sides. For instance: “God created the integers, the rest is the creation of man” — Kronecker, 19th century sums up one particular perspective.

To me, it’s probably a mix of both - some fantastic results in imaginary numbers show up as describing key electromagnetic effects many decades after they were first ‘discovered’ by theoretical mathematicians.

NB: My original comment led with a pejorative, which was rightly flagged.

dang1mo ago

> Who hurt you bro?

Please don't respond to a bad comment by breaking the site guidelines yourself. That only makes things worse.

https://news.ycombinator.com/newsguidelines.html

1 more reply

dang1mo ago

Can you please make your substantive points without fulminating, as the site guidelines request (https://news.ycombinator.com/newsguidelines.html)?

We're trying for curious conversation here, and you've clearly got something interesting to say, but when you put it this aggressively, curiosity gets fried (https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...)

zkmon1mo ago· 2 in thread

>> but it was definitely a non-trivial extension of those ideas, and for a PhD student to find that extension it would be necessary to invest quite a bit of time digesting Isaac’s paper

The "non-trivial" is for human abilities. The weights lifted by a crane are also "non-trivial". People keep getting amazed at machine's abilities. Just like a radio telescope can see things humans can't, microscope can see the detail humans can't, we need not be amazed. The sensory perception of patterns is at different level for AI. It's a machine.

svnt1mo ago

Too many people are wrapped around the ego axle thinking (assuming) their ideas are both them and somehow unique and special.

It usually takes dissolving that, often through difficult experiences, before they can see it as a machine, something that could be separated from them.

dag1001mo ago

I think the more pressing issue is that there isn't really much space left for humans in the economy if thinking can also be automated.

adammdaw1mo ago· 2 in thread

This is certainly interesting, though I would say that based on my understanding of how the current models work combinatorial problems would be an area where they could be particularly successful. They are pretty good at combinatorial creativity - its the exploratory and transformational aspects that are still pretty tricky, and I expect would come to bear in other areas of mathematics.

nxobject1mo ago

I wonder as well whether large-but-finite contexts can handle algebraic questions that require traversing up and down levels of abstraction, at least not without "thrashing".

hodgehog111mo ago

Indeed, analysis is a bit more loose in its arguments, and so I've found LLMs tend to make more mistakes there.

fulafel1mo ago· 2 in thread

Link to source blog post: https://gowers.wordpress.com/2026/05/08/a-recent-experience-...

dang1mo ago

That's the top link (i.e. that the title is linked to), no?

fulafel1mo ago

Indeed, the body in the post made me think it was a url-less submission.

1 more reply

dares25731mo ago· 2 in thread

I think the biggest advantage of ChatGPT compared to Claude is that there are fewer things outside the model itself, such as KYC, account bans, etc.

solenoid09371mo ago

This is just grossly misinformed.

OAI and Anthropic both require KYC for models of similar intelligence. They both do account bans if the classifiers fire wrong. You simply hear about it less with OAI because Codex has fewer prosumers.

tmp104232884421mo ago

Can you name any instance of OpenAI being as trigger-happy with bans as Anthropic has been in the past few months? Codex may have fewer prosumers, but they've added a lot in that time.

1 more reply

SubiculumCode1mo ago· 2 in thread

I honestly can't say this isn't AGI anymore. AGI shouldn't be a bar so taboo that it has to be at the extreme capability in every domain. What human is?

This is as AGI as it needs to be to get my vote. And it's scary.

MrScruff1mo ago

It's ASI with jagged intelligence, which is probably what it will remain for a while.

It still sounds to me like remarkable automation rather than something that's expanding the frontier of human knowledge, for now at least.

agiipullor1mo ago

to quote Demis Hassabis, "these models can solve frontieer problems in math, but also fail in really dumb ways at trivial questions - the car wash question".

jagged AGI

bambax1mo ago· 2 in thread

> quite a lot of perfectly good human mathematics consists in putting together existing knowledge and proof techniques

Creativity is connecting ideas from different domains and see if something from one field applies to another. I do think AI is overhyped generally; but a major benefit from AI could be that after ingesting all the existing human knowledge (something no single human can ever hope to achieve) it would "mix and connect" it and come up with novel insights.

Most published research sits ignored and unread; AI can uncover and use everything.

imiric1mo ago

> Creativity is connecting ideas from different domains and see if something from one field applies to another.

That's true. The question is whether the produced pattern has any value. LLMs are incapable of determining this, and will still often hallucinate, and make random baseless claims that can convince anyone except human domain experts. And that's still a difficult challenge: a domain expert is still needed to verify the output, which in some fields is very labor intensive, especially if the subject is at the edge of human knowledge.

The second related issue is the lack of reproducibility. The same LLM given the same prompt and context can produce different results. This probability increases with more input and output tokens, and with more obscure subjects.

The tools are certainly improving, but these two issues are still a major hurdle that don't get nearly as much attention as "agents", "skills", and whatever adjacent trend influencers are pushing today.

And can we please stop calling pattern matching and generation "intelligence"? This farce has gone on long enough.

agiipullor1mo ago

> And can we please stop calling pattern matching and generation "intelligence"

thats literally what an IQ test tests - abstract pattern matching. but I guess you dont like IQ tests either

1 more reply

slopinthebag1mo ago· 2 in thread

AI generated article btw.

Maybe if you find AI to be doing stuff you find impressive, the stuff you were doing wasn't that impressive? Worth ruminating on your priors at least.

hodgehog111mo ago

This is beyond ridiculous to say considering whose blog this is.

For those that don't know, this is Timothy Gowers. He is one of the most accomplished mathematicians in the world. Like Terence Tao, he is considered one of the world leaders in mathematics and tends to have good judgement in where the field is going.

Even without that knowledge, no, this article is certainly not AI generated. It has none of the tells.

reasonableklout1mo ago

What makes you think either the tweet or blog post are AI generated?

bustermellotron1mo ago· 1 in thread

I saw Tim Gowers give a talk at the AMS-MAA joint meeting in Seattle about ten years ago where he predicted that in 100 years humans would no longer be doing research mathematics. I wonder if he’s adjusted his timeline.

At the time I thought the key missing tool was a natural language search that acted like mathoverflow, where you could explain your problem or ideas as you understood them and get references to relevant literature (possibly outside your experience or vocabulary).

34qJhah1mo ago

And Teichmüller thought that Germany would win WW2 and volunteered for the Eastern Front.

Being a gifted mathematician does not make you right. In fact, mathematicians have a lot of bizarre theories.

iTokio1mo ago· 1 in thread

On complex problems with lengthy proofs, the first step that I would have done is to ask 5.5 pro in a new, unrelated, session, to be very critical, to try to find flaws in the arguments.

And certainly not to send it to a fellow colleague to ask its opinion first.

LLMs are certainly becoming capable to code, find vulnerabilities, solve mathematical problems, but we need to avoid putting their works in production, or in front of other humans, without assessing it by any possible mean.

Otherwise tech leads, maintainers, experts get overwhelmed and this is how the « AI slop » fatigue begins.

To be clear I’m talking about this step:

> That preprint would have been hard for me to read, as that would have meant carefully reading Rajagopal’s paper first, but I sent it to Nathanson, who forwarded it to Rajagopal, who said he thought it looked correct.

NitpickLawyer1mo ago

> but we need to avoid putting their works in production, or in front of other humans, without assessing it by any possible mean.

I think this is good advice in general, maybe with an emphasis on public vs. private, friendly contact. Having 0 thought AI slop thrown at you out of the blue is rude. "could have been a prompt" indeed. But having a friend/colleague ask for a quick glance at something they know you handle well is another story for me.

If I've worked on a subject for a few years, and know the particulars in and out, I'd have no trouble skimming something that a friend or a colleague sent me. I am sparing those 5-10 minutes for the friend, not for what they sent. And for an expert in a particular domain, often 5 minutes is all it takes for a "lgtm" or "lol no".

dabinat1mo ago· 1 in thread

I feel like this experiment was successful because those prompting the AI were knowledgeable enough to ask the right questions and verify the output was correct. This shows that there is still a place for expertise, even if the LLM does the actual research.

colechristensen1mo ago

I feel my input to LLMs is most valuable in the initial idea, big picture design tweaks, and the vast majority of my usefulness is negative feedback. This looks wrong, you've gotten off track, you're cheating with workarounds, you're falling into a rabbithole, etc.

incrediblylarge1mo ago· 1 in thread

A month ago my PhD supervisor told me it rips on proofs but he also said it's useless when formalising arguments in Lean - is this still the case?

vjerancrnjak1mo ago

Nope. Codex formalizes much better than any tool with exception of Aristotle from Harmonic.

https://github.com/vjeranc/fixed-rtrt

M3 module was formalized fully purely from experimental data and from a nudge by earlier versions of codex in 15-30 minutes in a simple write/compile/fix-first-error loop. I was a bit surprised how fast it picked up the pattern but given there was a paper from '70s it became clear why later.

ionwake1mo ago· 1 in thread

one thing I was wondering, is, if LLMs are word completions seemingly coming up with new solutions could this just be because stuff that was kept secret and now - is no longer is due to ingestion? I dont know enough about it tho

dist-epoch1mo ago

why would you keep secret this particular mathematical idea? it's not extraordinarily important, it's not on the path to some other major result, doesn't seem useful in financial trading. even author calls it good reasonable problem for a PhD thesis.

rklampp1mo ago· 1 in thread

Gowers has always been a proponent of Lean (naturally). He receives funding from the "AI for Math" fund, which is sponsored by a fund that is a front organization for venture capitalists:

https://www.renaissancephilanthropy.org/

The "brighter future" of course is that everyone is redundant and all capital is further concentrated.

It is always Gowers, Tao and Lichtman (math.ínc startup) who are pushing these technologies.

logicprog1mo ago

> It is always Gowers, Tao and Lichtman (math.ínc startup) who are pushing these technologies.

In your mind does this mean that they are lying, or driven by motivated reasoning and cognitive bias, or whatever you'd like to say?

Because I feel like people bring up these facts as a way to discount everything that these people are saying, but whether or not they've chosen to align themselves with AI aligned venture capital funding or not. The question is really, did what they say is happening happen or not? Are these capabilities real or not?

To my mind, mathematics is pretty definitely, externally, objectively verifiable, so it would be easy to catch them in a lie. In the case of the Erdös problem that was recently solved in a novel and productive way, it wasn't even initiated by them and the chat GPT transcript is public for all to see. And the proof could easily be verified by other people, for instance.

In addition, I think it's unlikely that they're not explaining things as they honestly see them and also doing their due diligence to make sure that they are seeing them as close to correctly as possible. Because their positions with these organizations not to mention their entire reputation and life's work and passion depends on their reputation in academic mathematics. If they were to give that up by falsifying these claims or not verifying them sufficiently, they would lose everything.

I think it's also worth pointing out that it is totally possible for someone to align themselves with such organizations after the fact because they agree with them instead of being bought out by such organizations. Otherwise, it would be possible to dismiss the opinion of anyone working at any NGO dedicated to being against AI and denying AI's capabilities or whatever, as well by the same logic of their salary being paid by an organization dedicated to pushing those ideas.

momojo1mo ago

Sorry, I'm reposting a comment I made yesterday that seems fitting:

> This reminds me of Antirez's "Don't fall into the anti-AI hype". In a sentence: These foundation models are really good at optimizing these extremely high level, extremely well defined problem spaces (ie multiply matrices faster). In Antirez's case, it's "make Redis faster".

readgrounded1mo ago

Quantitative finance went through a smaller version of this in the 2010s. The apprenticeship was building a Black-Scholes pricer from scratch, then a vol surface, then a calibration loop. Sweat problems that taught you what the math meant. Then libraries got good, platforms got good, and a junior could be productive without ever reeeeeally knowing how it worked. On some level yes the finer detailed knowledge is going to be lost because it gets locked in but in some ways we do get a "higher level api" to presumably solve more difficult problems.

locknitpicker1mo ago

From the article:

> Conversely, for problems where one’s initial reaction is to be impressed that an LLM has come up with a clever argument, it often turns out on closer inspection that there are precedents for those arguments, so it is still just about possible to comfort oneself that LLMs are merely putting together existing knowledge rather than having truly original ideas. How much of a comfort that is I will not discuss here, other than to note that quite a lot of perfectly good human mathematics consists in putting together existing knowledge and proof techniques.

This is exactly what leads me to believe that the real impact of LLMs in human history is yet to come. My work as a researcher was mostly spent on two classes of workloads: reading papers that were recently published to gather ideas and keep up with the state of the art, and work on a selection of ideas gathered from said papers to build my research upon. It turns out that LLMs excel at the most critical component of both workloads: parsing existing content and use it when prompting the model to generate additional content based on specific goals and constraints. I mean, papers are already a way to store and distribute context.

adaml_6231mo ago

"It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour"

This comment about time is very interesting to me. I know it's "just" doing mathematical proofs but the possibilities of speeding up planning, proposals and decision making in the physical world should excite people.

lysecret1mo ago

There is a great recent episode of latent space about a similar topic it’s worth a watch even with the click baiti thumbnail and title https://youtu.be/9d899Ram9Bs?is=pQMoVmlWVsTNKfRK

dpweb1mo ago

Don’t quite agree with the implication that if the answer to a problem is readily available (from an LLM) there is no use in the struggle to find the solution.

However I think it’s very important to approach such questions objectively, or at least self uninterested, and not as one who’s worried about one’s job or sense of self worth threatened by LLM technology.

The value is in the development of one’s own mental faculties. In math classes they tell you you have to work the problems. Even if LLMs become capable of solving entire classes of problems that that set expands over time, the value in developing one’s ability never goes out of style.

goopthink1mo ago

An interesting takeaway is that heretofore most of that advances have been not from “invention” but from a breadth of visibility. LLMs have been able to be “creative” because of the volume of work that they cover and can draw lines and associations between, not in discovering things that did not exist previously (though an argument can be made that something like AlphaFold was “discovering” and “intuiting” associations that were not explicit anywhere previously, uniquely found by the AI… but I’d argue back something about the bitter lesson and we’d go on for more than a few threads).

Somewhat ironic then, to not make this more explicit in an article about solving a combinatorial problem.

eranation1mo ago

Like coding, if you get inspired by AI for a novel idea, and can reproduce the same result independently (could code the same thing by hand) or at least understand and check every single argument (self review your code, test on your machine) and get it peer reviewed (code review, but with a real human) then I don’t see why the industry accepts the latest iteration of ChatGPT being 99% written by codex, but rejects a valid math result inspired by it.

arjie1mo ago

The question of where the creative input is was a big thing around Experiments in Musical Intelligence and co-composing. But it seems perhaps that it’s a transient state we needn’t spend too much effort it. The machine has failed to disappoint repeatedly. Perhaps this is as far as it gets or perhaps we will be like people in Catching Crumbs by the Table by Ted Chiang where almost all science is interpretation of papers by vastly greater intellects.

iandanforth1mo ago

I found the section on publishing very interesting. Even if the quality of the output is up to snuff, where should it go? Arxiv doesn't allow AI written work. The author proposes that only work that has been certified by human should be published. However, now the field is in the same boat as software engineering where we are facing a glut of pull requests and not enough time and people to review them.

robeym1mo ago

I think progressing as humans is something to be proud of. I care less about who gets credit and more about what we can now do.

I also do not think this makes people less capable of solving hard problems. The bar just moves up. More people can now work on harder problems with better tools.

If the goal is credit or proving real skill, then focus on harder problems, like ones AI can't reach

highfrequency1mo ago

> LLMs have got to the point where if a problem has an easy argument that for one reason or another human mathematicians have missed (that reason sometimes, but not always, being that the problem has not received all that much attention), then there is a good chance that the LLMs will spot it.

amelius1mo ago

Makes sense as a mathematician basically has two powers (1) using their intuition and (2) an enormous amount of mental stamina. A mathematician builds their intuition by reading maths books. It is thus not surprising that an LLM is well equipped to take over the tasks of the mathematician.

casey21mo ago

I think mathematicians like LLMs because this is the first time we have something like a computer for the kinds of math most people do, high level, hand wavy abstractions that are (relatively) easy for people to grok but hard to explain to traditional computers.

__rito__1mo ago

> So maybe there should be a different repository where AI-produced results can live.

Does the author know about CAISc 2026 [0]?

[0]: https://caisc2026.github.io

MagicMoonlight1mo ago

ChatGPT pro is garbage. It’ll spend 20 minutes on an answer, doing all kinds of ridiculous things like writing scripts… instead of just outputting plaintext.

And then the answer isn’t even right.

chalr1mo ago

There have always been attempts at settling all mathematics by using mechanized approaches. Often by mathematicians who already had made an impact and then wanted an automated approach.

The Bourbaki group was one of the first who attempted a mechanized approach (using pen and paper still of course) to set theory and were literally accused of wanting to end all mathematics. The approach was largely ignored in practice.

Gowers and a handful of others who work on computerized approaches also seem to want to end human mathematics and have sharecropper mathematics for a monthly tithe. So far they are largely ignored in practice.

zingar1mo ago

The post talks about LLM+human contributions being recognized in some different category from human-only. But is it possible to spot the difference between the two?

ortusdux1mo ago

An important lesson in web/blog design - I cannot for the life of me figure out who this author is (using only the website).

tmp104232884421mo ago

It's interesting that ChatGPT Pro is the real deal that can write novel physics or math papers (for certain values of novel), while Claude Pro is crap that, depending on the A/B test, may not even provide Claude Code or at the very least doesn't provide Opus. Shows how LLM naming conventions are currently a mess.

theptip1mo ago

> what should we do with this kind of content? Had the result been produced by a human mathematician, it would definitely have been publishable, so I think it would be wrong to describe it as AI slop. On the other hand, it seems pointless even to think about putting it in a journal, since it can be made freely available, and nobody needs “credit” for it (except that Isaac deserves plenty of credit for creating the framework on which ChatGPT could build). I understand that arXiv has a policy against accepting AI-written content, which makes good sense to me. So maybe there should be a different repository where AI-produced results can live. But various decisions would need to be made about how it was organized.

Interesting question, I guess a starting point is “moltbook”, but perhaps a better one is something like GitHub, where Lean proofs and preprints can go, and trending items can get boosted.

I also think that posting this stuff on x or bluesky has merit, but again the existing paradigm doesn’t quite work; perhaps you can create a completely separate identity for your agent (à la Moltbook) but I think you want some sort of reputational association with the human piloting the agent, at least for now. (Maybe eventually there are enough agents critically engaging with content so that “interesting” results get agent likes, and so we’ll-piloted agents stand on their own merit.)

alpama1mo ago

This is scary. Ai is growing faster than our knowledge. We are not prepared

jacktu1mo ago

That's a real shift. The value of an open problem used to be that it was unsolved. Now an open problem needs to be unsolvable by something that can read the entire literature and try a hundred approaches in an hour.

zuogl1mo ago

The HTML generation is surprisingly good because the training corpus for markup is cleaner than most programming languages.

OsamaJaber1mo ago

the bottleneck isn't generation, it's verification

richard_chase1mo ago

Dude needs to step down off his pedestal before he gets knocked down.

zkmon1mo ago

sexylinux1mo ago

Unfortunately it still does create errors.

This is of enormous importance but still is being actively ignored by many professionals or dismissed as as a minor issue.

Our emotional human brains are very enthusiastic about these new kind of "intelligent" products ("partners") and we want to believe so hard that they are finally "there" that we tend to ignore how big of a problem it is that LLMs carry a fundamental design problem with them that will make them produce errors even when we use a grotesque amount of resources to build "bigger" versions of them. The potential for errors will never go away with the current AI architecture.

This is a fundamental paradigm shift in computing. Instead of putting a lot of energy into building an architecture that will produce reliable results, we are now maximizing on a system / idea that will never give us 100% reliable results.

Basically it is just a marketing stunt. Probably the computer science guy building it knew very well that he would still need some fundamental break troughs to get to a real product, but the marketing guy saw that there is still potential to make a lot of money by selling a product that will produce correct results only 80% of the time.

The marketing guy was right and marketing is now dominating science, but humanity will pay a big price for that.

Putting enormous amounts of money into a fundamentally flawed system that we can not optimize to produce reliably error free results is just stupid.

The big achievement of "classical" computing is that the results are reliably error free. We have still some known issues eg. with floating point math and bad blocks on disk / bit flipping etc. but these are observable and we can handle / avoid them. Generally "non-ai-computing" was made so reliable, that we can depend on it for many very important things. This came not by accident but was created by a lot of people who put a lot of resources into research to achieve that result.

LLMs introduce a level of uncertainty and unreliability into computing that makes them practically useless.

Because if you have enough knowledge to verify the result and AI is only quicker in producing the result, what is the point then putting so much resources in it (besides making money by re-centralizing computing, of course). Verifying a lot of results that have been produced quicker is still slow, so the people who are now just AI verifiers should just produce the results themselves, makes the whole process quicker.

AI is only of value if it can produce results about things that you or your organization does not know anything about. But these results you can not verify and therefore potentially wrong results can be fatal for you, your organization and all the people that are affected by actions generated based on these wrong results.

Many people have already been killed because decision makers are not able to follow that very simple logic.

So we can still create "interesting and enjoyable results", but finally it is a gigantic miss-allocation of resources of historic idiocy. It fits, of course, very well in a timeline where grifters are on top of societies around the world.

It is a fundamentally wrong path that should not be followed and scientists around the world should articulate exactly that instead of producing marketing blog posts for a system with such fatal inherent issues.

j / k navigate · click thread line to collapse

535 comments

275 comments · 58 top-level

ziotom781mo ago· 54 in thread

nopinsight1mo ago

Moreover, there's no reason to believe the progress of LLMs, which couldn't reliably solve high-school math problems just 3–4 years ago, will stop anytime soon.

You might want to track the progress of these models on the CritPt benchmark, which is built on *unpublished, research-level* physics problems:

https://critpt.com/

Frontier models are still nowhere near solving it, but progress has been rapid.

* o3 (high) <1.5 years ago was at 1.4%

* GPT 5.4 (xhigh), 23.4%

* GPT-5.5 (xhigh), 27.1%

* GPT-5.5 Pro (xhigh) 30.6%.

https://artificialanalysis.ai/evaluations/critpt.

FrojoS1mo ago

> there's no reason to believe the progress of LLMs [...] will stop anytime soon

Wrong. Every advancement has followed a s curve. Where we are on that curve is anyones guess. Or maybe "this time its different".

14 more replies

civvv1mo ago

There are many indications that model progress is slowing down, so that is not entirely accurate.

3 more replies

Davidzheng1mo ago

Deep think still makes many many many more mistakes than gpt 5.5 pro on math

maximamas1mo ago

jillesvangurp1mo ago

2 more replies

ziotom781mo ago

tags2k1mo ago

illiac7861mo ago

It’s also because it is so annoying to have to manage the memory of the LLM with custom prompts/instructions manually.

timschmidt1mo ago

illiac7861mo ago

yeah, I should have been more specific: I meant the type of learning that mentoring fosters, the long term learning.

2 more replies

kybernetikos1mo ago

However, I think it's important to remember that LLMs are embedded in larger systems, and those larger systems do learn.

2 more replies

freedomben1mo ago

stingraycharles1mo ago

> Using the word “Mentoring” is anthropomorphic and subconsciously makes you think it will learn.

Heck the whole name of “machine learning” suggests these things can actually learn. “reasoning” suggests that these things can reason, instead of being fancy, directed autocomplete. Etc.

In other news: data hydration doesn’t actually make your data wet. People use / misuse words all the time, and that causes their meaning to evolve.

kasey_junk1mo ago

And that can be very hard to do given the ui we most interact with them in is a chat session.

1 more reply

DoctorOetker1mo ago

> ... that something as smart as an LLM does not learn.

it's not because model providers don't want to provide user specific continual learning, that we don't know how to do it.

it would be a lot more expensive to host user-specific model weights, and would prevent amortizing the weights over many requests in batches...

_the_inflator1mo ago

I agree and put it this way: LLMs sound so convincing presenting you the work it does rose colored and promising to give you more if you keep going.

There is a 50/50 chance that it turns out to be right or letting you jump of the cliff.

Only the trip stays the same beautiful 5 star plus travel.

Also, spotting an error and telling LLM makes it in most cases worse, because the LLM wants to please you and goes on to apologize and change course.

The moment I find myself in such a situation I save or cancel the session and start from scratch in most cases or pivot with drastic measures.

Gemini to me is the most unpredictable LLM while GPT works best overall for me.

Reasoning doesn’t help much in the Coding domain for me because it is very high level and formally right what the LLM comes up with as an explanation.

MattPalmer10861mo ago

Reusing the same prompt several times is something I've started doing too. The contrast is often illuminating.

In one case, it made a thoroughly convincing argument that an approach was justified. The second time it made exactly the opposite argument, which was equally compelling.

I now see LLMs as persuasion machines.

scotty791mo ago

Before AI happened I watched youtube. Occasionally I encountered there very convincing arguments. Same person often made very convincing arguments on many subjects.

But noticed that the closer the domain they were talking about was to my area of competence the less convincing their arguments were. There were more holes, errors and wrong conclusions.

I recalibrated my bs meter thanks to that.

Since AI came I successfully used this strategy of being extremely cautious towards convincing arguments to not become mislead by AI.

For me basically AI achieved the threshold of useful reliability for any domain that humans are reliable at.

I don't really care about sycophancy. I might have a slight advantage that I don't talk to AI in my native language. So its responses don't have a direct line to my emotions.

eitally1mo ago

For this sort of thing, using multiple LLMs is extremely helpful.

taneq1mo ago

Ever since they started getting really sycophantic, I’ve been presenting my ideas as “my co-worker says this is a good approach but I disagree, can you help me convince him that it’s wrong?”

pbhjpbhj1mo ago

>LLM wants to please you

I was using Copilot and asked it a question about a PDF file (a concept search). It turned out the file was images of text. I was anticipating that and had the text ready to paste in.

Instead, it started writing an OCR program in python.

I stopped it after several minutes.

Often Copilot says it can't do something (sometimes it's even correct), that's preferential to the try-hard behaviour here.

freedomben1mo ago

> Gemini to me is the most unpredictable LLM while GPT works best overall for me.

miki1232111mo ago

I think that ultimately, the largest change brought on by LLMs will be due not to their intelligence, but to their tenacity.

You don't need that much intelligence to do that, you just need somebody who's willing to dedicate their life to knowing everything there is to know about that guy from Louisiana.

With humans, the amount of money you'd need to pay such a person just isn't worth the reward. With LLMs, it may very well be.

mixtureoftakes1mo ago

please, sign up for a paid plan of either chatgpt or claude. gemini is while close, still noticeably behind

you deserve opinions shaped by interactions with the best tools that are out there.

wg01mo ago

Gemini feels deep and philosophical. Especially for product management. Tell him you're a product manager and we're a team of two.

But regular reminder - All LLMs can be wrong all the time. I only work with LLMs in domains I'm expert in OR I have other sources to verify their output with utmost certainty.

wafflemaker1mo ago

Or when you don't care about results being very correct.

smartmic1mo ago

> I only work with LLMs in domains I'm expert in

This. Should become a general rule for any non-trivial use of LLM in a professionel setting.

1 more reply

peyton1mo ago

Seriously, it’s not worth reaching for less intelligence. Use Extended Pro 100% of the time for things you’d spend the amount of time GP spent writing their post.

cubefox1mo ago

Gemini is certainly not behind Claude in terms of physics.

ainch1mo ago

hodgehog111mo ago

ChatGPT and Gemini are actually fairly comparable.

Quothling1mo ago

stasomatic1mo ago

Are you using this agent hive for any repeatable tasks? What you described, superficially, seems like a one off. Genuinely curious.

1 more reply

cyanydeez1mo ago

At some point, a rubicon will be crossed where these systems can't fallback to a human operator and will fail spectacularly.

pbhjpbhj1mo ago

It is troubling. It suggests a plateauing of human understanding.

regularfry1mo ago

It absolutely is where the learning is, that's pretty well established brain science.

1 more reply

regularfry1mo ago

What that means practically is that we've got a generation - 25 years or less - to evolve these things not to need the fallback. If such a thing is possible.

leptons1mo ago

We're on the road to Idiocracy.

wccrawford1mo ago

recursivecaveat1mo ago

northzen1mo ago

Hi ziotom! I wonder about you work in 3D Cifford Algebras. May you share some links to the research you do? I also have interest in this topic I research on my own.

Just in case if you don't want to disclose your name my email is northzen@gmail.com

port111mo ago

quantummagic1mo ago

> Gemini’s smug...

Anthropomorphizing these systems is dangerous, whether coming from the bullish or bearish perspective. The output is statistically generated by a machine lacking the capability to be smug.

2 more replies

danielparsons1mo ago

I find exactly the same for legal analysis. Great at ideation and proofreading but frequently misunderstands concepts and hallucinates conclusions from faulty premises.

wslh1mo ago

I assume that once LLMs are trained with large [synthetic] information about 3D Clifford algebras it will work better.

tasuki1mo ago

> in 3D Clifford algebras it repeatedly confuses exponential of bivectors and of pseudoscalars.

jiggawatts1mo ago

Ironically, it's sort of the other way around! Every frontier chatbot since GPT 4 (at least) has had a pretty good understanding of even very esoteric technical concepts.

Bivectors and pseudoscalars (in a 3D context) are "just" signed areas and volumes. Easy!

Back around the GPT 3, 3.5, and 4.0 era I used to ask the bots to explain "counterfactual determinism", which is one of the most complex topics I personally understand.

Then I would lie to the bot about it, and see if it corrected me or not.

This test is useless now, the frontier models can't be fooled any longer on such "basic" concepts.

eth0up1mo ago

Any experience with NotebookLM?

Mine has been epically bad.

DeathArrow1mo ago

I don't think the experience with Gemini will be the same when using GPT.

wood_spirit1mo ago

Chiming in to agree but clarify that the latest sota models are no better than Gemini.

I look forward to the day one a small open model that we can run ourselves outperforms the sum of all today’s models. That’s when enough is enough and we can let things plateau.

energy1231mo ago

Basically all Erdos problems that get solved with AI use ChatGPT 5.* Pro, not Gemini/Opus.

1 more reply

ed_balls1mo ago

intern that never sleeps

ieieaaa1mo ago

LLM’s are the most powerful tool invented to search across a huge information space in response to human input.

That’s all they are. They don’t ‘know’ anything intrinsically and do know ‘know’ what reasoning even is.

NotOscarWilde1mo ago· 27 in thread

As a TCS assistant professor from Eastern Europe, I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models.

As a nail to the coffin, I was "denied" all Claude Opus recently as part of Microsoft's clampdown on individual (and academic) use of Copilot.

(Chagpt 5.5 Plus does not seem sufficient for any deeper investigations into new research topics, I've tried.)

Apologies for the rant.

vthallam1mo ago

@NotOscarWilde drop your email here, I will reach out and happy to get you a pro account for a few months so you can try 5.5 pro.(work at OAI)

teiferer1mo ago

NotOscarWilde1mo ago

You're absolutely right (pun intended).

An aside: It was a very nice gesture and completely unexpected by me, so even if it doesn't work out, it made my day. I personally believe that kind gestures have a lot of power.

1 more reply

Scea911mo ago

2 more replies

snayan1mo ago

I'd argue that the OpenAI dude/dudettes level of generosity is appropriate given the circumstances.

NotOscarWilde1mo ago

This requires a major "dox" of myself, but I am really grateful for the offer, so these are my academic contacts:

https://pastebin.com/hNYrCjhL

I probably will erase the contents in a few days.

Even if you just drop an email and it doesn't work out, I appreciate this gesture so much. Thank you.

vthallam1mo ago

Got the contact, will reach out tomorrow, you can delete them.

thierrydamiba1mo ago

Shoutout to you-I will match it if they need other resources. (I don’t work at OAI, just think this is cool)

alsetmusic1mo ago

You know what, I'm ashamed that I didn't think of this. I'll sponsor three months. Email in my hn profile. I don't understand the math in the article, but I'd love to help you make progress in it.

1 more reply

NotOscarWilde1mo ago

Thank you.

lelanthran1mo ago

This doesn't solve the problem, though: having the ability to finish a field of study without paying a toll to a token provider.

layer81mo ago

And what should he do after the few months?

johndough1mo ago

alsetmusic1mo ago

It’s a classic example of the best positioned people being in the best position to keep reaping all the rewards.

huijzer1mo ago

I know the example, but as a counter-argument: often more expensive boots are not more durable. It’s about spending time to learn to spot the quality.

NotOscarWilde1mo ago

> Learning to do more with less money isn’t as bad as many people think.

1 more reply

m_mueller1mo ago

bambax1mo ago

johndough1mo ago

bambax1mo ago

Depends very much on usage! If you connect it to tools like Cursor, etc. then yes a subscription is probably cheaper -- although, you'd have to subscribe to each provider if you want to use them all.

But if you ask questions occasionally, (and don't resend, for example, your whole codebase with each request), then the API feels really cheap, even for the frontier models.

tasuki1mo ago

nerdsniper1mo ago

I'm not trying to shame here, just curious whether this is completely unattainable for most researchers in your area.

Computer01mo ago

It appears that in their country someone in their position makes about 50k usd annually. I make a similar amount in my country and cannot justify it.

ziotom781mo ago

qq661mo ago

Paste what you want me to ask 5.5 Pro and I'll paste you the response.

bwfan1231mo ago

> I always am a little jealous of the biggest names in math having such an easy access to the expensive, long thinking models

I am starting to see folks saying - ok, so LLMs can do this, what value have you added ? modulo llm is becoming the norm.

iberator1mo ago

Its good. You should work hard if you are in the public sector! Using claude (cheating) and getting check from government is unethical

pmontra1mo ago· 17 in thread

It's a very long post with a mix of technical (math) and philosophical sections. Here are the most striking points to reflect upon IMHO.

Training must start from the basics though. Of course everybody's training in math starts with summing small integers, which calculators have been doing without any mistake since a long time.

The point is perhaps confirmed by another comment further down in the post

bambax1mo ago

tempaccount50501mo ago

_vertigo1mo ago

doginasuit1mo ago

2 more replies

kobe_bryant1mo ago

so how would an LLM being able to do your job help you afford a place to live

slibhb1mo ago

According to the blog post linked in the OP, the LLM-generated results were read, understood, and confirmed by the mathematician whose work they built on.

1 more reply

joe_mamba1mo ago

> They literally have no value whatsoever; they're a passthrough; they're invisible.

Then middle management also have no value, since they're also a passthrough between upper management and ICs, yet they never went extinct.

1 more reply

panarky1mo ago

I delegate writing a binary executable to the compiler and the linker.

I don't know or understand the binary executable.

I don't own the binary executable, I don't understand it, I can't explain it, it's not my work.

I'm a passthrough; I'm invisible.

I have literally no value whatsoever.

kerabatsos1mo ago

But perhaps we should regard it as a major achievement.

lmpdev1mo ago

I mean in the same way getting Wolfram Alpha to solve a really hard/ugly differential equation I suppose

aspenmartin1mo ago

Insane that we have a system capable of making innovative math proofs and people dismiss it as unimpressive

1 more reply

palata1mo ago

I feel like you slightly miss both points.

> Training must start from the basics though.

Sure, but the point is that at some point (e.g. when starting a PhD) one needs to do research, not learn the basics. And LLMs make that harder, because they solve the "easy research" part.

> People pay coders to build stuff that they will use to make money and I can happily use an AI to deliver faster and keep being hired.

Again, that's true but missing the point: if you never get to be a "good coder", you will always be a "bad vibe coder". Maybe you can make money out of it, but the point was about becoming good.

sdeframond1mo ago

> Would we regard that as a major achievement of the mathematician? I don’t think we would.

1. Does it matter, really? 2. Is it very different from previous computer-aided proofs, philosophically?

agnishom1mo ago

1. It matters because there are human mathematicians who pride themselves for their mathematical achievements. Mathematics is art to them.

1 more reply

layer81mo ago

It matters because most mathematicians thrive on the recognition of their achievements. If what you do any mediocre mathematician could have done, that takes away motivation and fulfillment.

andai1mo ago

> Training must start from the basics though. Of course everybody's training in math starts with summing small integers, which calculators have been doing without any mistake since a long time.

The alternative would be like, just asking ChatGPT to do your math homework and then "verifying" it by looking at it and saying "yeah, that looks okay." What are you going to learn?

We do stuff by hand for a reason.

coderenegade1mo ago

1 more reply

Jweb_Guru1mo ago· 16 in thread

So I think it will be a while before the impressive capabilities of these models really percolate into our lives as programmers, unless you're one of the lucky ones given unlimited access to 5.5 Pro.

elAhmo1mo ago

I swear that people have said the same thing with effectively every new model that came out in the last six months.

fluidcruft1mo ago

The reality is likely that everyone is hitting similar barriers and the solutions are somewhat generalizable and get added to training new models.

Eventually people will reach the new limits and the cycle repeats.

fasterik1mo ago

> I swear that people have said the same thing with effectively every new model

Jweb_Guru1mo ago

Many people may have, but I certainly haven't.

hackable_sand1mo ago

They did. The scam continues.

1 more reply

y1n01mo ago

> This jives with what I've experienced

Just as an fyi, the word you are looking for is jibes. Jive is something else entirely.

jibe1mo ago

I'm with you!

1 more reply

pfdietz1mo ago

Excuse me stewardess, I speak jive.

jfaat1mo ago

I'm going to start using malapropisms so people know I didn't use an llm to write things

boring-human1mo ago

Cut me some slack, Jack.

billfor1mo ago

Blame The Bee Gees: https://en.wikipedia.org/wiki/Jive_Talkin'

1 more reply

bicepjai1mo ago

Interesting I did not know that I would have used jives :) thanks

refulgentis1mo ago

That ship sailed looooong ago.

1 more reply

idiotsecant1mo ago

The only thing worse than complaining about this is being the guy complaining about the guy complaining about this. So congratulations on being second most annoying.

1 more reply

Forgeties791mo ago

> can wrangle into solving tedious, but straightforward, problems correctly. It still makes a ton of mistakes and needs to be very rigidly guided,

Jweb_Guru1mo ago

MrDrDr1mo ago· 15 in thread

charleshn1mo ago

Yes, they can.

Some people like to parrot "next token prediction", "LLMs can only interpolate", and other nonsense, but it is obviously not true for many reasons, in particular since we introduced RL.

Humans do not have the monopoly on generating novel ideas, modern AI models using post training, RL etc can come to them in the same way we do, exploration.

See also verifier's law [0]: "The ease of training AI to solve a task is proportional to how verifiable the task is. All tasks that are possible to solve and easy to verify will be solved by AI."

This applied to chess, go, strategy games, and we can now see it applying to mathematics, algorithmic problems, etc.

It is incredibly humbling to see AI outperform humans at creative cognitive tasks, and realise that the bitter lesson [1] applies so generally, but here we are.

[0] https://www.jasonwei.net/blog/asymmetry-of-verification-and-...

[1] http://www.incompleteideas.net/IncIdeas/BitterLesson.html

vld_chk1mo ago

energy1231mo ago

2 more replies

jdub1mo ago

charleshn1mo ago

> Some people like to parrot "next token prediction", "LLMs can only interpolate", and other nonsense

Thank you for illustrating my point.

eterm1mo ago

My own take, and it's veering into the Philosophy of Mathematics, but there's a debate about whether Mathematics is "Invented" or "Discovered".

If it's "invented", then it requires ingenuity.

If it's "discovered", then it was always already there, just waiting for the right connections to be made for it to be uncovered and represented in a way we can understand.

layer81mo ago

This also holds for physical inventions. You discover a working way to build some functioning mechanism. It’s a process of discovery of what is possible in the physical world.

Whether you portray somethings as a discovery or as an invention is more a matter of degree, a matter of from which angle one is looking at it.

MrDrDr1mo ago

1 more reply

eiieue1mo ago

Mathematical objects are an invention of the mind - they are abstract objects that only an entity who can process abstractions can make sense of.

There is no ‘discovery’ here nor was it waiting to be found. The human has to sacrifice and pursue the path of exploring reality and thereby is inherently inventing.

Humans built up mathematics iteratively from smaller bases extending into large ones. Is this what LLM’s do? Of course not - They are fed with vast amounts of information from the off.

LiamPowell1mo ago

humanfromearth91mo ago

This works really well.

Now, it's clear that I have no idea how much of this is something we would consider new and original, and how much is a kind of systematic, but not novel, easy of thinking.

realtalk_sp1mo ago

jasfi1mo ago

It's about the ability to combine ideas in novel ways, without breaking the rules in relevant frameworks. Sometimes the idea may even be to contradict existing theories where they are weak.

ikari_pl1mo ago

How do you define a new idea?

To me, it's rearranging the information you had in a way that hasn't been applied or published before.

That's literally what LLMs are built for.

eiieue1mo ago

Theres a simple test for this.

Limit the knowledge an llm to some point in time at which a discovery was made. And check to see if the llm could produce the discovery.

If you think OAI hasn’t already tried this then think again - they have every incentive to do so and announce it to the world.

MinimalAction1mo ago· 13 in thread

hodgehog111mo ago

As someone who is much further down the track, I would kindly suggest you drop that line of thought. I've seen far too many brilliant and ambitious people drop into depression because of it.

Don't work for the possibility of eternal glory though, it just doesn't exist anymore.

MinimalAction1mo ago

1 more reply

whatever1201mo ago

You are worthy. You will hone your skills in grad school and be able to command these AIs better than somebody who hasn’t struggled with hard problems for a long time.

jlarcombe1mo ago

A depressing thought that all that work is just so you can "command AIs better"

folderquestion1mo ago

It could happen than the AI, in a near future, is not something external but just a part of your brain, so you retain the glory.

2 more replies

alexashka1mo ago

All that work to kick a ball into a net.

Nobody looks at this species and goes hm, rational and reasonable :)

helloplanets1mo ago

"If you value intelligence above all other human qualities, you're gonna have a bad time." - Ilya Sutskever, 2023

timedude1mo ago

ionwake1mo ago

I feel bravery transcends time better than the odd scientific breakthrough which are often attributed to one, but whose roots came from a "lesser" unknown

kranke1551mo ago

try meditation.

MinimalAction1mo ago

Thanks. This might help. Are you suggesting any particular form?

2 more replies

alexashka1mo ago

> I always believed that my work speaks for itself and transcends beyond my limited time on this cosmic experience

Any statement preceded by the word 'believe' is a coping mechanism.

> This notion of immortality was just a small intangible bonus I hoped for when I jumped into grad school

Any statement preceded by the word 'hope' is a coping mechanism.

> AI is making me feel less worthy

Worth comes from understanding, not achievement.

MinimalAction1mo ago

I strongly disagree on beliefs and hopes are coping mechanisms. Coping from what? Beliefs and hopes are what they are.

But I agree worth should be derived from understanding, not through achievement.

1 more reply

mxwsn1mo ago· 8 in thread

pmontra1mo ago

I replied to a comment about AI in sports and I build on that.

In the case of math, the human can lead the LLM on the right track, point it to a problem or to another one. So it deserves some praise.

Then the team that built the car, cared about the horse, built the AI might deserve even more praise but we tend to care more about the single most visible human.

butlike1mo ago

dmbche1mo ago

Could you win an F1 race with the latest winning car against F1 drivers?

3 more replies

djeastm1mo ago

>Would we regard that as a major achievement of the mathematician? I don’t think we would

For some reason this reminds me of AI images and a domain like comedy.

So if a mathematician comes up with an amazing result that an LLM "did", I think they could still get a bit of credit for prompting it to do it and being its guide.

But whereas the first person could perhaps be called a comedian and not an artist, would the mathematician still be called a mathematician or something else?

gobdovan1mo ago

5424581mo ago

> given all the billionaire mathematicians

We just call those ones “quant traders”.

ComplexSystems1mo ago

bambax1mo ago

It may not be a major achievement by the mathematician (although it's debatable) but it would still be a major result.

few1mo ago· 8 in thread

This made me a little sad

mentos1mo ago

I watched the movie '21' (2008) for free on YouTube yesterday.

The opening of the movie features the MIT campus full of students navigating its grounds and all the promise and status that higher education brings. [0]

Gave me the same sense of sadness realizing how much will fall to AI.

[0] - https://youtu.be/0lsUsWdkk0Y?si=TJl7f_b1RcWcDqF8&t=278

lexoj1mo ago

Not free in my country, didn’t know YouTube was broadcasting full movies in certain regions as you imply.

1 more reply

vessenes1mo ago

jdale271mo ago

hodgehog111mo ago

bananaflag1mo ago

Now repeat that for every sort of human achievement

bel81mo ago

Machines are comming even after table tennis :(

https://www.youtube.com/watch?v=VVEzgYxDdrc

pmontra1mo ago

We care about sports with humans.

1 more reply

robot-wrangler1mo ago· 7 in thread

A very interesting comment from Baez, I'll just quote part of it.

> In other words, mathematicians may need to adjust to a transformation from a scarcity economy to an abundance economy.

https://gowers.wordpress.com/2026/05/08/a-recent-experience-...

zarzavat1mo ago

There are three species of mathematicians:

The first species is the pure problem solver. Tao is the poster child for this group. Their currency is interesting problems and solutions to those problems.

The third species is the applied mathematician. They see mathematics as a means to an end, they have some problem outside of mathematics and they want to use mathematics to solve it.

It seems like the first group (the problem solvers) are the most immediately threatened by AI, although so far AI is better at solving problems than finding new conjectures.

robot-wrangler1mo ago

Anyway, for now, dreamers will probably find more inspiration by cross-pollination with colleagues from different disciplines, or just going for a walk.

lhd11mo ago

I was expecting Grothendieck. Conway is hardly the poster child for theory building.

1 more reply

pocksuppet1mo ago

https://en.wikipedia.org/wiki/Paradox_of_value

qwrahg1mo ago

I note that it is always the same online pundits (even if they are distinguished academics) who push anything new.

Meanwhile Wiles and Perelman stayed offline and solved real problems.

robot-wrangler1mo ago

I don't necessarily think engaging in the personalities is interesting, but I'm struggling to see what is the beef here. Is it personalities? Pure vs applied math? Or AI?

DoctorOetker1mo ago

would Wiles be willing to transcribe his proof for the metamath verifier? it can be done offline indeed...

1 more reply

einrealist1mo ago· 7 in thread

Anyone spotting the issue here? What did that really cost?

sidkshatriya1mo ago

einrealist1mo ago

intexpress1mo ago

Not necessarily. Humans brains use a tiny amount of power. Most of the human cost would be due to the very high cost of housing in many locations.

1 more reply

hackable_sand1mo ago

Are you guys being intentionally obtuse?

Do you understand the difference between an apple and an apple tree?

colordrops1mo ago

einrealist1mo ago

Did I praise our animal agriculture anywhere?

colordrops1mo ago

We have a wide audience.

kang1mo ago· 4 in thread

5.5pro is amazing but this implication might not be true & is the core argument of this piece.

AI will prove all sort of things - interesting, boring & incorrect.

To sort it will be the task of the PhD.

layer81mo ago

kang1mo ago

Verification on its own is not research, but judgement is research.

ianm2181mo ago

Verification is generally a much lower bar than solution generation. I don’t think it’s likely sorting out the right from wrong will end up being this huge PhD level effort.

kang1mo ago

Verification & solution generation are both part of problem generation & defining the passing test - judgement.

YeGoblynQueenne1mo ago· 4 in thread

If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening.

whimsicalism1mo ago

thank you for the morning laugh

YeGoblynQueenne1mo ago

https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...

famouswaffles1mo ago

>If you can sic ChatGPT on a mathematics problem and it can solve it without your input, that's a different matter but that's not what's happening.

I mean that has happened so yeah ?

https://www.scientificamerican.com/article/amateur-armed-wit...

Actual GPT transcript. Zero such input https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...

That there are multiple of these stories in the last few months by the latest set of models (there are even more than these 2) should provoke this sort of consideration and discussion.

YeGoblynQueenne1mo ago

And btw there's scholarship that indicates vibe-solving is not yet ready to replace mathematicians like Timothy Gowers:

First Proof

https://arxiv.org/abs/2602.05192

See Appendix A for initial results.

1 more reply

globular-toast1mo ago· 4 in thread

I wish people would stop generating stuff they don't understand only to forward it to someone who does. Something about that really rubs me the wrong way.

hodgehog111mo ago

Also if he did send me complete junk, I would still parse it for multiple days to see what is there.

globular-toast1mo ago

1 more reply

auggierose1mo ago

Lol. If Gowers sends you a piece of math he doesn't quite understand because he thinks that you might, that is something you celebrate.

frozenseven1mo ago

You are criticizing a Fields Medalist for consulting with another mathematician.

CharlesLau1mo ago· 4 in thread

Is the assessment system of undergraduate mathematics education no longer effective?

margalabargala1mo ago

Graduate? Yes.

whatever1201mo ago

How should graduate school be changed then? Specifically for mathematics

2 more replies

dyauspitr1mo ago

crocdundae1mo ago

1 more reply

TrackerFF1mo ago· 3 in thread

If we look at it a bit cynically: Some researcher will be able to pump out a lot more papers by spending money on compute, than a couple of years of training students.

Interesting times. But also so much uncertainty. I feel terrible for the students that will have to decide now what they want to do, with all this knowledge.

robot-wrangler1mo ago

ndkap1mo ago

PhD students are already using AI models to work for them. Most of the PhD candidates I know have $200 Claude Max plan which they use to their fullest.

odyssey71mo ago

It’s not as if institutions have been lavishing PhD students with money up until now.

As an afterthought budget item, those funds aren’t exactly attractive targets to raid for pursuing an expensive, different process.

robot-wrangler1mo ago· 3 in thread

electriclove1mo ago

Marketing campaign???

energy1231mo ago

robot-wrangler1mo ago

> We know this because most of the Erdos proofs made by AI have been done by amateurs prompting GPT 5.* Pro

> which is too conspiratorial of an explanation for my liking regardless

OpenAI is not, in fact, open. Why do they deserve the benefit of the doubt?

1 more reply

quinndupont1mo ago· 3 in thread

[flagged]

vessenes1mo ago

I don’t love the tone here, but I do think you get at a key question in mathematical philosophy.

There are lots of strong feelings on both sides. For instance: “God created the integers, the rest is the creation of man” — Kronecker, 19th century sums up one particular perspective.

NB: My original comment led with a pejorative, which was rightly flagged.

dang1mo ago

> Who hurt you bro?

Please don't respond to a bad comment by breaking the site guidelines yourself. That only makes things worse.

https://news.ycombinator.com/newsguidelines.html

1 more reply

dang1mo ago

Can you please make your substantive points without fulminating, as the site guidelines request (https://news.ycombinator.com/newsguidelines.html)?

zkmon1mo ago· 2 in thread

>> but it was definitely a non-trivial extension of those ideas, and for a PhD student to find that extension it would be necessary to invest quite a bit of time digesting Isaac’s paper

svnt1mo ago

Too many people are wrapped around the ego axle thinking (assuming) their ideas are both them and somehow unique and special.

It usually takes dissolving that, often through difficult experiences, before they can see it as a machine, something that could be separated from them.

dag1001mo ago

I think the more pressing issue is that there isn't really much space left for humans in the economy if thinking can also be automated.

adammdaw1mo ago· 2 in thread

nxobject1mo ago

I wonder as well whether large-but-finite contexts can handle algebraic questions that require traversing up and down levels of abstraction, at least not without "thrashing".

hodgehog111mo ago

Indeed, analysis is a bit more loose in its arguments, and so I've found LLMs tend to make more mistakes there.

fulafel1mo ago· 2 in thread

Link to source blog post: https://gowers.wordpress.com/2026/05/08/a-recent-experience-...

dang1mo ago

That's the top link (i.e. that the title is linked to), no?

fulafel1mo ago

Indeed, the body in the post made me think it was a url-less submission.

1 more reply

dares25731mo ago· 2 in thread

I think the biggest advantage of ChatGPT compared to Claude is that there are fewer things outside the model itself, such as KYC, account bans, etc.

solenoid09371mo ago

This is just grossly misinformed.

tmp104232884421mo ago

Can you name any instance of OpenAI being as trigger-happy with bans as Anthropic has been in the past few months? Codex may have fewer prosumers, but they've added a lot in that time.

1 more reply

SubiculumCode1mo ago· 2 in thread

I honestly can't say this isn't AGI anymore. AGI shouldn't be a bar so taboo that it has to be at the extreme capability in every domain. What human is?

This is as AGI as it needs to be to get my vote. And it's scary.

MrScruff1mo ago

It's ASI with jagged intelligence, which is probably what it will remain for a while.

It still sounds to me like remarkable automation rather than something that's expanding the frontier of human knowledge, for now at least.

agiipullor1mo ago

to quote Demis Hassabis, "these models can solve frontieer problems in math, but also fail in really dumb ways at trivial questions - the car wash question".

jagged AGI

bambax1mo ago· 2 in thread

> quite a lot of perfectly good human mathematics consists in putting together existing knowledge and proof techniques

Most published research sits ignored and unread; AI can uncover and use everything.

imiric1mo ago

> Creativity is connecting ideas from different domains and see if something from one field applies to another.

And can we please stop calling pattern matching and generation "intelligence"? This farce has gone on long enough.

agiipullor1mo ago

> And can we please stop calling pattern matching and generation "intelligence"

thats literally what an IQ test tests - abstract pattern matching. but I guess you dont like IQ tests either

1 more reply

slopinthebag1mo ago· 2 in thread

AI generated article btw.

Maybe if you find AI to be doing stuff you find impressive, the stuff you were doing wasn't that impressive? Worth ruminating on your priors at least.

hodgehog111mo ago

This is beyond ridiculous to say considering whose blog this is.

Even without that knowledge, no, this article is certainly not AI generated. It has none of the tells.

reasonableklout1mo ago

What makes you think either the tweet or blog post are AI generated?

bustermellotron1mo ago· 1 in thread

34qJhah1mo ago

And Teichmüller thought that Germany would win WW2 and volunteered for the Eastern Front.

Being a gifted mathematician does not make you right. In fact, mathematicians have a lot of bizarre theories.

iTokio1mo ago· 1 in thread

On complex problems with lengthy proofs, the first step that I would have done is to ask 5.5 pro in a new, unrelated, session, to be very critical, to try to find flaws in the arguments.

And certainly not to send it to a fellow colleague to ask its opinion first.

Otherwise tech leads, maintainers, experts get overwhelmed and this is how the « AI slop » fatigue begins.

To be clear I’m talking about this step:

NitpickLawyer1mo ago

> but we need to avoid putting their works in production, or in front of other humans, without assessing it by any possible mean.

dabinat1mo ago· 1 in thread

colechristensen1mo ago

incrediblylarge1mo ago· 1 in thread

A month ago my PhD supervisor told me it rips on proofs but he also said it's useless when formalising arguments in Lean - is this still the case?

vjerancrnjak1mo ago

Nope. Codex formalizes much better than any tool with exception of Aristotle from Harmonic.

https://github.com/vjeranc/fixed-rtrt

ionwake1mo ago· 1 in thread

dist-epoch1mo ago

rklampp1mo ago· 1 in thread

Gowers has always been a proponent of Lean (naturally). He receives funding from the "AI for Math" fund, which is sponsored by a fund that is a front organization for venture capitalists:

https://www.renaissancephilanthropy.org/

The "brighter future" of course is that everyone is redundant and all capital is further concentrated.

It is always Gowers, Tao and Lichtman (math.ínc startup) who are pushing these technologies.

logicprog1mo ago

> It is always Gowers, Tao and Lichtman (math.ínc startup) who are pushing these technologies.

In your mind does this mean that they are lying, or driven by motivated reasoning and cognitive bias, or whatever you'd like to say?

momojo1mo ago

Sorry, I'm reposting a comment I made yesterday that seems fitting:

readgrounded1mo ago

locknitpicker1mo ago

From the article:

adaml_6231mo ago

"It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour"

lysecret1mo ago

There is a great recent episode of latent space about a similar topic it’s worth a watch even with the click baiti thumbnail and title https://youtu.be/9d899Ram9Bs?is=pQMoVmlWVsTNKfRK

dpweb1mo ago

Don’t quite agree with the implication that if the answer to a problem is readily available (from an LLM) there is no use in the struggle to find the solution.

goopthink1mo ago

Somewhat ironic then, to not make this more explicit in an article about solving a combinatorial problem.

eranation1mo ago

arjie1mo ago

iandanforth1mo ago

robeym1mo ago

I think progressing as humans is something to be proud of. I care less about who gets credit and more about what we can now do.

I also do not think this makes people less capable of solving hard problems. The bar just moves up. More people can now work on harder problems with better tools.

If the goal is credit or proving real skill, then focus on harder problems, like ones AI can't reach

highfrequency1mo ago

amelius1mo ago

casey21mo ago

__rito__1mo ago

> So maybe there should be a different repository where AI-produced results can live.

Does the author know about CAISc 2026 [0]?

[0]: https://caisc2026.github.io

MagicMoonlight1mo ago

ChatGPT pro is garbage. It’ll spend 20 minutes on an answer, doing all kinds of ridiculous things like writing scripts… instead of just outputting plaintext.

And then the answer isn’t even right.

chalr1mo ago

There have always been attempts at settling all mathematics by using mechanized approaches. Often by mathematicians who already had made an impact and then wanted an automated approach.

zingar1mo ago

The post talks about LLM+human contributions being recognized in some different category from human-only. But is it possible to spot the difference between the two?

ortusdux1mo ago

An important lesson in web/blog design - I cannot for the life of me figure out who this author is (using only the website).

tmp104232884421mo ago

theptip1mo ago

Interesting question, I guess a starting point is “moltbook”, but perhaps a better one is something like GitHub, where Lean proofs and preprints can go, and trending items can get boosted.

alpama1mo ago

This is scary. Ai is growing faster than our knowledge. We are not prepared

jacktu1mo ago

zuogl1mo ago

The HTML generation is surprisingly good because the training corpus for markup is cleaner than most programming languages.

OsamaJaber1mo ago

the bottleneck isn't generation, it's verification

richard_chase1mo ago

Dude needs to step down off his pedestal before he gets knocked down.

zkmon1mo ago

sexylinux1mo ago

Unfortunately it still does create errors.

This is of enormous importance but still is being actively ignored by many professionals or dismissed as as a minor issue.

The marketing guy was right and marketing is now dominating science, but humanity will pay a big price for that.

Putting enormous amounts of money into a fundamentally flawed system that we can not optimize to produce reliably error free results is just stupid.

LLMs introduce a level of uncertainty and unreliability into computing that makes them practically useless.

Many people have already been killed because decision makers are not able to follow that very simple logic.

j / k navigate · click thread line to collapse