GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2 (opens in new tab)

(arrowtsx.dev)

542 pointsoshrimpton7d ago273 comments

273 comments

143 comments · 38 top-level

wolttam6d ago· 24 in thread

> it is clear that actual intelligence has plateaued significantly.

> Moving forward, the industry cannot continue to train bigger and bigger models since their intelligence not only plateaus but often will get worse

These are wild claims - why are we concluding that bigger models and more data = more hallucination? That’s actually the opposite of what’s been happening over the last couple years. Some models may still hallucinate more but they all hallucinate much less than the original 175B ChatGPT which was smaller and trained on (much) less data than anything current.

Edit: My mention of data comes from this quote:

> A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling

My take on the current situation: it seems clear that the industry has seen that there is still a lot left to squeeze out of sub-1T models. But for that you do need more, high-quality data in the distribution which you want to unlock capabilities for.

an0malous6d ago

> why are we concluding that bigger models and more data = more hallucination?

That’s not what your quotes said. They said bigger models = plateau in intelligence, nothing about more data or increased hallucinations

The relevant quote for what you’re talking about would be:

> It’s been proven that when a model is trained on large volumes of highly factual and non-theoretical data, it learns to always have an answer.

So there’s two separate claims: 1) bigger models have plateauing results 2) models trained on larger amounts of factual data have a higher hallucination rate

I’m pretty sure #1 is well known, I think OpenAI’s own research on scaling laws showed diminishing returns on parameter count and training data volume years ago. I don’t know what the support for #2 is besides for the actual post contents.

jmalicki6d ago

I find these internet arguments talking about LLMs as if they are trained by reading the internet to be wild.

Yes, pretraining still exists. But for the past few years, pretraining by reading the internet is just the initial bootstrapping of LLM training. The RL training they get from bespoke training data, with very very different characteristics than what these armchair analyses claim, dominates these days.

5 more replies

themgt6d ago

That’s not what your quotes said. They said bigger models = plateau in intelligence, nothing about more data or increased hallucinations ... I’m pretty sure #1 is well known

Well known in a multiverse branch where Fable was a dud?

1 more reply

coffeefirst6d ago

Yeah #2 may be incidental. Suppose one lab focused on bigger, and another on reinforcement training geared towards factual accuracy over sycophancy. You could easily wind up with a model from the second lab that is less powerful but more accurate.

I can’t prove it but I suspect there’s a bit of that going on.

1 more reply

ifwinterco6d ago

#2 is not that surprising from first principles if the way you made the bigger model was by feeding it poorer quality training data because it’s the only way you can get enough

claytongulick6d ago

My impression is that the fundamental issue is that LLMs attempt to extract reasoning (executive execution) from data (relationship between tokens).

There's an open question about whether this is theoretically possible, but it doesn't seem like it to me.

Human generated data is an effect of reasoning. Attempting to extract executive function from it is kind of like taking an anti-derivative of a function.

This has always seemed like the root of hallucinations to me. It sort of follows the parallels to lossy compression that a lot of people draw. You're extracting some characteristics by observing the relationship between tokens, and then trying to argue that those characteristics are equivalent to the thing that generated the original tokens.

Surely there's some sort of overlap there, but viewed that way, it seems obvious that more and more parameters and scaling won't solve the fundamental problem. There's only so much meaning you can extract from token relationships.

It's like trying to derive the shape of a flame from the smoke it produces.

The original intelligence that created those tokens was driven by a whole universe of inputs, from hormones to starlight to gravity, not to mention all of the strange things about consciousness and parapsychology that is so poorly understood.

The machines are definitely useful for a certain class of tasks - those that don't require much executive function, and the useful work mostly involves pattern matching.

The problem is, we seem to be mistaking effect for cause and imagining that these things have greater capabilities than they'll ever posess.

The investors that don't understand this are indeed going to learn a bitter lesson.

coldtea5d ago

>The original intelligence that created those tokens was driven by a whole universe of inputs, from hormones to starlight to gravity

Still inputs, that in the end changed something about synapses and their activation. And whether doesn't have a strong enough local effect to be material to the those operations, can be ignored too. E.g. gravity might kill you via a fall or a tide drowning you, but might have zero influence in your thinking at the brain operation level, aside from some influence that can be expressed in weights and such.

1 more reply

bilater6d ago

Yeah not only is it totally unsubstantiated, the benchmarks are getting less useful to really show the difference between these models. Big model smell is still a thing and GLM 5.2 while impressive is not Fable class.

Here is something I would like people to chew on. Perhaps the smartest researchers in the world across multiple labs know more about this than we do? Perhaps they are aware of issues like the data wall and diminishing marginal returns. And perhaps they are being honest when they tell you there is no wall?

coldtea5d ago

>Perhaps the smartest researchers in the world across multiple labs know more about this than we do?

Perhaps the smartest researchers in the world across multiple labs follow the money, and don't make waves that go against them getting their paychecks?

That's part of what makes them smartest.

nathan_compton6d ago

Are the smartest researchers in the world out there saying there isn't a wall? I don't know of any people doing the actual R&D who frequently make outrageous claims.

2 more replies

eurekin6d ago

> A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling

I'm pretty sure it's mostly due to the training data quality. No idea, why this never gets mentioned in those discussions.

It was obvious right from the get go, that the scaling law just enabled some abilities, that were described by the underlying data and allowing the ANN to abstract it in the latent space.

aurareturn6d ago

Aren't hallucinations also heavily influenced by compute and memory capacity? IE. Companies can spend more time to verify results in an agentic format, spend more thinking tokens, and less quantization. All of these heavily depend on compute and memory but are proven to decrease hallucinations.

Maybe GPT 5.5 is heavily nerfed due to lack of compute, memory, and energy?

I agree that it's farfetched to conclude that bigger models have pleateued.

dominotw6d ago

article specifically talks about this. deepseek spending significant test time with worse results than klm

1 more reply

utopiah6d ago

>> it is clear that actual intelligence has plateaued significantly.

> These are wild claims -

Indeed, it is not clear there was any actual intelligence at any point.

A lot of generated content sure, sometimes even useful, but not necessarily anything more.

ozgung6d ago

What is the definition of "actual intelligence"? How does it differ from regular intelligence and non-intelligence?

If someone can "design a custom asyncio event loop policy in that overrides get_child_watcher()", I would call that person intelligent. Does that mean that person is not actually intelligent but a mere content creation machine?

Traditionally if you can create content, this shows you're intelligent. Created content is often called "intellectual" property. If a person can understand complex ideas and make connection between them, that is considered intellectual work. You have to be intelligent to do intellectual work. If a person can solve problems, this is also called intelligence. If the person can solve more complex problems, that person is said to have higher intelligence. This is often measured with a scale called IQ (Intelligence Quotient). There are other types of intelligence but they are basically the variations of the same ability. Most definitions of intelligence also involve an ability to adapt into the environment.

Since intelligence is such a broad concept what exactly is the difference between the actual intelligence and AI, other than one is natural and the other one is artificial?

I understand being anti-AI because of the very real societal concerns. But ignoring what is in front of you is not a solution.

1 more reply

resters6d ago

to train models to be smarter than they are, one needs examples and cases to train on, and once you get close to the top percentiles of human reasoning there is extremely little such material available.

You can create contrived logic problems, but they often turn into language games because English is not formal logic.

And you can train on "monty hall" style problems, but those too are language games that are intriguing to humans but obvious when framed slightly differently.

In other words, model trainers are fighting against the overwhelming mediocrity of the training corpus (all of the recorded human output from history).

As models improve, the next phase will be models co-designed with humans to overcome these limits. The way we use language and the process we use to problem solve (we currently call this "orchestration") will evolve as part of this. Meatspace metaphors map badly when we have massive context and don't need the same limits. How different is hallucination from extrapolation, etc.

Much of the skepticism and confusion about LLMs is no different than a person of average intelligence hearing a highly intelligent person explain something and considering the explanation gibberish, then arrogantly accusing the intelligent person of being unhelpful.

Much like dogs were domesticated from wolves to have traits that make them good around humans, LLMs will evolve around our limits, around our arrogance, around our aesthetic biases and prejudices. Intelligence and rationality is fundamentally not what most humans want from an LLM.

madduci6d ago

Isn't that the case of over fitting? You have more data, but when you ask something that's not in that data, hallucinations happen

goodness4all4d ago

I’m studying the root of probabilities and it’s impossible to have models without probable “hallucinations”. If we hit truth 20 times and 1 miss, we still would not consider it truth. This is the mathematical foundation these models are built & trained upon. Probability is our way of life, yet truth is subjective in life. Why use AI when you could just use a database if humans want determinism? The statistical mirror IS the power in AI.

coldtea6d ago

>These are wild claims - why are we concluding that bigger models and more data = more hallucination?

Because that's what they measured in this case.

blurbleblurble6d ago

How do we know gpt 5.5 is a bigger model

Phelinofist6d ago

Since it was created by _Open_AI surely it's really open and we can check, right? SCNR

harrall6d ago

In cognitive science, it appears your brain has two modes of thinking:

- A very parallel type of computation that is fast and generally accurate and integrates hundreds of variables. It’s sometimes labeled as intuition or system 1 thinking.

- A much slower, step by step, analytical type, commonly linked with your pre-frontal cortex (one of the newest parts of the brain). Sometimes called system 2 thinking.

Maybe the way the universe works is that all computation more or less is one of those two types. In which case, an LLM alone is only the first part, which is often right but its results also cannot ever be proven.

stevemk14ebr6d ago

An LLM is not thinking, assuming and relating it to thought and universal truths is nonsense.

3 more replies

dominotw6d ago

you mixed two random quotes from the article to create a strawman.

ofcourse you knew what you were doing but disappointing that this was top comment.

stalfie6d ago· 22 in thread

One thing I wonder about hallucinations, is that it seems on the surface that it is an easy problem for RLVR to target. Since you're already generating enormous amounts of reasoning traces which are verified by correct answers, just have "don't know" as an option as a valid answer, and on problems where none of the thousands of reasoning traces led to a correct answer, just promote the traces that led to the "don't know" answer as training data. Essentially teaching the model that "I don't know" is a valid answer.

Sam Altman himself had a blog post about this a while ago that seemed to suggest this thought, so I guess it's obvious to everyone. But if that is so I assume it's just not as easy in practice.

wongarsu6d ago

Because nearly all benchmarks measure "accuracy" by giving you a point for a correct answer, and 0 points for everything else. If you have 100 questions you are 10% certain on, answering "I don't know" to all of those leads to 0 points, answering all of them as if you are confident leads to an expected value of 10 points. So that's what most AIs are trained to do

AA-Omniscience is the only AI benchmark I know of where randomly guessing gets you a lower average score than answering all questions with "I don't know"

jampekka6d ago

AA-Omniscience Index gives +100 for correct, 0 for "I don't know" and -100 for incorrect.

For your scenario the confident confident strategy will give average of -90. Saying I dont't know to all will give 0.

A lot of models have negative AA-Omniscience Index.

They also do have AA-Omniscience Accuracy and AA-Omniscience Hallucination Rate that handle "I don't knows" differently.

https://artificialanalysis.ai/evaluations/omniscience

nutjob26d ago

It should be 1 for correct, 0 for don't know and -1 for wrong.

They are much better incentives. In real life a wrong answer is much more damaging than a don't know.

6 more replies

macleginn6d ago

The main problem here is that hallucination suppression doesn’t generalise. We can penalise models for incorrect answers on a wide range of questions, but this doesn’t lead to the emergence of a coherent worldview, which, coupled with logical abilities, is the only true remedy against hallucinations. With current architectures, hallucinations will likely persist on open-domain tasks forever.

embedding-shape6d ago

> We can penalise models for incorrect answers on a wide range of questions, but this doesn’t lead to the emergence of a coherent worldview, which, coupled with logical abilities, is the only true remedy against hallucinations

I don't think anyone is trying to add "a coherent worldview" by reducing hallucinations, not sure how that even realistically could be aim.

What people want, is for the models to stop giving confident answers that are clearly incorrect. Yes, it won't lead to "a coherent worldview", but it'll at least stop wasting people's time if the model said "You know what, what you said doesn't make sense / isn't clear, is what you mean .... ?" or even "I'm not sure" or "I don't know".

Currently, if you have the wrong starting point, ask the model to do something, they more often than not just go ahead and do that, misunderstandings or not. They seem optimized to never push back, unless you prompt for that, and most seem to favor "I'm just gonna assume X" rather than taking a step back and figuring out how to not assume. Again, unless you prompt against that behaviour/steering it into a different workflow.

1 more reply

smallerize6d ago

I think the trouble is in the outputs of the LLM and how it's interpreted by the tooling. The output is a distribution of probabilities of all possible next tokens. Even if the probability of every token is very low, the output gets normalized so that the sum of all probabilities is 1. So after that step, it's hard to see if the model was strongly preferring certain tokens or if you're just looking at amplified noise.

Training an extra "don't know" token means you have to build a moat between every other token. Between "yes" and "no", you don't have a muddled noisy area where both "yes" and "no" have relatively high probabilities, you need a new peak where "don't know" is higher. Then you just have new muddled areas between "yes" and "don't know", and "don't know" and "no". That requires even more finesse to train another answer in between.

Instead, you could check whether multiple options are about equally likely. But then you have to check if they are actually synonyms, like are the top two choices "Genève" and "Geneva", which is a good sign that the model knows the answer? Or are the top two "yes" and "no"?

omneity6d ago

It’s not as simple. I trained an LLM before on exactly this, to scratch the itch of this question.

The task was simple, using the MS-MARCO[0] dataset which contains queries, search results, answers, I made a training set that has:

1. Questions paired with real results supporting them (mixed with some irrelevant results), and a correct answer

2. Questions paired only with irrelevant results, with the answer “No answer present”

The dataset was huge (close to 1M samples), and I trained using different techniques, from SFT (just mimicking the dataset) to DPO (good answer contrasted with a bad answer for the same user query) to GRPO (verifier that checks my annotations whether an answer was present or not)

Lo and behold, this didn’t reduce hallucination, rather made it much worse. Now the model started claiming “No answer present” even when it is, or even when the question didn’t need search results in the first place (simple stuff like what is X+Y).

Now you could argue that my training was basic compared to what frontier labs could do. Yet I think it hints at a more profound limitation. LLMs are finicky and don’t have a neat understand of things from first principles (list of search results, check relevance of result to user query, if answers are below a certain threshold of relevance then don’t consider them to answer …).

tl;dr: not as simple as one might think, perhaps not attainable at all.

0: https://huggingface.co/datasets/microsoft/ms_marco

jonathanhefner6d ago

Thank you for sharing! Based on your experience, do you think a two-model system might fare better? For example, two models in serial where the second model is trained to "sniff out" potential hallucinations and fact check them (and possibly iterate with the first model)?

1 more reply

maxbond6d ago

If you could write that reward function you wouldn't need an LLM, you'd just query the reward function to answer any question. You can create a benchmark and check that automatically, but you can't solve this in the general case. The model can do well on the benchmark but still give overconfident answers in areas the benchmark doesn't cover.

You can definitely tune a model to say "I don't know" more often but it will cost you performance, the model will reject some questions that it could answer meaningfully. In the degenerate case the model could collapse predicting that sequence always or almost always.

stalfie6d ago

I guess so. Just to be clear, I was talking about post-training methods for reasoning models here, not pre-training. I think "model as a judge" should actually do okay as a "sentiment analysis" style reward for expressing uncertainty. So if none of the thousands of reasoning traces you generate reach the validated answer, you run a judge to rate uncertainty and put those reasoning traces back into the training pool.

But I guess my logic breaks down here a bit, because if there is such a thing as a validated answer, then the correct answer is in fact never uncertainty. The correct answer is to continue post training until the model gets it right. So perhaps the real answer is to create RLVR tasks where the valid answer is "I don't know" and nothing else, like this benchmark does. Or maybe that doesn't work either, no matter how many you create.

I feel as though there is some kind of philosophical lesson to be had from how hard hallucinations are to get rid of. Maybe, similarly to humans, successful models are often "arrogant" in a sense. Perhaps you just never solve an Erdös problem without some degree of self deception that it's possible for you to do so. In this line of thinking, greatness in humans is actually not related to humility, but just being so good that you actually get things right when you try. Expressing humility is of course something great people tend to do, but I'm referring to what happens under the hood.

If you squint a bit, that's kinda the trend with models. The useful ones are not that much less likely to hallucinate, they are just good enough that they tend to get it right. This comparison is of course probably not even remotely correct, but at least it's fun to anthropomorphize a bit.

roenxi6d ago

If we had a theoretical technique to identify the true and objective reality we'd use it in the courts and laboritories. There is no such technique, but what we do have is 2 techniques that seem work:

1) Has a certain standard of evidence been met?

2) Are the related arguments free of logical inconsistencies?

We can train the LLMs to do 2, and maybe even 1 to some extent (exactly what quality of evidence a computer can practically gather is limited). But that isn't going to get rid of hallucinations, for the same reason courts are hit-and-miss or the conclusions of studies often aren't very reliable. These techniques help, but sometimes they still get people to say things that, on close inspection, turn out to be nonsense. And those best-effort approaches are too much to expect for most questions an LLM will be handed which are informal, low stakes and don't need strong supporting evidence or logical rigour.

I think it is underestimated how many LLM-style hallucinations people themselves have. It just isn't obvious because most humans have a strategy of only repeating what the herd says after it has been socially vetted, which makes their individual eccentricities less obvious.

TLDR; I don't think it looks like an easy problem for RLVR, it looks technically unsolvable. Even making progress requires a philosophical breakthrough on the nature of truth so that the objective function can be established.

stalfie6d ago

Well, I'd argue that this depends on the field you're investigating. Sometimes you have a way to identify objective reality and sometimes you don't. In mathematics the majority of the field is verifiable in this way. Coding a bit less as it's intersubjective, as and the ideal methodology is subject to taste.

But even in muddy fields of reality like medicine, there are objective facts to be found. When someone comes into an ER with chest pain, you often find a true, undeniable reason for why that is happening. If their lung has collapsed, a coronary artery is clogged or the aortic artery is dissecting, even if you don't find that out it tends to be clear in retrospect. The area of reality that becomes muddy is when use proxy signals to try to figure out who gets promoted to expensive/harmful examinations we can make final conclusions from, or the cases that don't fit cleanly into one bucket or the other. But very often, the gold standard truly is golden.

Of course, many realms of reality cannot be verified in this way. But I'd argue that there are quite a few that can.

1 more reply

amelius6d ago

But if an LLM says "I don't know" should you pay for the tokens?

guerrilla6d ago

Why not? It did the work. Why should you expect it to be omniscient?

We can rank them based on how much they know and people will gravitate towards those that do know more.

It's a market after all.

1 more reply

skillina6d ago

"I don't know" has positive value, presumably you could prompt further to learn more about where it got stuck. It also increases the value of correct answers, by improving confidence that answers are actually correct.

"Confidently incorrect" has negative value. At best, a human realizes the answer is wrong and At worst, the incorrect information makes is not identified and can cause untold damage. By having the potential to be so severely wrong, it lessens the value of correct answers because there is a lower confidence value on their output.

embedding-shape6d ago

Depends on what your understanding of the product is.

If someone sold you a "Solved all your problems" machine, and it suddenly doesn't solve all your problems, then probably no, you shouldn't pay.

But the way I'm being sold LLMs, is basically "A text generator that gives your plausible-sounding human text that sometimes hallucinates and gets things wrong, based on your input", then regardless of what the outcome is, I still made use of the "Input > Output" part, which is what I bought into, so I should still pay for that.

Now of course bunch of people will say they been sold the former, but the companies themselves seem to be selling the latter. That's my perspective from a person who doesn't follow "influencers" and what not though, which seem to be selling the public on the former rather than the latter.

3 more replies

nutjob26d ago

'I don't know' is the correct answer for infinitley more questions than those that can be answered.

ludwik6d ago

I would be very willing to pay more! The choice between “you may get a correct answer, or you may get lied to, without a clear way to distinguish between the two” and “you may get a correct answer, or a clear indication that the answer was not found” is pretty clear. One is a much more useful tool than the other. I don’t see any real incentives for companies making LLMs to keep their AI factually unreliable. (Full disclosure: I work for one, but I’m definitely not in the rooms where such decisions would be made.)

maxbond6d ago

Would you rather pay for a nonsensical explanation?

cyanydeez6d ago

the problem is the null answer will stop the "markov" chain.

so, thats all.

BDPW6d ago

You dont have to literally send a null token. Train it to generate text that summarizes the evidence that is there but the uncertainty of the final answer to a prompt.

make36d ago

Transformers are not Markovian, their whole point is arguably to be the reverse of Markovian, to efficiently make it so the new tokens are a function of all previous tokens

aesthesia7d ago· 20 in thread

Hallucination rate scores are a little tricky to interpret because they're conditional on the model not knowing the answer. That means they don't measure the probability of your encountering a hallucination in everyday use, since that also depends on the probability of the model not knowing the answer, as well as how well your distribution of tasks aligns with the distribution tested in the eval.

I'd also hesitate to attribute this difference in hallucination rates purely to model size. Yes, GLM-5.2 hallucinates much less frequently than DeepSeek-V4 Pro with twice as many parameters, but DeepSeek-V4 Flash is less than half the size of GLM-5.2 and tops the AA-Omniscience hallucination index. Opus 4.8, which is likely larger than DeepSeek-V4 Pro, has a 36% hallucination rate on the index, above GLM-5.2's 28%, but way below the DeepSeek numbers. Opus also has a 47% accuracy rate vs GLM-5.2's 25%. If you use these numbers to calculate the absolute hallucination rate (i.e., the number of hallucinated responses divided by the total number of responses), you get 19% for Opus and 21% for GLM-5.2.

So yes, all else equal larger models may be more prone to hallucination in scenarios where they don't know the answer, but there are a lot of other factors that affect hallucination rates, and it's not totally clear that this is the main metric that's worth tracking.

ComputerGuru6d ago

I’m not disagreeing with you but at the same time, models don’t “know” anything in that binary sense. I’m not trying to get in the woods here, I genuinely mean that what you pass off as a simple explanation is actually incredibly nuanced. A fact appeared once in training data , a fact never appeared in the training data, a fact appeared ten times, a fact appeared a thousand times. Which does the model know? Facts aren’t stored as-is, they’re all broken down into their components and compressed in the weights. “Similar” facts that didn’t appear an overwhelming number of times get bundled together and eventually conflated. But then what is a similar fact? Which facts were entirely ablated vs which were bundled together with others effectively poisoning the pool but also giving it inference strength? The model doesn’t know anything and can never know what it knows or doesn’t know.

unshavedyak6d ago

I often wonder how humans "know" things. I suspect (ignorant armchair) we have some ability to signal strength of those facts, via repetition. Without this layer of introspection i imagine LLMs can never avoid hallucination.

It obviously breaks down with humans too, given we so easily hallucinate and confuse things we "know". However i still suspect we're more reliable at probing information we've experienced vs not. Even if the case of poisoned knowledge, eg a crime scene accidentally implying information to a witness that the witness doesn't actually know, we still "know" that poisoned information via incorrect inference. Ie we "experienced" it.

Wonder what architecture would allow for this style of information/weight probing for an LLM.

1 more reply

in-silico7d ago

Additionally, maybe it's easier for a model to realize that it doesn't know the answer when the question is easier.

If Opus gets all but the hardest questions right, it might have a higher hallucination rate because the questions it gets wrong are the questions where verification or hallucination detection are the most difficult

sudosysgen7d ago

This is missing a common failure mode, which is information past the knowledge cutoff. If you need info past that time they'll fail no matter how big or small the model is, so the hallucination rate can matter independently of the knowledge base. If all use-cases had a uniform risk of falling out of support, this would be a valid argument, but since it's often the case that a datapoint is guaranteed to fall out of support, the absolute ability to recognize that is crucial.

reinitctxoffset7d ago

Hallucination should be called "failure to ground".

Something about the cost model of US near frontier has the cattle prod out whenever a model is uncertain but thrashes on whether to search. Search flinch is roughly all hallucination.

I don't even wait for the model's turn, if there's a man page or Hoogle hit, stuff the last prefix cache cut point. You come out ahead.

gymbeaux7d ago

Those numbers are abysmal. Should we really be using LLMs to write our code? I have a theory- LLMs can spit out code that gets the job done and looks ok, maybe even great, but contains small “anomalies” that compound over time. An enterprise app developed entirely with LLM-happy devs might end up virtually unmaintainable.

I’m not sure how to explain it, but the more I see LLM-written code the more I feel it’s bad code doing a good job of masquerading as good code. I think this take will become less-hot in the next year or two when we see enterprise greenfield projects that were created entirely with LLM “assistance” go to prod. I think we’ll find that the code is difficult for humans to read, understand, debug, and extend- and I think the larger the codebase the harder it will be for LLMs to maintain. More opportunity for hallucination, larger context windows needed, more tokens bought and spent for smaller and smaller code changes. I think the more code an LLM writes for an app, the worse that codebase becomes.

andybak6d ago

I can't help but feel that people continually underestimate how bad human written code becomes over time. The exception is probably single-person passion projects or open source projects that maintain quality governance over time.

I strongly suspect most closed source code developed under commercial or internal pressure is pretty awful after a few years of development.

All LLM code has to do is suck less than existing code. And that's presuming the code quality doesn't improve as the models, the harnesses and our ways of working with them improve.

5 more replies

xvinci6d ago

Not my observation. If you never look at the code and dont have basic guardrails in place (linters, architecture tests, some guidelines for best practices) - probably.

But as soon as you do minimal reviews and high-level corrections, applications turn out just fine.

Can there be bugs? Sure. That's the price of not reading or understanding every line. It should depend on the criticality of your software how much of these you tolerate and how much you don't (reviewing, understanding, testing everything 100% like you were used to if you had written it yourself will kill most if not all of your gained speed)

But I never got the impression of unmaintainability or unfixable bugs.

Actually the other side around: A really good cleanup pass, architectural changes, or bugfixes are seldom more than a few prompts and 2 hours away, provided your overall base is decent and you actually gave a fuck from the start.

3 more replies

realaleris1496d ago

Take a look at a sufficiently old random internal repo which was not written with LLMs and compare.

My observation is that they are equally bad and hard to maintain or even more so than the new ones.

One thing I’ve noticed is that the LLM assisted ones have a lot more comments which is nice but take more time to read.

2 more replies

csomar6d ago

Easy fix: Code's basically free now, so just pipe your errors straight into an LLM and get instant patches. Sure, the patches themselves are broken too, but no worries! just pipe those back in again. Code's disposable now, fresh code generated on every request.

On a more serious note, I think the problem will be the inability to handle/maintain the systems once they are too big and nobody has no idea what's inside of them or what they do.

1 more reply

realusername7d ago

> code that gets the job done and looks ok, maybe even great, but contains small “anomalies” that compound over time

They clearly are only assistants for the moment, you can use them to do work ... but only if you could do the said work yourself alone in the first place.

1 more reply

coldtea5d ago

>I have a theory- LLMs can spit out code that gets the job done and looks ok, maybe even great, but contains small “anomalies” that compound over time. An enterprise app developed entirely with LLM-happy devs might end up virtually unmaintainable.

For most enterprise apps, being "unmaintainable" would be an improvement.

rienbdj6d ago

I have a theory that LLM generated code in a highly modular style (simple data, pure functions) will be easier to “recover” by a human team when the LLM gets muddled. So Haskell, basically.

Foobar85687d ago

Have you worked with enterprise apps? The ones I have used for decades are hot garbages.

1 more reply

andix6d ago

I guess you can test that on hypotheticals. Ask about things after the knowledge cut off that never happened. Or ask things that are genuinely unsolvable.

grayhatter7d ago

> Hallucination rate scores are a little tricky to interpret because they're conditional on the model not knowing the answer. That means they don't measure the probability of your encountering a hallucination in everyday use, since that also depends on the probability of the model not knowing the answer, as well as how well your distribution of tasks aligns with the distribution tested in the eval.

Do you have a cite for this?

If a human makes up some bullshit lie, I wouldn't accuse them of making it up only if they actually knew the correct answer. If you don't know, the only correct answer is I don't know. Any other answer is made up bullshit. Why is it only a hallucination if and only if the LLM contains the answer? If you make something up it's still wrong. It shouldn't matter if you could give the correct answer. You didn't, and instead invented some bullshit instead?

Follow up question, how can I apply this rule set to the next test I have to take? I'd love to be able to use "I didn't know" as the excuse for why I made something up.

edit:

> and it's not totally clear that this is the main metric that's worth tracking.

I don't know, the rate at which some model is willing to make up something feels useful. If the argument I see repeated on HN so much is that it's impossible to completely get rid of hallucinations; being able to choose a model that's less likely to invent some lie seems like a positive trait, no?

Either way, I'm happy to agree that a restrictive definition, where a lie doesn't count as a hallucination iff the model doesn't know the answer feels strictly, infinitely less useful than an exact error rate. What percentage of emitted tokens are misleading would be useful for me. Anyone know any group that's attempted to quantify the global error rate?

aesthesia7d ago

This isn't quite the point. When comparing two different models' hallucination rates, the denominator is different. The evaluation works more or less like this: for each question, the model has the option to answer or abstain, so there are three possible outcomes: the model answers and gets it right, the model answers and gets it wrong (hallucination), or the model abstains. The hallucination rate is (model answers wrong) / (model answers wrong or abstains). So if a model A has 50 correct answers, 20 incorrect answers, and 30 abstentions, its hallucination rate is 40%, while a model with 20 correct answers, 20 incorrect answers, and 60 abstentions has a hallucination rate of 25%, even though it hallucinated exactly the same number of times. This is why hallucination rate is incomplete as a metric: it says nothing about the accuracy rate.

1 more reply

jpalomaki6d ago

As human I also give wrong answers if if I know the right one. Sometimes I also give answers even when I don’t really know them.

When pushed, I then start thinking and realise my mistake. System 1 vs 2?

2 more replies

luuundonjk6d ago

there is a difference between a human knowingly bullshitting and being confident because he misremembers something

1 more reply

sgc7d ago

Since models just output the the most probable tokens and you can never accuse them of doing anything other than making it all up, I would like to see these tests run with a prompt that attempts to mitigate hallucination and finishes with something like: "Telling me that you don't have the relevant information or that the task is impossible is extremely useful to me and a valid answer", and see how much that changes the scoring - as well as the usefulness of the answers. There are so many skills like context7 that can be tweaked to improve these results as well.

In other words, you shouldn't choose the model that hallucinates the least without detailed prompting, since a well-crafted agents.md clause should go a long way to improving output, and almost certainly the top scoring order will be different. To the point that I don't find this type of raw comparison useful beyond maybe 'make sure you test that one with more explicit prompts'.

2 more replies

spwa46d ago· 7 in thread

Why is everyone expecting LLMs to be like the Star Trek computer? I wonder if anyone's ever measured what the hallucination rate of a human is.

flexagoon6d ago

Because AI company executives and devoted vibecoders constantly make egregious claims like "programming is fully solved" and even straight up "hallucinations don't exist on frontier models"

verdverm6d ago

We don't have to listen to these people and can form our own perspectives. Following bad leaders is something to avoid

1 more reply

__natty__6d ago

Because this is how LinkedIn “specialists” promotes LLM. The same specialists shouting about crypto a few years ago, then specialists about nft and now about how coding, architecture, accounting, law, medicine and basically every white collar job is solved and you just need enough money to pay for Opus/GPT.

master-lincoln6d ago

Yeah it has been looked at e.g. in [0]. They separate that from lying, but I think for the LLM context it should be included. To me the difference is humans do not bullshit at the same rate and I can find out over time who tends to bullshit more and exclude that persons info from my pool.

> Why is everyone expecting LLMs to be like the Star Trek computer?

Because they are often marketed as magic AIs, not as mere language models.

[0] https://bpspsychub.onlinelibrary.wiley.com/doi/10.1111/bjso....

glouwbug6d ago

It’s not a lie if everyone collectively believes it

bravetraveler6d ago

Marketing, essentially

oshrimptonOP6d ago

I would be so curious to find a comprehensive benchmark on this, humans do have an unfortunate ahem Dunning-Kruger effect ahem tendency to do this

solid_fuel7d ago· 5 in thread

> It’s been proven that when a model is trained on large volumes of highly factual and non-theoretical data, it learns to always have an answer. DeepSeek V4 Pro (1.6T params, 49B active, 44 AA Intelligence Index score) has a ludicrous 94% hallucination score on the AA-Omniscience benchmark, meaning on questions that it couldn’t figure out, it only stated that it didn’t know around 6% of the time, and the rest it confidently hallucinated an answer. GLM-5.2 scored a 28% hallucination rate, Opus 4.8 was 36%, Fable 5 was 48%, and GPT-5.5 was 86%.

Wow! I already knew from previous research shared here that hallucinations are a fundamental problem for LLMs and likely to be unfixable, just like prompt injection, but I didn't realize the hallucination rates were so bad!

Everyone has been acting like the best models only hallucinate in edge cases, but even the best performing one mentioned here - GLM-5.2 - has a hallucination rate of 28% when it doesn't "know" the answer to something.

That said, I think the title on the blog - "Bigger models are not the way" is probably more fitting and touches on what should be even bigger news. If bigger models and bigger training sets have already stopped producing proportional returns, then it seems likely we are already near the top of the S-curve. That's huge news, considering the valuation of companies like OpenAI and xAI is largely based around the (absurd) idea of ever increasing scaling from these models.

SeriousM6d ago

There is no concept of "knowledge" in LLM as it is on Wikipedia.

The question-tokens define the answer-tokens. That's it. The art relies in clustering the relevant weights together.

baq6d ago

If it were that simple we’d all be talking with sql and yet this isn’t happening.

Circuits which emerge in the layers during training are much more complicated than a simple Bayesian relation.

tempaccount4206d ago

> There is no concept of "knowledge" in LLM as it is on Wikipedia.

There can be, you don't know if the closed source models aren't using something like DeepSeek's Engram.

1 more reply

solid_fuel6d ago

Correct, LLMs are not ontologically capable of “knowing”. That is why I put “know” in quotes.

oshrimptonOP7d ago

Agreed on the title, my bad! But yeah, I've had some truly terrible experiences using these "frontier" models in coding agents especially, where they just fabricate facts about codebases.

taffydavid6d ago· 4 in thread

> For the non technical, this is like asking a delivery driver to drop off packages at 3 houses at the same time without ever stopping the truck.

I'm already hallucinating about how this could work and it involves catapults

m3h6d ago

Or we could simply hallucinate that the packages are there at the three houses.

Hallucinations all the way down...

boofus6d ago

Nobody said the 3 houses needed to be on separate properties. Just throw the 3 packages from the moving truck at the one address where all 3 live.

Being an LLM is easy!

sigmoid106d ago

In the end it's just Boltzmann brains.

https://en.wikipedia.org/wiki/Boltzmann_brain

Lionga6d ago

[flagged]

frankohn6d ago· 3 in thread

I think hallucination rates are not a matter of model size but depends on the training of the model. They have been trained on a huge corpus of material that had overwhelmingly well formed questions and we'll formulated and correct answers. This is typically the case of books where the material is highly curated by experts in the field. In a book you never see a question which admit no answer and the book just reasoning and explaining why and how the question has no answer. Neither you will see a good question and the book explaining candidly it doesn't know the answer , because the way the book material is curated the author will omit discussing the question for which it has no answers.

In addition, I think that during HFRL, the labs has a bias for interesting answers that admit a solution and under represent the "bad" questions that admit no good answer. In addition they probably do less effort to HFRL on questions the model should admit it doesn't know.

As humans we have been trained all our lives, in the real world, to be confronted with questions we don't know the response right away and we learned to very quickly assess that we don't know or that we are not sure about the answer.

Another thing we have and LLM have not is fear. We have an amygdala in our brain, separated from the logic thinking part, that can raise a signal of fear so that we get much more carefully about what we say. On the other LLM has no fear organ like the amygdala and just learn to respond based on the patterns in it's training corpus. It never "fears" looking bad or being fired because it gave a wrong answer so it can merrily give perfectly wrong answers.

So, we see hallucination rates can be improved with training but currently the lab are not optimizing for that because there is an high stake race to get the most intelligent and capable model.

Alternatively I can see creating a separate amygdala-like organ for an LLM and that organ may asynchronously fires signal, based on the user prompt and the LLM thinking trace, to inject into the LLM reasoning a fear signal so that it can steer it's answer to something more safe.

oshrimptonOP6d ago

I'd definitely agree that it isn't directly model size, but there is the fact that a larger model in terms of parameter count needs a large amount of training data to not overfit or underfit. So I think this race to the top of "max training data size" has kind of led to unintentional overfitting, not catastrophically, but enough to trigger this perceived omniscience within the model

leobg6d ago

Skinner would say it is not so much about emotions like fear or greed, but about consequences.

frankohn6d ago

Yes, that's when we are mindful and we see the arise in our mind but we don't directly act out of it but we understand it and reason about our options and the consequences.

However the fear has to arise in the first place, to raise the alert.

wiether6d ago· 3 in thread

Purely anecdotal, but when OpenAI removed Codex-5.3 from the ChatGPT sub and forced me to move to GPT-5.5, the result was far worse than what I was enjoying with Codex.

And, of course, it was burning 10 times more tokens for this output.

fvv6d ago

I have the opposite experience with codex 5.3 I had to use 5.2 to design and 5.3-codex to execute , while 5.4 was a better in both, and 5.5 ( all used xhigh) is even better

oshrimptonOP6d ago

Yeah they are 100% in the wrong for removing the fine tuned codex models. It makes sense why they wouldn't want to allocate so many resources towards fine tuning but still the enshittification of GPT models is real

embedding-shape6d ago

Huh, the fine-tuned "codex" variants always seemed like "quick specific edit" prototypes that weren't meant for real use. They worked OK when you were very specific, but besides that, nowhere close to GPT5.X and the other "real" models.

1 more reply

remix20006d ago· 3 in thread

Calling llm slop "hallucinating" is so counter-productive imo. After all, LLMs are just a variant of markov chains and as such this technology isn't able to discern falsehoods from truths. It's like trying to use a barometer to tell the time.

hit8run6d ago

You are also just a variant of markov chains wired in your brain. So what you complaining about?

remix20006d ago

Well the difference here is that you're overly simplifying complex biology and many other factors whereas llms are in fact actually simple mathematical models. As always, the devil lies in the details. Dismissing intricacies is a useful tool for daydreamers, not so much for engineers.

1 more reply

__natty__6d ago

And often it’s not perfect either. Just because one is true it doesn’t dismiss the other

xlii6d ago· 2 in thread

My anecdotal experience differs (though I hold ground that LLM evaluations are highly subjective and benchmarks are just as useful for LLMs as they are for dating websites users).

GLM 5.2 tends to stray way more than and 5.1. It also hallucinates you things subtly: morphs requirements, makes unfounded conclusions. This output is not something I experienced in any model I seen so far.

In coding it's especially annoying because it steers whole request. E.g. I give instruction: "make we a Rust-WASM-Canvas app" and GLM 5.2 goes like "Oh user surely doesn't mean that. I'll better build Dioxus app instead".

LaurensBER6d ago

GLM 5.2 is great but it heavily detoriates once the context window gets past 200k tokens.

I've had more success with creating a plan first and then implementing it in (short-lived) sub-agents.

Ironically good software architecture patterns (small functions, single responsibility) heavily impact the performance of these models as well. They do surprisingly well in well architectured codebases.

They do very poorly in anything that's a mess where Opus and GPT 5.5 still get reasonable performance.

oshrimptonOP6d ago

Yeah the benchmark for sure isn't perfect and without super rigid prompting it is far too easy for it to get off course. 28% hallucination rate isn't nothing either

giancarlostoro6d ago· 2 in thread

I wonder if this is what a “Minimally Viable LLM” looks like. I often wonder how much of an LLM do you need before you can just shove a bigger context Window and any dynamic knowledge content to it like a PDF or markdown file to give it knowledge outside of its training data. I feel like LLMs don’t need more data they just need to be refined.

x3cca6d ago

You might be interested in this model. It's a densely trained on math whuch let's it punch way higher than it should https://github.com/WeiboAI/VibeThinker

giancarlostoro5d ago

Cant open the link without an account is it private or is that just GitHub being annoying?

gcanyon6d ago· 2 in thread

> it is clear that actual intelligence has plateaued significantly

N=1, but I disagree strongly. I'm writing a hard-science science fiction story, and the physics of it is at (and frankly, beyond) my skillset. The story's plot has had to change over a dozen times as I realized errors in my application of physics in the story.

Throughout, I've been reviewing the physics with LLMs, mainly Gemini 3.1 Pro Preview, but also with Claude and OpenAI. Often I have the LLMs debate each other -- "My friend [another model] said XYZ about the physics, is that right or wrong?" In almost all cases, Gemini explains why the other models are wrong, and when I send its explanation to them, they concede it is right and they are wrong.

As I said, I did the above checks literally dozens of times as I wrote the story. And everything was dialed in: no further issues claimed by anyone, me or the LLMs.

Not with Fable. I managed to get it to review the story while it was running, and it listed out something like ten issues: some minor, some general knowledge-based, and two that were impressive:

1. It pointed out where Gemini (and I, and other LLMs) had missed a , resulting in values about 152 times larger than they should have been. I sent that to Gemini and it fully conceded that it had been wrong all along. 2. It pointed out a simple inconsistency in the application of special relativity (I thought I had that at least dialed in, but no :-/ ) that affected a very specific plot point. The story is novella-length, about 28,000 words long, and this is a point that was mentioned in the first two pages, and then not again until the very last page. And it's obvious, once you realize it. And I missed it. Gemini missed it. Claude and ChatGPT missed it.

Only Fable found it. Again, N=1, but that was a remarkable run I got out of it in the couple days it was available.

Bolwin6d ago

Hah, I noticed the same thing writing fiction with fable. Most models seem to go into a sort of "storytelling mode" where they forget their PhD level smarts. I had a character who is doing repair on a satellite. Most models would give you a half-baked explanation with some technical terms - half of them right half of them wrong.

Fable gave a description so deep that even I couldn't figure out what was going on and had to ask it to give me a simpler explanation.

gcanyon6d ago

Nice to hear N=2. I'm really hoping Fable comes back soon.

In my case two people are making very-near-light-speed trips to a star 20-ish light years away. Originally, I had one leaving a month earlier and making the journey with a Lorentz factor of 40, while the protagonist takes the same trip at > 200.

The former experiences a trip of 6 months, the latter something like 25 days. And I wrote it as if that meant that the protagonist would get there months ahead. But both of them will take hours to a day over the time light takes, and the one who leaves a month earlier will get almost a month before.

That error sat in my manuscript for two months of back and forth with other models. Fable found it on the first go.

LMK if you want to trade manuscripts!

metalspot6d ago· 2 in thread

hallucination is good for tasks that have an external oracle like computer programming

dgellow6d ago

Could you explain what you mean? That feels like a waste of processing to me. Yes the model will correct itself once it eventually run a compiler/linter. But that's still wasted time and compute

sometimelurker6d ago

ehh ur right but there's a lot of nuance here. if you have a system that doesn't hallucinate a ton and is still very "creative" that's great, and probably much better than a hallucinating system regardless of its creativity. I'm reminded of theoremproving LLMs working in lean producing millions of slop proofs until one works, but if you have something like that simple RLVR should fix it (external oracle can be the judge for the RL.

stcg5d ago· 1 in thread

In the referenced benchmark GLM-5.2 (max) got 25% of all questions correct. GPT-5.5 (xhigh) got 57% correct.

https://artificialanalysis.ai/evaluations/omniscience

I'd much rather have some answer that I can verify than no answer to verify.

I don't want a model that says "I don't know", because I will verify the answer anyway.

bwfan1235d ago

> I don't want a model that says "I don't know", because I will verify the answer anyway.

Few people actually review answers or code. Because they have been sold the myth that these models can do it all. The main problem is that LLMs dont have causal models, and as a result, their reasoning is a high probability word salad and not a logically sound argument. Particularly on tricky corner cases which it hasnt encountered. I would still agree with you that sometimes hallucinations are actually useful as it provides a strawman, and having even a hallucinated answer to spar with is better than a "dont know".

andai6d ago· 1 in thread

This implies that bigger models are more likely to hallucinate? That doesn't match my experience.

npilk6d ago

I think it implies they are more likely to hallucinate if they don't know the answer. So a big model will return the correct answer more often than a small one, but in the cases where it doesn't, it will be more likely to make something up instead of saying "I don't know".

aubanel6d ago· 1 in thread

> Bigger is not better

The article uses the example of GLM being smaller than DeepSeek, yet better on hallucinations as "smaller can be good too"

But the GLM family itself is scaling up fast: GLM-5.x family is 754B, double the previous generation of GLM-4.x

> comes within just 4 points of GPT-5.5 and 9 points of Fable 5

9 percentage points IS a big difference

CuriouslyC6d ago

If we're hand waiving how an open source model from a Chinese lab that you can use a nearly unlimited amount for <100/mo's 9% difference from the premier, unavailable, expensive when it was available American frontier model, we've already lost.

EbNar6d ago· 1 in thread

The fact that a huge amount uf parameters may lead to worse hallucinations is something I didn't think of. Would this somewhat imply that DeepSeek V4 flash should be less prone tho these issues?

oshrimptonOP6d ago

Surprisingly not! It is the biggest hallucinator on the AA Omniscience Index just 2pp away from V4 Pro. I think this is partially due to the fact that Flash was trained on >32T tokens just like Pro deapite being almost 10x smaller - it seems somewhat likely it was overfit.

EbNar6d ago· 1 in thread

The fact that a huge amount uf parameters may lead to worse hallucinations is something I didn't think of. Would this somewhat imply that DeepSeek V4 flash should be less prone to these issues?

verdverm6d ago

small models cannot encode so many facts, they will hallucinate more out-of-box

a key method to help with hallucinations is to provide good sources when asking questions (context engineering / knowledge base)

nextaccountic7d ago· 1 in thread

>GPT-5.5 and DeepSeek V4 Pro are two of the clearest hallucination leaders, despite being absolutely huge. Because of their immense size they simply did not learn how to say “I don’t know” or recognize intricate logical and technical fallacies. While it is true that a multi-trillion parameter model will always beat a lightweight consumer model on paper (today at least), the commoditization of these huge models is blurring the line between benchmark performance and actual real-world truthfulness and accuracy.

What about using two models, with a smaller model used for this kind of negative reasoning?

bastawhiz7d ago

Now you need a third model to decide if the two other models disagree

1 more reply

hereme8886d ago

Artificial Analysis says GPT-5.5 xhigh scores highest on AA-Omniscience accuracy. The article focuses on rate instead of overall accuracy. Those are different things: a model can answer more questions correctly overall while still being worse at abstaining when wrong.

Curiously, this post and article is the only submission and interaction the OP has made, and these claims support the product he's intending to release.

hyperpape6d ago

> A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling. The limits of this paradigm were put on the world’s stage when Claude Fable 5 was restricted by the US government just three days after its release, marking the first US AI ban stemming from national security. One of the biggest models in the world was banned because a single jailbreak was too much of a risk.

Such a weird thing to start with. The legal status of Fable does not mean that it's not intelligent. If anything, the problem is the opposite, someone thinks it's too intelligent (and/or that Anthropic wouldn't share its last gen intelligent models on the terms the government demanded).

nghnam5d ago

I’d be careful about reading too much into these numbers. The test only looks at cases where the model doesn’t know the answer, so it doesn’t show how often users will actually see hallucinations.

chazeon6d ago

GPT-5.5 must have serious issues; it is fast, but quality-wise, it is just not good. It read one LaTeX paper (which is not long) and can spell my name wrong. This is GPT-5.5-high.

cwillu7d ago

Please don't editorialize titles unless the original title is misleading.

czk6d ago

if you're benchmaxxing then maybe bigger doesnt always mean better, but for general intelligence and big model smell, that couldn't be further from the truth

the oss models are impressive but it's pretty clear how quickly they fall off when you try to use them outside of a narrow set of problems they benchmarked well on when compared to opus/5.5

orbital-decay6d ago

DS v4 is an undertrained snapshot, which is mentioned in their model card. The full version is supposed to be released later and have multimodal input. That said, hallucination rate likely depends on the training policy and different optimization tradeoffs a lot more than on the scale.

raincole6d ago

> meaning on questions that it couldn’t figure out, it only stated that it didn’t know around 6% of the time, and the rest it confidently hallucinated an answer.

From how they measure it, a model that simply answers "I don't know." to any prompt would be the one hallucinates the least. So it's not surprising at all that a smaller model can perform better.

stevenhubertron6d ago

The more I have been using 5.2 the more I have been impressed with it. And I’ve just been using the usually neutered ollama version.

zuzuen_16d ago

I think we need better classification and taxonomy on erroneous LLM behaviors than the catch-all term "hallucinate"..

gitaarik6d ago

This reminds me of the Missing Dollar Riddle [1], where the listener is deliberately put on a wrong thinking path, to fool it.

With your own logical thinking you might never come to this confusion, and if you never heard this riddle before, you might be tricked by it.

But as we grow in life, and get experience, we learn about these riddles and aren't fooled as easily anymore.

Maybe it'll work like that for LLMs too?

[1]: https://en.wikipedia.org/wiki/Missing_dollar_riddle

brown_munda6d ago

GLM 5.2 is really impressive at design as well. Overall loving it.

anArbitraryOne6d ago

It's fine if it hallucinates, as long as it sounds overconfident

ecommerceguy6d ago

It's very much looking like OpenAI will be bailed out, along with all the other Capex'ers. I say this because the trump admin (I feel partially at fault because I voted for him) has indicated they will be bailing out the entire ai stargate from intel and amd to amazon and anthropic. I know alot of everyday folks that absolutely hate - passionately HATE - anything and everything tech bro. Downvote all you want, that's the reality. They see Palantir et al as evil and demonic.

nathan_compton6d ago

Synthesizing a bunch of stuff I've read here lately, it seems like if OpenAI and Claude have actually found product market fit (generating code) then the question of hallucination is going to get less attention in the future. If the real money is in code generation (where there is a relatively clear acceptance criteria of at least "it runs and does what I wanted as far as I can tell") then there doesn't seem to be a lot of juice in pulling ones hair out on hallucination of facts.

It seems like for agentic coding, just making sure the AI can find the relevant documentation to establish a ground truth is probably sufficient.

Note that I'm distinguishing here between hallucination of what you might call "free facts" and hallucination of material which deviates from what is in the context itself. The latter seems both a tractable problem and one which will improve coding agent functionality. But the former seems like its no longer on the critical path, probably because its hard.

dgellow6d ago

> One of the biggest models in the world was banned because a single jailbreak was too much of a risk.

We really don't know what the actual reason is given the politics at play. I would bet more on the Trump administration looking for any excuse to punish Anthropic

Naveja6d ago

loving glm 5.2 personally

metalman5d ago

to paraphrase the title, "in the land of the insane, those who are meerly delusional will rule"

abracadobre6d ago

This is where I asked GPT 5.5

"they say u hallucinate 3x more than GLM 5.2, whats your comeback to this? do i need to dump u? $article"

j / k navigate · click thread line to collapse

273 comments

143 comments · 38 top-level

wolttam6d ago· 24 in thread

> it is clear that actual intelligence has plateaued significantly.

> Moving forward, the industry cannot continue to train bigger and bigger models since their intelligence not only plateaus but often will get worse

Edit: My mention of data comes from this quote:

> A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling

an0malous6d ago

> why are we concluding that bigger models and more data = more hallucination?

That’s not what your quotes said. They said bigger models = plateau in intelligence, nothing about more data or increased hallucinations

The relevant quote for what you’re talking about would be:

> It’s been proven that when a model is trained on large volumes of highly factual and non-theoretical data, it learns to always have an answer.

So there’s two separate claims: 1) bigger models have plateauing results 2) models trained on larger amounts of factual data have a higher hallucination rate

jmalicki6d ago

I find these internet arguments talking about LLMs as if they are trained by reading the internet to be wild.

5 more replies

themgt6d ago

That’s not what your quotes said. They said bigger models = plateau in intelligence, nothing about more data or increased hallucinations ... I’m pretty sure #1 is well known

Well known in a multiverse branch where Fable was a dud?

1 more reply

coffeefirst6d ago

I can’t prove it but I suspect there’s a bit of that going on.

1 more reply

ifwinterco6d ago

#2 is not that surprising from first principles if the way you made the bigger model was by feeding it poorer quality training data because it’s the only way you can get enough

claytongulick6d ago

My impression is that the fundamental issue is that LLMs attempt to extract reasoning (executive execution) from data (relationship between tokens).

There's an open question about whether this is theoretically possible, but it doesn't seem like it to me.

Human generated data is an effect of reasoning. Attempting to extract executive function from it is kind of like taking an anti-derivative of a function.

It's like trying to derive the shape of a flame from the smoke it produces.

The machines are definitely useful for a certain class of tasks - those that don't require much executive function, and the useful work mostly involves pattern matching.

The problem is, we seem to be mistaking effect for cause and imagining that these things have greater capabilities than they'll ever posess.

The investors that don't understand this are indeed going to learn a bitter lesson.

coldtea5d ago

>The original intelligence that created those tokens was driven by a whole universe of inputs, from hormones to starlight to gravity

1 more reply

bilater6d ago

coldtea5d ago

>Perhaps the smartest researchers in the world across multiple labs know more about this than we do?

Perhaps the smartest researchers in the world across multiple labs follow the money, and don't make waves that go against them getting their paychecks?

That's part of what makes them smartest.

nathan_compton6d ago

Are the smartest researchers in the world out there saying there isn't a wall? I don't know of any people doing the actual R&D who frequently make outrageous claims.

2 more replies

eurekin6d ago

> A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling

I'm pretty sure it's mostly due to the training data quality. No idea, why this never gets mentioned in those discussions.

It was obvious right from the get go, that the scaling law just enabled some abilities, that were described by the underlying data and allowing the ANN to abstract it in the latent space.

aurareturn6d ago

Maybe GPT 5.5 is heavily nerfed due to lack of compute, memory, and energy?

I agree that it's farfetched to conclude that bigger models have pleateued.

dominotw6d ago

article specifically talks about this. deepseek spending significant test time with worse results than klm

1 more reply

utopiah6d ago

>> it is clear that actual intelligence has plateaued significantly.

> These are wild claims -

Indeed, it is not clear there was any actual intelligence at any point.

A lot of generated content sure, sometimes even useful, but not necessarily anything more.

ozgung6d ago

What is the definition of "actual intelligence"? How does it differ from regular intelligence and non-intelligence?

Since intelligence is such a broad concept what exactly is the difference between the actual intelligence and AI, other than one is natural and the other one is artificial?

I understand being anti-AI because of the very real societal concerns. But ignoring what is in front of you is not a solution.

1 more reply

resters6d ago

You can create contrived logic problems, but they often turn into language games because English is not formal logic.

And you can train on "monty hall" style problems, but those too are language games that are intriguing to humans but obvious when framed slightly differently.

In other words, model trainers are fighting against the overwhelming mediocrity of the training corpus (all of the recorded human output from history).

madduci6d ago

Isn't that the case of over fitting? You have more data, but when you ask something that's not in that data, hallucinations happen

goodness4all4d ago

coldtea6d ago

>These are wild claims - why are we concluding that bigger models and more data = more hallucination?

Because that's what they measured in this case.

blurbleblurble6d ago

How do we know gpt 5.5 is a bigger model

Phelinofist6d ago

Since it was created by _Open_AI surely it's really open and we can check, right? SCNR

harrall6d ago

In cognitive science, it appears your brain has two modes of thinking:

- A very parallel type of computation that is fast and generally accurate and integrates hundreds of variables. It’s sometimes labeled as intuition or system 1 thinking.

- A much slower, step by step, analytical type, commonly linked with your pre-frontal cortex (one of the newest parts of the brain). Sometimes called system 2 thinking.

stevemk14ebr6d ago

An LLM is not thinking, assuming and relating it to thought and universal truths is nonsense.

3 more replies

dominotw6d ago

you mixed two random quotes from the article to create a strawman.

ofcourse you knew what you were doing but disappointing that this was top comment.

stalfie6d ago· 22 in thread

Sam Altman himself had a blog post about this a while ago that seemed to suggest this thought, so I guess it's obvious to everyone. But if that is so I assume it's just not as easy in practice.

wongarsu6d ago

AA-Omniscience is the only AI benchmark I know of where randomly guessing gets you a lower average score than answering all questions with "I don't know"

jampekka6d ago

AA-Omniscience Index gives +100 for correct, 0 for "I don't know" and -100 for incorrect.

For your scenario the confident confident strategy will give average of -90. Saying I dont't know to all will give 0.

A lot of models have negative AA-Omniscience Index.

They also do have AA-Omniscience Accuracy and AA-Omniscience Hallucination Rate that handle "I don't knows" differently.

https://artificialanalysis.ai/evaluations/omniscience

nutjob26d ago

It should be 1 for correct, 0 for don't know and -1 for wrong.

They are much better incentives. In real life a wrong answer is much more damaging than a don't know.

6 more replies

macleginn6d ago

embedding-shape6d ago

I don't think anyone is trying to add "a coherent worldview" by reducing hallucinations, not sure how that even realistically could be aim.

1 more reply

smallerize6d ago

omneity6d ago

It’s not as simple. I trained an LLM before on exactly this, to scratch the itch of this question.

The task was simple, using the MS-MARCO[0] dataset which contains queries, search results, answers, I made a training set that has:

1. Questions paired with real results supporting them (mixed with some irrelevant results), and a correct answer

2. Questions paired only with irrelevant results, with the answer “No answer present”

tl;dr: not as simple as one might think, perhaps not attainable at all.

0: https://huggingface.co/datasets/microsoft/ms_marco

jonathanhefner6d ago

1 more reply

maxbond6d ago

stalfie6d ago

roenxi6d ago

1) Has a certain standard of evidence been met?

2) Are the related arguments free of logical inconsistencies?

stalfie6d ago

Of course, many realms of reality cannot be verified in this way. But I'd argue that there are quite a few that can.

1 more reply

amelius6d ago

But if an LLM says "I don't know" should you pay for the tokens?

guerrilla6d ago

Why not? It did the work. Why should you expect it to be omniscient?

We can rank them based on how much they know and people will gravitate towards those that do know more.

It's a market after all.

1 more reply

skillina6d ago

embedding-shape6d ago

Depends on what your understanding of the product is.

If someone sold you a "Solved all your problems" machine, and it suddenly doesn't solve all your problems, then probably no, you shouldn't pay.

3 more replies

nutjob26d ago

'I don't know' is the correct answer for infinitley more questions than those that can be answered.

ludwik6d ago

maxbond6d ago

Would you rather pay for a nonsensical explanation?

cyanydeez6d ago

the problem is the null answer will stop the "markov" chain.

so, thats all.

BDPW6d ago

You dont have to literally send a null token. Train it to generate text that summarizes the evidence that is there but the uncertainty of the final answer to a prompt.

make36d ago

Transformers are not Markovian, their whole point is arguably to be the reverse of Markovian, to efficiently make it so the new tokens are a function of all previous tokens

aesthesia7d ago· 20 in thread

ComputerGuru6d ago

unshavedyak6d ago

Wonder what architecture would allow for this style of information/weight probing for an LLM.

1 more reply

in-silico7d ago

Additionally, maybe it's easier for a model to realize that it doesn't know the answer when the question is easier.

sudosysgen7d ago

reinitctxoffset7d ago

Hallucination should be called "failure to ground".

Something about the cost model of US near frontier has the cattle prod out whenever a model is uncertain but thrashes on whether to search. Search flinch is roughly all hallucination.

I don't even wait for the model's turn, if there's a man page or Hoogle hit, stuff the last prefix cache cut point. You come out ahead.

gymbeaux7d ago

andybak6d ago

I strongly suspect most closed source code developed under commercial or internal pressure is pretty awful after a few years of development.

All LLM code has to do is suck less than existing code. And that's presuming the code quality doesn't improve as the models, the harnesses and our ways of working with them improve.

5 more replies

xvinci6d ago

Not my observation. If you never look at the code and dont have basic guardrails in place (linters, architecture tests, some guidelines for best practices) - probably.

But as soon as you do minimal reviews and high-level corrections, applications turn out just fine.

But I never got the impression of unmaintainability or unfixable bugs.

3 more replies

realaleris1496d ago

Take a look at a sufficiently old random internal repo which was not written with LLMs and compare.

My observation is that they are equally bad and hard to maintain or even more so than the new ones.

One thing I’ve noticed is that the LLM assisted ones have a lot more comments which is nice but take more time to read.

2 more replies

csomar6d ago

On a more serious note, I think the problem will be the inability to handle/maintain the systems once they are too big and nobody has no idea what's inside of them or what they do.

1 more reply

realusername7d ago

> code that gets the job done and looks ok, maybe even great, but contains small “anomalies” that compound over time

They clearly are only assistants for the moment, you can use them to do work ... but only if you could do the said work yourself alone in the first place.

1 more reply

coldtea5d ago

For most enterprise apps, being "unmaintainable" would be an improvement.

rienbdj6d ago

I have a theory that LLM generated code in a highly modular style (simple data, pure functions) will be easier to “recover” by a human team when the LLM gets muddled. So Haskell, basically.

Foobar85687d ago

Have you worked with enterprise apps? The ones I have used for decades are hot garbages.

1 more reply

andix6d ago

I guess you can test that on hypotheticals. Ask about things after the knowledge cut off that never happened. Or ask things that are genuinely unsolvable.

grayhatter7d ago

Do you have a cite for this?

Follow up question, how can I apply this rule set to the next test I have to take? I'd love to be able to use "I didn't know" as the excuse for why I made something up.

edit:

> and it's not totally clear that this is the main metric that's worth tracking.

aesthesia7d ago

1 more reply

jpalomaki6d ago

As human I also give wrong answers if if I know the right one. Sometimes I also give answers even when I don’t really know them.

When pushed, I then start thinking and realise my mistake. System 1 vs 2?

2 more replies

luuundonjk6d ago

there is a difference between a human knowingly bullshitting and being confident because he misremembers something

1 more reply

sgc7d ago

2 more replies

spwa46d ago· 7 in thread

Why is everyone expecting LLMs to be like the Star Trek computer? I wonder if anyone's ever measured what the hallucination rate of a human is.

flexagoon6d ago

Because AI company executives and devoted vibecoders constantly make egregious claims like "programming is fully solved" and even straight up "hallucinations don't exist on frontier models"

verdverm6d ago

We don't have to listen to these people and can form our own perspectives. Following bad leaders is something to avoid

1 more reply

__natty__6d ago

master-lincoln6d ago

> Why is everyone expecting LLMs to be like the Star Trek computer?

Because they are often marketed as magic AIs, not as mere language models.

[0] https://bpspsychub.onlinelibrary.wiley.com/doi/10.1111/bjso....

glouwbug6d ago

It’s not a lie if everyone collectively believes it

bravetraveler6d ago

Marketing, essentially

oshrimptonOP6d ago

I would be so curious to find a comprehensive benchmark on this, humans do have an unfortunate ahem Dunning-Kruger effect ahem tendency to do this

solid_fuel7d ago· 5 in thread

SeriousM6d ago

There is no concept of "knowledge" in LLM as it is on Wikipedia.

The question-tokens define the answer-tokens. That's it. The art relies in clustering the relevant weights together.

baq6d ago

If it were that simple we’d all be talking with sql and yet this isn’t happening.

Circuits which emerge in the layers during training are much more complicated than a simple Bayesian relation.

tempaccount4206d ago

> There is no concept of "knowledge" in LLM as it is on Wikipedia.

There can be, you don't know if the closed source models aren't using something like DeepSeek's Engram.

1 more reply

solid_fuel6d ago

Correct, LLMs are not ontologically capable of “knowing”. That is why I put “know” in quotes.

oshrimptonOP7d ago

Agreed on the title, my bad! But yeah, I've had some truly terrible experiences using these "frontier" models in coding agents especially, where they just fabricate facts about codebases.

taffydavid6d ago· 4 in thread

> For the non technical, this is like asking a delivery driver to drop off packages at 3 houses at the same time without ever stopping the truck.

I'm already hallucinating about how this could work and it involves catapults

m3h6d ago

Or we could simply hallucinate that the packages are there at the three houses.

Hallucinations all the way down...

boofus6d ago

Nobody said the 3 houses needed to be on separate properties. Just throw the 3 packages from the moving truck at the one address where all 3 live.

Being an LLM is easy!

sigmoid106d ago

In the end it's just Boltzmann brains.

https://en.wikipedia.org/wiki/Boltzmann_brain

Lionga6d ago

[flagged]

frankohn6d ago· 3 in thread

So, we see hallucination rates can be improved with training but currently the lab are not optimizing for that because there is an high stake race to get the most intelligent and capable model.

oshrimptonOP6d ago

leobg6d ago

Skinner would say it is not so much about emotions like fear or greed, but about consequences.

frankohn6d ago

Yes, that's when we are mindful and we see the arise in our mind but we don't directly act out of it but we understand it and reason about our options and the consequences.

However the fear has to arise in the first place, to raise the alert.

wiether6d ago· 3 in thread

Purely anecdotal, but when OpenAI removed Codex-5.3 from the ChatGPT sub and forced me to move to GPT-5.5, the result was far worse than what I was enjoying with Codex.

And, of course, it was burning 10 times more tokens for this output.

fvv6d ago

I have the opposite experience with codex 5.3 I had to use 5.2 to design and 5.3-codex to execute , while 5.4 was a better in both, and 5.5 ( all used xhigh) is even better

oshrimptonOP6d ago

embedding-shape6d ago

1 more reply

remix20006d ago· 3 in thread

hit8run6d ago

You are also just a variant of markov chains wired in your brain. So what you complaining about?

remix20006d ago

1 more reply

__natty__6d ago

And often it’s not perfect either. Just because one is true it doesn’t dismiss the other

xlii6d ago· 2 in thread

My anecdotal experience differs (though I hold ground that LLM evaluations are highly subjective and benchmarks are just as useful for LLMs as they are for dating websites users).

LaurensBER6d ago

GLM 5.2 is great but it heavily detoriates once the context window gets past 200k tokens.

I've had more success with creating a plan first and then implementing it in (short-lived) sub-agents.

They do very poorly in anything that's a mess where Opus and GPT 5.5 still get reasonable performance.

oshrimptonOP6d ago

Yeah the benchmark for sure isn't perfect and without super rigid prompting it is far too easy for it to get off course. 28% hallucination rate isn't nothing either

giancarlostoro6d ago· 2 in thread

x3cca6d ago

You might be interested in this model. It's a densely trained on math whuch let's it punch way higher than it should https://github.com/WeiboAI/VibeThinker

giancarlostoro5d ago

Cant open the link without an account is it private or is that just GitHub being annoying?

gcanyon6d ago· 2 in thread

> it is clear that actual intelligence has plateaued significantly

As I said, I did the above checks literally dozens of times as I wrote the story. And everything was dialed in: no further issues claimed by anyone, me or the LLMs.

Not with Fable. I managed to get it to review the story while it was running, and it listed out something like ten issues: some minor, some general knowledge-based, and two that were impressive:

Only Fable found it. Again, N=1, but that was a remarkable run I got out of it in the couple days it was available.

Bolwin6d ago

Fable gave a description so deep that even I couldn't figure out what was going on and had to ask it to give me a simpler explanation.

gcanyon6d ago

Nice to hear N=2. I'm really hoping Fable comes back soon.

That error sat in my manuscript for two months of back and forth with other models. Fable found it on the first go.

LMK if you want to trade manuscripts!

metalspot6d ago· 2 in thread

hallucination is good for tasks that have an external oracle like computer programming

dgellow6d ago

Could you explain what you mean? That feels like a waste of processing to me. Yes the model will correct itself once it eventually run a compiler/linter. But that's still wasted time and compute

sometimelurker6d ago

stcg5d ago· 1 in thread

In the referenced benchmark GLM-5.2 (max) got 25% of all questions correct. GPT-5.5 (xhigh) got 57% correct.

https://artificialanalysis.ai/evaluations/omniscience

I'd much rather have some answer that I can verify than no answer to verify.

I don't want a model that says "I don't know", because I will verify the answer anyway.

bwfan1235d ago

> I don't want a model that says "I don't know", because I will verify the answer anyway.

andai6d ago· 1 in thread

This implies that bigger models are more likely to hallucinate? That doesn't match my experience.

npilk6d ago

aubanel6d ago· 1 in thread

> Bigger is not better

The article uses the example of GLM being smaller than DeepSeek, yet better on hallucinations as "smaller can be good too"

But the GLM family itself is scaling up fast: GLM-5.x family is 754B, double the previous generation of GLM-4.x

> comes within just 4 points of GPT-5.5 and 9 points of Fable 5

9 percentage points IS a big difference

CuriouslyC6d ago

EbNar6d ago· 1 in thread

The fact that a huge amount uf parameters may lead to worse hallucinations is something I didn't think of. Would this somewhat imply that DeepSeek V4 flash should be less prone tho these issues?

oshrimptonOP6d ago

EbNar6d ago· 1 in thread

The fact that a huge amount uf parameters may lead to worse hallucinations is something I didn't think of. Would this somewhat imply that DeepSeek V4 flash should be less prone to these issues?

verdverm6d ago

small models cannot encode so many facts, they will hallucinate more out-of-box

a key method to help with hallucinations is to provide good sources when asking questions (context engineering / knowledge base)

nextaccountic7d ago· 1 in thread

What about using two models, with a smaller model used for this kind of negative reasoning?

bastawhiz7d ago

Now you need a third model to decide if the two other models disagree

1 more reply

hereme8886d ago

Curiously, this post and article is the only submission and interaction the OP has made, and these claims support the product he's intending to release.

hyperpape6d ago

nghnam5d ago

chazeon6d ago

GPT-5.5 must have serious issues; it is fast, but quality-wise, it is just not good. It read one LaTeX paper (which is not long) and can spell my name wrong. This is GPT-5.5-high.

cwillu7d ago

Please don't editorialize titles unless the original title is misleading.

czk6d ago

if you're benchmaxxing then maybe bigger doesnt always mean better, but for general intelligence and big model smell, that couldn't be further from the truth

the oss models are impressive but it's pretty clear how quickly they fall off when you try to use them outside of a narrow set of problems they benchmarked well on when compared to opus/5.5

orbital-decay6d ago

raincole6d ago

> meaning on questions that it couldn’t figure out, it only stated that it didn’t know around 6% of the time, and the rest it confidently hallucinated an answer.

From how they measure it, a model that simply answers "I don't know." to any prompt would be the one hallucinates the least. So it's not surprising at all that a smaller model can perform better.

stevenhubertron6d ago

The more I have been using 5.2 the more I have been impressed with it. And I’ve just been using the usually neutered ollama version.

zuzuen_16d ago

I think we need better classification and taxonomy on erroneous LLM behaviors than the catch-all term "hallucinate"..

gitaarik6d ago

This reminds me of the Missing Dollar Riddle [1], where the listener is deliberately put on a wrong thinking path, to fool it.

With your own logical thinking you might never come to this confusion, and if you never heard this riddle before, you might be tricked by it.

But as we grow in life, and get experience, we learn about these riddles and aren't fooled as easily anymore.

Maybe it'll work like that for LLMs too?

[1]: https://en.wikipedia.org/wiki/Missing_dollar_riddle

brown_munda6d ago

GLM 5.2 is really impressive at design as well. Overall loving it.

anArbitraryOne6d ago

It's fine if it hallucinates, as long as it sounds overconfident

ecommerceguy6d ago

nathan_compton6d ago

It seems like for agentic coding, just making sure the AI can find the relevant documentation to establish a ground truth is probably sufficient.

dgellow6d ago

> One of the biggest models in the world was banned because a single jailbreak was too much of a risk.

We really don't know what the actual reason is given the politics at play. I would bet more on the Trump administration looking for any excuse to punish Anthropic

Naveja6d ago

loving glm 5.2 personally

metalman5d ago

to paraphrase the title, "in the land of the insane, those who are meerly delusional will rule"

abracadobre6d ago

This is where I asked GPT 5.5

"they say u hallucinate 3x more than GLM 5.2, whats your comeback to this? do i need to dump u? $article"

j / k navigate · click thread line to collapse