483

Apple study exposes deep cracks in LLMs’ “reasoning” capabilities (arstechnica.com)

submitted 2 days ago by misk@sopuli.xyz to c/technology@lemmy.world

107 comments fedilink hide all child comments

top 50 comments

sorted by: hot top controversial new old

[-] FlyingSquid@lemmy.world 5 points 14 hours ago* (last edited 14 hours ago)

The part of the study where they talk about how they determined the flawed mathematical formula it used to calculate the glue-on-pizza response was mindblowing.

^(I^ ^did^ ^not^ ^read^ ^the^ ^study.)^

[-] nutsack@lemmy.world 32 points 1 day ago

cracks? it doesn't even exist. we figured this out a long time ago.

[-] Halcyon@discuss.tchncs.de 41 points 1 day ago

They are large LANGUAGE models. It's no surprise that they can't solve those mathematical problems in the study. They are trained for text production. We already knew that they were no good in counting things.

[-] Flocklesscrow@lemm.ee 25 points 1 day ago

"You see this fish? Well, it SUCKS at climbing trees."

[-] zbyte64@awful.systems 4 points 16 hours ago

That's not how you sell fish though. You gotta emphasize how at one time we were all basically fish and if you buy my fish for long enough, those fish will eventually evolve hands to climb!

[-] Flocklesscrow@lemm.ee 2 points 16 hours ago

"Premium fish for sale: GUARANTEED to never climb your trees"

[-] CombatWombat1212@lemmy.ml 56 points 1 day ago

So do I every time I ask it a slightly complicated programming question

[-] Saik0Shinigami@lemmy.saik0.com 19 points 1 day ago

And sometimes even really simple ones.

[-] werefreeatlast@lemmy.world 9 points 1 day ago

How many w's in "Howard likes strawberries" It would be awesome to know!

[-] Saik0Shinigami@lemmy.saik0.com 9 points 1 day ago* (last edited 1 day ago)

So I keep seeing people reference this... And I found it curious of a concept that LLMs have problems with this. So I asked them... Several of them...

Outside of this image... Codestral ( my default ) got it actually correct and didn't talk itself out of being correct... But that's no fun so I asked 5 others, at once.

What's sad is that Dolphin Mixtral is a 26.44GB model...
Gemma 2 is the 5.44GB variant
Gemma 2B is the 1.63GB variant
LLaVa Llama3 is the 5.55 GB variant
Mistral is the 4.11GB Variant

So I asked Codestral again because why not! And this time it talked itself out of being correct...

Edit: fixed newline formatting.

[-] realitista@lemm.ee 2 points 1 day ago* (last edited 1 day ago)

Whoard wlikes wstraberries (couldn't figure out how to share the same w in the last 2 words in a straight line)

load more comments (3 replies)

load more comments (4 replies)

[-] anon_8675309@lemmy.world 88 points 1 day ago

Did anyone believe they had the ability to reason?

[-] Furbag@lemmy.world 8 points 17 hours ago

Like 90% of the consumers using this tech are totally fine handing over tasks that require reasoning to LLMs and not checking the answers for accuracy.

[-] Aeri@lemmy.world 28 points 1 day ago

People are stupid OK? I've had people who think that it can in fact do math, "better than a calculator"

[-] LodeMike@lemmy.today 37 points 1 day ago

Yes

load more comments (11 replies)

[-] N0body@lemmy.dbzer0.com 55 points 1 day ago

The tested LLMs fared much worse, though, when the Apple researchers modified the GSM-Symbolic benchmark by adding "seemingly relevant but ultimately inconsequential statements" to the questions

Good thing they're being trained on random posts and comments on the internet, which are known for being succinct and accurate.

[-] blind3rdeye@lemm.ee 23 points 1 day ago

Yeah, especially given that so many popular vegetables are members of the brassica genus

[-] VantaBrandon@lemmy.world 4 points 23 hours ago

Definitely true! And ordering pizza without rocks as a topping should be outlawed, it literally has no texture without it, any human would know that very obvious fact.

load more comments (2 replies)

[-] emerald@lemmy.blahaj.zone 45 points 1 day ago

statistical engine suggesting words that sound like they'd probably be correct is bad at reasoning

How can this be??

[-] Siegfried@lemmy.world 19 points 1 day ago

I would say that if anything, LLMs are showing cracks in our way of reasoning.

[-] MoogleMaestro@lemmy.zip 12 points 1 day ago

Or the problem with tech billionaires selling "magic solutions" to problems that don't actually exist. Or how people are too gullible in the modern internet to understand when they're being sold snake oil in the form of "technological advancement" when it's actually just repackaged plagiarized material.

load more comments (1 replies)

load more comments (3 replies)

[-] jabathekek@sopuli.xyz 202 points 2 days ago

[-] WhatAmLemmy@lemmy.world 82 points 2 days ago

The results of this new GSM-Symbolic paper aren't completely new in the world of AI research. Other recent papers have similarly suggested that LLMs don't actually perform formal reasoning and instead mimic it with probabilistic pattern-matching of the closest similar data seen in their vast training sets.

WTF kind of reporting is this, though? None of this is recent or new at all, like in the slightest. I am shit at math, but have a high level understanding of statistical modeling concepts mostly as of a decade ago, and even I knew this. I recall a stats PHD describing models as "stochastic parrots"; nothing more than probabilistic mimicry. It was obviously no different the instant LLM's came on the scene. If only tech journalists bothered to do a superficial amount of research, instead of being spoon fed spin from tech bros with a profit motive...

[-] jimmy90@lemmy.world 1 points 19 minutes ago

i think it's because some people have been alleging reasoning is happening or is very close to it

[-] ObviouslyNotBanana@lemmy.world 45 points 2 days ago

It's written as if they literally expected AI to be self reasoning and not just a mirror of the bullshit that is put into it.

[-] Sterile_Technique@lemmy.world 38 points 2 days ago

Probably because that's the common expectation due to calling it "AI". We're well past the point of putting the lid back on that can of worms, but we really should have saved that label for... y'know... intelligence, that's artificial. People think we've made an early version of Halo's Cortana or Star Trek's Data, and not just a spellchecker on steroids.

The day we make actual AI is going to be a really confusing one for humanity.

load more comments (18 replies)

load more comments (3 replies)

load more comments (2 replies)

[-] whotookkarl@lemmy.world 17 points 1 day ago* (last edited 1 day ago)

Here's the cycle we've gone through multiple times and are currently in:

AI winter (low research funding) -> incremental scientific advancement -> breakthrough for new capabilities from multiple incremental advancements to the scientific models over time building on each other (expert systems, LLMs, neutral networks, etc) -> engineering creates new tech products/frameworks/services based on new science -> hype for new tech creates sales and economic activity, research funding, subsidies etc -> (for LLMs we're here) people become familiar with new tech capabilities and limitations through use -> hype spending bubble bursts when overspend doesn't keep up with infinite money line goes up or new research breakthroughs -> AI winter -> etc...

[-] KingThrillgore@lemmy.ml 27 points 1 day ago* (last edited 1 day ago)

I feel like a draft landed on Tim's desk a few weeks ago, explains why they suddenly pulled back on OpenAI funding.

People on the removed superfund birdsite are already saying Apple is missing out on the next revolution.

[-] BreadstickNinja@lemmy.world 16 points 1 day ago

"Superfund birdsite" I am shamelessly going to steal from you

load more comments (1 replies)

[-] Gradually_Adjusting@lemmy.world 94 points 2 days ago* (last edited 2 days ago)

One time I exposed deep cracks in my calculator's ability to write words with upside down numbers. I only ever managed to write BOOBS and hELLhOLE.

LLMs aren't reasoning. They can do some stuff okay, but they aren't thinking. Maybe if you had hundreds of them with unique training data all voting on proposals you could get something along the lines of a kind of recognition, but at that point you might as well just simulate cortical columns and try to do Jeff Hawkins' idea.

load more comments (5 replies)

[-] RaoulDook@lemmy.world 20 points 1 day ago

I hope this gets circulated enough to reduce the ridiculous amount of investment and energy waste that the ramping-up of "AI" services has brought. All the companies have just gone way too far off the deep end with this shit that most people don't even want.

[-] thanks_shakey_snake@lemmy.ca 18 points 1 day ago

People working with these technologies have known this for quite awhile. It's nice of Apple's researchers to formalize it, but nobody is really surprised-- Least of all the companies funnelling traincars of money into the LLM furnace.

load more comments (2 replies)

load more comments

this post was submitted on 15 Oct 2024

483 points (96.4% liked)

Technology

58711 readers

3997 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

founded 1 year ago

MODERATORS