We're choosing to call LLMs (and the little "harness" programs that query them in loops and execute their output) "AI", even though it doesn't make much sense.
I absolutely love this technology but these aren't autonomous intelligences. They're little programs executing Bash scripts from JSON output.
Our ideas about AI were naive. We thought passing a basic Turing test would require human-like intelligence. It turned out to be possible with fairly basic statistical text generation, because fooling humans is easy.
It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning. The irony is that the startup founders most worried about "AI" have created so much hype and funding that we may very well figure out how to build "real" AI.
netdevphoenix 27 minutes ago [-]
> We're choosing to call LLMs (and the little "harness" programs that query them in loops and execute their output) "AI", even though it doesn't make much sense.
Most people are laypeople who have no idea what's on the other side of their fave chatbot page. As far as laypeople are concerned, AI has always been a talking machine. The literature and filmography has reinforced this idea. So as soon as a talking machine emerged, people applied those fictional concepts onto reality.
Tech people should have known better than to jump on this bandwagon.
ACCount37 20 hours ago [-]
We've chosen to call Deep Blue and Half-Life 1 NPCs "AI" too.
It boggles my mind that this "b-b-but it's not actual real AI" whine is even a thing. Were people saying this living in the cave for the past 5 decades of AI research?
jacobgold 20 hours ago [-]
You can call your little doggy "AI" if it makes you happy.
But when you call something "AI" and it tells you to walk instead of drive to the car wash, you're not talking about the "AI" science fiction authors were dreaming of.
TeMPOraL 20 hours ago [-]
That's a bad and tired example as it confuses people just as much as it does (or rather, did?) LLMs.
andai 19 hours ago [-]
Yeah it's a trick question, the human error rate for it was about 30% (higher depending on the country).
The thing there though is that, if a human were given time to think about it, they'd probably go "hang on a minute", and with the LLMs that didn't seem to happen. They just kept confidently reasoning down the absurd path.
That reminds me, I recently had an AI write a ton of tests proving the "correctness" of a feature it had implemented completely backwards. (I noted that if I had been using a language that required formal proofs, that wouldn't have helped either: it would have just provided a formal proof for the absurd implementation!)
jacobgold 18 hours ago [-]
> Yeah it's a trick question, the human error rate for it was about 30% (higher depending on the country).
Error rate doesn't prove anything. The nature of the errors is what matters.
pixl97 12 hours ago [-]
To err is human, it also seems that to err is AI.
interstice 7 hours ago [-]
Seems like we need to update that other saying - to err is human but to really f** things up you need an AI
jacobgold 20 hours ago [-]
It's one example that points out a major (possibly fundamental) flaw. I can point to prompt injection as another example. There are tons more if you're interested.
Are you actually claiming LLMs operate based on human-like intelligence?
ACCount37 19 hours ago [-]
We're on Hacker News. Do I really have to point out the existence of social engineering to you? Or that scamming old people out of their life savings is a profitable enough activity that there are entire call centers dedicated to the task?
Humans keep overestimating just how high the bar of "human-like intelligence" is.
jacobgold 19 hours ago [-]
Drawing the conclusion that "humans fail" and "models fail", so they must be similar, is very wrong.
You could have humans calculate 2+2 all day and get a surprisingly high error rate. That reveals a flaw in how humans operate.
LLMs fail for entirely different reasons. Their mistakes don't imply they're human-like at all.
It's not about the error rate.
ACCount37 19 hours ago [-]
You're saying that a class of mistakes points out a "major (possibly fundamental) flaw". I'm pointing out some very similar classes of mistakes in humans - well known, well documented and widely exploited. They just keep paying the "IRS" in gift cards, buying lottery tickets and getting the captain's age wrong.
If you're using the existence of flaws in LLMs to deny the claim of intelligence to them, then why do "generally intelligent" humans exhibit some impressively similar-looking flaws?
And, if we're talking about that conspicuous similarity - do they actually fail "for entirely different reasons"? Or do you just want the reasons to be "entirely different" - and not the same reasons viewed at a different angle?
Because the similarities between humans falling for trick questions or scams, and LLMs falling for adversarial questions or prompt injections don't look coincidental to me at all.
One of the oldest patterns in scamming is overwhelming and confusing the victim. Numerous prompt injection methods seek to overwhelm and confuse an LLM - if an LLM can't keep track of things, can't grasp what's going on, it's far more likely to lose track of what's a prompt and what's data, overlook past instructions or go past its behavioral guardrails.
And humans who fall for trick questions like "1kg of feathers" or "captain's age" due to shallow attention and naive pattern matching? They fail in surprisingly similar ways to how LLMs fail on SimpleBench tasks that are filled with overwhelming adversarial distractors. Many "trick questions" are tricky to humans and LLMs alike - to the point that it's unlikely to be coincidental.
jacobgold 18 hours ago [-]
> If you're using the existence of flaws in LLMs to deny the claim of intelligence to them...
That's not the point at all. It's the fact that they fail in ways completely unlike humans.
You also have the burden of proof reversed. Its on you to prove these LLM agents are human-like intelligences if that's your claim. No one can prove this because it's false.
ACCount37 16 hours ago [-]
You are the one claiming that "they fail in ways completely unlike humans" insistently. Now go cough up some proof. I'll wait.
Jensson 10 hours ago [-]
If they didn't you wouldn't need the operator, you'd have replaced all your programmers with no drawbacks by now. As long as we keep hiring humans that is all the evidence you need that these AI fails in ways humans don't.
gghackernewsgg 13 minutes ago [-]
Junior developers fail in different ways than senior developers too; that's why seniors oversee juniors. But this doesn't necessarily mean that the senior's and junior's intelligences differ in kind
Your argument "AI needs supervision, therefore it fails in different ways than its operator does" holds.
Your argument "AI fails in different ways than its operator, therefore the AI's intelligence is different in kind" doesn't hold.
jacobgold 11 hours ago [-]
This isn't even controversial. The proof is available to anyone who uses these systems:
They hallucinate tool state, drift from the objective while seeming to comply, switch languages randomly (Cyrillic or Japanese characters in output), confuse tasks they've planned for completed ones, and of course follow prompt injections embedded in files or web pages.
vidarh 2 hours ago [-]
I switch languages "randomly" all the time when I think about something in another one of the languages I know. Some word will trigger it and before I know it I will continue in the other language.
In fact just the other day I commented on it to my fiancee after I randomly switched to French because we were discussing a trip and I mentioned a French location and pronounced it in French, and suddenly I was in "French mode" entirely unintentionally and it took a sentence before I realised.
That you think this is unique to LLM's suggests you simply don't know the diversity of human thought as well as perhaps you think you do. That's fine - none of us have a very complete view of that.
TeMPOraL 4 hours ago [-]
So just like me, including the prompt injections if you count "nerd sniping" as such?
(And in particular, switching languages on the fly is normal for people who speak more than one well, it's something you learn not to do for the sake of people less comfortable with the languages involved.)
Kim_Bruning 11 hours ago [-]
> Are you actually claiming LLMs operate based on human-like intelligence?
Ok, so we've established that it doesn't work like a human being. To paraphrase Dijkstra: The submarine doesn't swim.
But does it exactly sail either? An LLM doesn't exactly work like traditional deterministic software either, does it?
And yet it moves. You can put in data and ask it to process it, and you'll get an answer that's in some ballpark. Closer to quantum or stochastic computing perhaps, but that's not it either, is it? Or SAT-solving? Eh. It's its own computing approach. If you have a problem where the asking is hard but the verification is cheap, it might just be the right tool for the job.
orangecat 20 hours ago [-]
You can call your little doggy "AI" if it makes you happy.
Or you can keep calling them stochastic parrots as they solve decades-old open problems. The real question is how useful they are, and the answer "not at all" increasingly requires flat-earth levels of denial.
it tells you to walk instead of drive to the car wash, you're not talking about the "AI" science fiction authors were dreaming of
They sort of are. Think of Data from Star Trek TNG failing to understand figures of speech. Not that it's terribly relevant; humans regularly fall for tricks like "Paris in the the spring" or "where do you bury the survivors".
vidarh 2 hours ago [-]
If anything a large proportion of stories about AI in sci fi is about AI failing to understand humans in various ways.
jacobgold 19 hours ago [-]
> Or you can keep calling them stochastic parrots as they solve decades-old open problems.
I didn't use that phrase at all. But computers calculated digits of π to trillions of digits. With a chat interface for a Python math program would look like the most impressive math genius if you took it back a few decades.
> The real question is how useful they are...
That's not the "real question" but an entirely different question that is easily answered. Nothing I wrote suggested they're not incredibly useful.
> Data from Star Trek TNG failing to understand figures of speech.
These are just little instances of bad writing. Data is very much an attempt at displaying a human-like intelligence.
Kim_Bruning 3 hours ago [-]
> That's not the "real question" but an entirely different question that is easily answered. Nothing I wrote suggested they're not incredibly useful.
Oh, ok then. That does change things a bit. The impression I'm getting is that you were suggesting they're not. What's succinctly the thing you're objecting to?
Is it Anthropomorphization?
I mean, sure, but watch out : when defending on that axis, it's easy to slip into Anthropodenial, right? Frans de Waal (from the same science that invented "Don't Anthropomorphize" ) can tell you about it.
> With a chat interface for a Python math program would look like the most impressive math genius if you took it back a few decades.
Well, exactly. Whether any particular generation of AI or software is yes/no "Like A Human Being" is probably the least interesting question axis. It's all just anthropocentrism.
Is that the thing you're trying to lay your finger on?
lproven 52 minutes ago [-]
> The real question is how useful they are
No, it is not.
> and the answer "not at all" increasingly requires flat-earth levels of denial.
No, it does not.
For me, after ~25 years in the skeptics movement, I think the parallels with supplementary, complementary and alternative medicine are most useful.
I choose that term intentionally: its initials are S.C.A.M. and that's exactly what it is. As Tim Minchin and Alan Kay both noted, "we have a special term for alternative medicine that's been tested and shown to work. It's called 'medicine'."
If it worked, it'd be normal standard clinical medicine. But it doesn't work, and so it isn't.
And yet, SCAM is a multi-billion-dollar industry. People have ostensibly official qualifications like "ND", for "naturopathic doctor", even though that person is not a doctor and can't make you better from any kind of illness at all. Colleges teach it, millions use it, and yet, it does not work.
Which means we need to ask:
1. What does "It works! It's useful!" really mean?
2. How do we know it does not in fact work?
As a handy example, let's look at homeopathy.
Here's a quick list of things widely believed...
* It's traditional. It isn't. It was invented by Samuel Hahnemann in 1796.
* It's a kind of herbal medicine. It isn't. One widely-used ingredient is duck's liver ("Oscillococcinum"). Ducks are not herbs and neither are their livers.
* It's been proved to work. It hasn't.
We can go through the principles and prove it doesn't work even without going into a laboratory.
The principle is, "like cures like." A substance that causes symptoms like a given disease can treat that disease.
Fact: they can't.
Then we make that substance stronger by successive, succussive dilution.
Fact: it doesn't. That's why we say things are "watered down".
Succussive: you have to mix the diluted substance by banging the bottle against a copy of Hahnemann's book. Dude knew how to make money.
Fact: Dilution does not work.
That's why we call things "watered down." It makes them weaker.
Sufficiently high dilutions can be shown by statistics to have not a single molecule of the substance left, but that's OK because "water has a memory".
Fact: water does not have a memory.
We know from the principles it cannot work.
Relevance to AI: we know how the transformer algorithm works. It cannot think. Adding a few feedback loops for more plausible, but much more computationally expensive, answers does not miraculously add thinking, any more than banging a test tube of water and duck's liver magically mixes it better.
But people believe it, so it's been tested. It doesn't work. It doesn't work on people, or in vivo meaning when tested on animals, or in vitro meaning when tested in the lab on cell culture, or in silico which means in computational simulation.
*BUT!*
Most people get better from most things. This is called "reversion to the mean" and if it weren't so the first cold would have wiped out the cavemen.
What it can do, like all SCAM treatment, is make people feel better.
Being treated by a nice friendly doctor makes people feel better. It does not make them better -- it is only a state of mind.
That can sometimes marginally help gravely ill people rally, but only very rarely.
There is also the placebo effect, also much misunderstood.
This makes someone FEEL as if they'd had medicine if they think they've had medicine.
They do not get better. They just feel better for a bit. If they are ill, they remain ill. If they are dying, they still die.
But it might hurt less.
The placebo effect is very strong. Medicine from a person in a white coat works better than form the same person in street clothes.
Very big pills work better than smaller ones... but very small pills work better still, as a tiny pill suggests to people it's a very strong drug.
This is what "But AI works!" really means.
It makes people think they're doing less work -- in tests, they in fact do more, checking and fixing. Unless they don't check or fix, in which case, they are irresponsible fools.
It makes people think it can do amazing things because it can find prior art in its corpus they couldn't find -- or didn't look for, or know how to search for.
It does not save the need for skills.
Experienced practitioners can front-load the work with really detailed prompts which cover exceptions, edge cases, and things that novices don't know about. But the novices don't know that they don't know. (It enhances the illusion of competence. It helps the skilled more than it helps the unskilled, but neither realises, and it prevents the unskilled learning by trial and error. It reduces the supply of skilled workers.)
The reason AI works is the reason that people see the face of Jesus in slices of toast, as someone said recently.
ACCount37 20 hours ago [-]
Sure, just keep moving the goalposts. It's not a "real AI" because it can't take over the US military command and kick off WW3 and finish the survivors off with killer robots yet!
dofm 11 hours ago [-]
It can however select a girls school as a military target, and did.
_doctor_love 20 hours ago [-]
That's how it works though. The moment we have "AI" and see something working, it immediately ceases to be magic because "it's just a program after all." Aligning on a true definition of Artificial Intelligence is a very vexing problem.
jacobgold 20 hours ago [-]
We could've slapped a chat interface on calculators and called them "AI" because they can do superhuman math instantly. Most technical people would've thought that was stupid.
LLMs are the same kind of category mistake.
CamperBob2 19 hours ago [-]
Let's talk about category mistakes. You've been here since 2007, according to your other reply. You understand that calculators have as much to do with mathematics as telescopes have to do with cosmology. Right?
If someone unskilled at math brings a calculator to an international math competition, they will not succeed at solving many problems. Most likely, they will solve none at all. But if they bring a frontier LLM (and succeed at concealing it from the organizers), they can walk away with a gold medal. Such a feat requires intelligence... and if the contestant didn't provide the intelligence himself/herself, where'd it come from?
That means that analogies involving calculators are completely useless when the topic is AI. Calculators are not, and can never be, intelligent. LLMs are nothing even remotely like calculators.
jacobgold 18 hours ago [-]
> If someone unskilled at math brings a calculator to an international math competition, they will not succeed at solving many problems.
> But if they bring a frontier LLM (and succeed at concealing it from the organizers), they can walk away with a gold medal.
Of course you could win all kinds of math competitions with a concealed calculator. Maybe you'd need a fancy one, like a little SBC running Python. Anything complex and timed would be easy to win. You'd look like a genius to anyone who didn't know you had it.
> Such a feat requires intelligence... and if the contestant didn't provide the intelligence himself/herself, where'd it come from?
From computer software running on computer hardware, just like a calculator.
Calculating trillions of digits of pi also requires intelligence far beyond human capacity.
Computers displaying intelligence doesn't imply human-like intelligence. This is the source of confusion.
scarmig 14 hours ago [-]
The idea that a calculator, or a calculator with Python, or even a calculator with a proof assistant (and every book ever written on math) would help a random person at e.g. IMO or Putnam is fairly revealing.
jacobgold 12 hours ago [-]
The idea that anyone would think anyone else would think that is fairly revealing.
12 hours ago [-]
CamperBob2 13 hours ago [-]
Yeah, this is definitely one of those "Smile, nod, back away slowly, reach for doorknob" threads.
CamperBob2 18 hours ago [-]
Of course you could win all kinds of math competitions with a concealed calculator. Maybe you'd need a fancy one, like a little SBC running Python.
My mistake.
jacobgold 17 hours ago [-]
Taking issue with using a computer to power the calculator? If so, you should know that all modern calculators are computers under the hood.
Everything I've written about calculators applies to computers doing any kind of traditional deterministic processing, without anything like LLMs.
_doctor_love 19 hours ago [-]
Funny enough, calculators went through this exact same thing when they came out. "If the calculator can do math for the students, will they still learn?"
jacobgold 19 hours ago [-]
We didn't have confused people claiming calculators were human-like intelligences doing math.
Jensson 10 hours ago [-]
We did have that, people thought computers would overtake humans very soon when computers got better than humans at such things. It happens every single time computers do a new thing that previously humans were better at. Then 10 years later people see, oh that is just calculations, of course computers are better at that.
_doctor_love 19 hours ago [-]
Sorry I don't follow what argument you're making?
pixl97 12 hours ago [-]
They seem to be making a religious argument at this point. They don't like that the AI is a very poorly defined term and they are mad as hell about it to the point of irrational forum posting.
But when you call something "AI" and it tells you to walk instead of drive to the car wash
There are so many other sites. So many others. Why are you here?
jacobgold 19 hours ago [-]
> There are so many other sites. So many others. Why are you here?
I've been here since 2007 when HN launched.
You're confused about my objection. I don't like the term "AI" but I love the technology as much as almost anyone.
ahartmetz 20 hours ago [-]
Is this website called AI Faithful News?
20 hours ago [-]
pj_mukh 15 hours ago [-]
Is this a real debate? AI has a well-defined technical definition. It’s right there in the Wikipedia [1] . Yes it’s quite a broad umbrella of systems and algorithms but it’s all AI
> It boggles my mind that this "b-b-but it's not actual real AI" whine is even a thing.
As I understand it, a major reason it's a consistent chorus is because people don't want the "AI is here" talk to drown out (and thus slow the arrival or distribution of) speech/text/popular-understanding about actual strong AGI.
To make an analogy, it could be like this:
Some people were expecting 100 tulips (because they were told tulips are available and can be ordered), and they ordered them. They received 100 daisies. And were saying "OMG, THE TULIPS ARE HERE! THE TULIPS ARE HERE!"
A nearby observer might have said, "You know, those are daisies. Not tulips."
And 95% of people might have said back, "WE GOT 100 TULIPS! SAYS SO RIGHT HERE! THEY ARE BEAUTIFUL! STOP BEING A NAY-SAYER! THESE ARE BEAUTIFUL TULIPS!"
The 5% could just to think to themselves, and could get chastised by the crowd, if they were to say say it out loud: "Well, those are not nearly as beautiful as tulips. And if you don't take it up with the seller, you may never receive the real tulips you were after. Since you think or at least act as though you've been sold them already."
ACCount37 20 hours ago [-]
In my eyes that "chorus" is just insecurity talking.
If it's not "actual real AI", we can keep pretending that human intelligence is something distinct and special - and that what our computers are doing now is some sort of other, obviously fake and vastly inferior thing.
When Deep Blue won at chess, people didn't revise their estimates of AI capabilities upwards. They revised their estimates of how much intelligence is required to play chess at world level downwards, by a lot. Surely playing chess must have never required any intelligence in the first place!
Now, the list of things that "must have never required any intelligence in the first place" includes gems like "reading comprehension at high school level", "copywriting", "frontend work", "CTF tasks", "theory of mind", "arguing with people online" and more.
If the goalposts were moved far enough that the claim to "actual intelligence" is denied to a double digit percentage of human population, hasn't something gone wrong somewhere?
simonh 19 hours ago [-]
It has, but I think in both directions.
It used to be assumed that playing chess would require the same level of general purpose problem solving cognitive skills that the best chess players possess. But of course a Chess grandmaster that spend a few minutes learning Go can beat a Chess AI at Go with no trouble at all, because a chess AI is incapable of making effective moves in Go at all. Clearly those expectations were incorrect. Pointing that out isn't revisionism.
On the other hand, intelligence is an incredibly broad term. About as broad as a term can get. Arguably Eliza, or an Excel macro has some degree of decision making ability in some sense, it's just unbelievably primitive.
So, we need to be clearer what we mean by intelligence. We're learning that as we go along. At least now we have a few more bits of the map between us and an IF statement visible to us.
pixl97 11 hours ago [-]
>So, we need to be clearer what we mean by intelligence
I disagree in one sense. The word intelligence is burned, mostly useless at this point. I've been a strong proponent of new terms that break intelligence into much smaller subcategories so we can define what different software, humans, and animals have.
Jensson 10 hours ago [-]
> When Deep Blue won at chess, people didn't revise their estimates of AI capabilities upwards. They revised their estimates of how much intelligence is required to play chess at world level downwards, by a lot. Surely playing chess must have never required any intelligence in the first place!
You are wrong, many did temporarily revise their estimates of AI capabilities upwards, but then 10 years later they realized they were wrong and adjusted chess downward as you say.
We have seen that pattern over and over.
andai 19 hours ago [-]
"But it's not really doing arithmetic," he mumbled to himself, as he punched the numbers into his Busicom LE-120A.
20 hours ago [-]
Terr_ 4 hours ago [-]
I think you're conflating two separate issues:
1. The accuracy of the label.
2. The likelihood the label will cause problematic misunderstandings.
When my rice-cooker logic is advertised as "AI", that's a stretch, sure... But it's extremely unlikely to cause an investment bubble seeking the Rice Cooker Economic Singularity, incur protests from the Rice Cooker Emancipation League, or lead to weird folks in their basement seeking divine wisdom from its vaporous whispers.
CharlesW 20 hours ago [-]
> It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning.
We've called that "AGI" since the late 90s/early 00s (depending on whether you count first use or popularization). Even if AGI does come to pass, we'll still need "AI" since not all forms of AI will be AGI.
20 hours ago [-]
lopsotronic 17 hours ago [-]
What I'm seeing here, reading this thread, is that "intelligence" isn't a thing.
"Thing" in terms of a quantifiable that you can measure with tools and reason about, reproducibly. Everyone's got some idea what it is, so you get lots of different angles, but no one has an Intelligence Ruler we can hold up to a text output and say, yep, this one's got an INT of 14.
Seems to be the crux of the disagreement.
Kon5ole 6 hours ago [-]
It's deceptively undefined I'd say. People can argue under the impression that everyone shares their idea of what "intelligence" means, before realizing that their counterpart actually has an entirely different idea of what it means.
I'm leaning towards there being a divide between those who feel "intelligence" is entirely separate from "sentience" and those who feel that one implies the other.
andai 20 hours ago [-]
>but these aren't autonomous intelligences
Well, the labs are in a weird bind. They need to keep increasing autonomy so the agents can do increasingly complex, long-horizon tasks. But at the same time, they're closely guarding against autonomy in the sense of "pursuing its own goals."
Over the past year and a half especially, several labs have mentioned adding safeguards against self-replication, resistance to shutdown etc. (Notably, shortly after they all started bragging about involving them in the AI training loop itself, i.e. "self-improvement".)
My point here is that the autonomy of which you seek might be only a few small mutations away, but the labs are actively working to prevent such a mutation. I don't expect that situation to last for very long.
Not that I expect an AI lab will be overtaken by a rogue intelligence any time soon, but that as the cost of training goes down, I expect more "open minded" organizations and individuals to become involved.
It only takes one.
That's going to be the beginning of a new era of biology, and it's a little unsettling to think about.
wodenokoto 6 hours ago [-]
> It turned out to be possible with fairly basic statistical text generation, because fooling humans is easy.
I don't think it is fair to call a GPT model "fairly basic statistical text generation" - a Markov Text Generator I'd agree can be called basic statistics, but they are not fooling any humans in a Turing test.
> It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning.
No, AI would absolutely be apt for describing a computing reasoning like a child
davidpapermill 20 hours ago [-]
> It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning.
What would a frontier API have to be able to do to satisfy you?
jacobgold 20 hours ago [-]
Maybe just a very rigorous version of the Turing test? Modern LLMs can superficially simulate conversation but it's trivial to force them into revealing their non-human like intelligence.
They've been "patched" since but all models fail basic tests like "Should I walk or drive to the car wash which is 100 feet away" by recommending you walk.
So you'd just ask questions that require theory of mind, abstract and common sense reasoning, causal inference, learning novel rules, transferring knowledge novel situations, recognizing ambiguity, etc.
vidarh 31 minutes ago [-]
How old does a child need to be before you think they have "human like intelligence"?
continuational 20 hours ago [-]
If an alien lands on Earth and learns English, would you deem it non-intelligent if you can tell it apart from a human in conversation?
I think we should consider slime mold intelligent, and realise that it's a spectrum. Path finding is AI. There are probably forms of intelligence we have yet to discover.
jacobgold 18 hours ago [-]
If an alien landed we could decide whether it seems to have a human-like intelligence or not. It could be incredibly intelligent but very non-human-like.
davidpapermill 20 hours ago [-]
Can you give me one example that works on Claude right now?
I'm never sure whether this indicates "no reasoning present" or you've just hit an odd behaviour in the AI such that its reasoning fails. For example, you present a problem in a way that's dissimilar to the way problems are presented in its training set. That doesn't mean it's not reasoning, just it can only reason correctly in some circumstances.
jacobgold 20 hours ago [-]
The models are continually patched with training and post-training. All you have to do is find an area they haven't patched yet, and they'll be just as stupid. I run into deep technical examples every day where they fail in the most basic ways no human ever would.
I'm pretty sure most people building these models would admit they don't operate as human-like intelligences? It's baffling that anyone thinks they are.
davidpapermill 19 hours ago [-]
Yes I agree, they’re an alien kind of intelligence.
But that doesn’t mean they don’t reason.
jacobgold 19 hours ago [-]
I get what you're saying but this is kind of a semantic game.
These LLM models/agents absolutely do not reason in the sense that humans do, so you're quietly redefining the word.
You can say of course decide to call them an "alien kind of intelligence" that "reasons" but you could just as reasonably say that calculators are an "alien" kind of intelligence that "reasons" about math differently than us.
ACCount37 16 hours ago [-]
And what stops what AIs do from being "reasoning"? What's the elusive magic fairy dust of reasoning that humans put into their napkin notes, but AIs neglect to put into their chain of thought scratchpads?
Do you have a RealReasoningBenchmark, perhaps, that can reliably tell apart that fake mass produced token-flavored AI reasoning from the real, organic, 100% natural human reasoning?
penteract 15 hours ago [-]
If you had access to a bunch identical copies of me that couldn't communicate with each other, you'd be able to find many questions I would give stupid answers to. I suspect I'd come out of it looking worse than an LLM.
bgilroy26 20 hours ago [-]
I would walk
wx196 19 hours ago [-]
Me too. At least it doesn't say I need to wash my car.
Hilliard_Ohiooo 20 hours ago [-]
To answer for OP:
We are now calling text and image generators "intelligent" in the same way a spell checker is intelligent.
Whatever it's become, "AI" research started as a way to study digital neurology, or how to digitize a mind, not just how to generate data.
The Turing Test should have had a caveat, it needs to fool a, "non-stupid" person, and we still have not gotten even close to passing that version.
radial_symmetry 20 hours ago [-]
What exactly would a 'non-stupid' person do to catch the latest models on a Turing Test? Aside from being aware of AI 'tells' like em-dashes.
dominotw 20 hours ago [-]
If this was true. I would repalce myself with ai that pretends to be me on slack.
my coworkers would know almost immediately if i did that.
HeatrayEnjoyer 20 hours ago [-]
> my coworkers would know almost immediately if i did that.
The same would happen if you were replaced by any random human.
daveguy 16 hours ago [-]
But immitating others is about the only thing genAI does. Sometimes "others" is a 'programmer', sometimes "others" is an 'artist', but regardless, it still does it poorly.
dominotw 15 hours ago [-]
that would be a silly test then
pixl97 11 hours ago [-]
It sounds kind of like you made up a silly test then.
daveguy 17 hours ago [-]
Not OP, but I'd settle for something that actually learns, instead of being a static pile of linear algebra. Pretending it learns because you change the input (context) doesn't count.
contagiousflow 20 hours ago [-]
strawberry
dominotw 20 hours ago [-]
Can it produce a chart topping album if its given all the tools and the prompt "produce chart topping album" .
you might say almost no humans can do tht either but some human can but no ai can.
JacobAsmuth 20 hours ago [-]
Oh, you're talking about "AGI"! In the 90's we started using the term, you should catch up!
jacobgold 18 hours ago [-]
Sorry to tell a fellow Jacob that you're the one who is out of date. The kids are calling everything "AI" and they mean "AGI", and that's the complaint.
NamlchakKhandro 13 hours ago [-]
You mean the term is a brand name now? Hoover, vacuum cleaner.
peterashford 10 hours ago [-]
The field has been called Artificial Intelligence for what, 60 plus years now. Why is it a problem now?
lproven 49 minutes ago [-]
Because we've spent something like 2 trillion dollars on it, as we hurtle into global climate collapse and WW3.
runarberg 20 hours ago [-]
This no news for people who study philosophy, as it was known since the 1980s when John Searle described the Chinese room thought experiment.
Even Turing him self did envision the Turing test as something to pass as intelligence, but rather as a more useful replacement for the troubled term.
That said, I think your quest is doomed. There will never be a superior human-like intelligence. Forever is a long time, but my reasoning for believing this is the same reason Turing offered a replacement. Intelligence is way too vague to be useful as a measurement for anything. And if we ever discover something that is more intelligent them humans (by whichever definition of intelligence) we will simply redefine intelligence to exclude that.
vidarh 24 minutes ago [-]
The Searle's Chinese room thought experiment usually reveals more about those who think it rules out a machine intelligence than it does about AI.
It rests on a staunch unwillingness to even consider the possibility that a computational process encode intelligence and reasoning, in favour of looking for the intelligence in the medium the computation runs on, and going "a-ha!" when there is nothing that looks intelligent there.
I agree with you that there will certainly be people who just continuously redefine the words to avoid accepting that AI is intelligent or reasoning, exactly for that reason - people have avoided pinning down an objective, measurable definition of these terms for a very long time, at least in part because it leads to some very uncomfortable discussions.
In particular how to define them so that they don't exclude an uncomfortable proportion of humans, but at the same time won't include entities people don't want to include (be it certain animals, or AI)
To a lot of people, the notion that there isn't a clear binary divide between human and non-human is deeply disconcerting.
joe_the_user 20 hours ago [-]
I don't think you can say the Turing test has been passed in a computer versus determined humans setting. IE humans making a strategy effort to sort humans versus computers as well as humans motivated to distinguish themselves as humans, IE, people quiz the person or machine about "common sense, reasoning, etc." and people make an effort to exhibit that reasoning. I'd concede that creating such a competition would be challenging.
I find references to LLMs fooling humans in "casual conversations" [1] but that's not how I think the original Turing test was conceived - or at least that's not all versions that existed.
At the same time, before even LLMs appeared, the exact meaning of the test was under intense debate. The "Loebner Prize" [2] being awarded to fairly simple chatbots made serious computer scientists very embarrassed.
We're choosing to call LLMs (and the little "harness" programs that query them in loops and execute their output) "AI", even though it doesn't make much sense.
They fucking solve original math problems that you can't solve. They are indisputably intelligent, and they are indisputably artificial. That makes them indisputably "artificial intelligence." Denying that (or downvoting it, for that matter) is up there with denying evolution and the Moon landings.
It's time to start flying a different flag. You're making humans look stupid.
It turned out to be possible with fairly basic statistical text generation, because fooling humans is easy.
Yes, fooling humans is easy. Yet somehow we still consider ourselves qualified to say what is "intelligent" and what isn't, even though we can't seem to define the term.
It's a semi-valid reply to deliberately-provocative phrasing on my part, I suppose. I'm over it, don't ban him. :)
I do wish that people who aren't interested in, engaged with, and informed about technical progress in AI would find someplace else to signal their disinterest, disengagement, and disregard. But that's admittedly a me problem and not an HN problem.
gausswho 20 hours ago [-]
Towards the end, he approaches the subject of digital provenance, and muses why it's not a part of our expectations. I find the argument compelling:
> If a chatbot appears to be manipulative, mean, weird, or deceptive, what kind of answer do we want when we ask why? Revealing the indispensable antecedent examples from which the bot learned its behavior would provide an explanation: we’d learn that it drew on a particular work of fan fiction, say, or a soap opera. We could react to that output differently, and adjust the inputs of the model to improve it. Why shouldn’t that type of explanation always be available? There may be cases in which provenance shouldn’t be revealed, so as to give priority to privacy—but provenance will usually be more beneficial to individuals and society than an exclusive commitment to privacy would be.
Remember 'View Source'? And how bundling engines eventually made it irrelevant? What if every piece of content had a genuinely accurate and useful View Source?
Procrastes 20 hours ago [-]
"A.I."[1] like "technology"[2] is a term colloquially reserved for things that don't work yet. Once something works, we have to call it something else.
1. "Every time we figure out a piece of it, it stops being called AI; it becomes just computation." - Ray Kurzweil
2. "Technology n. - Something that doesn't work yet." - Douglas Adams
Diogenesian 18 hours ago [-]
To be clear the root cause of this phenomenon is that the task was solved using methods that obviously have nothing to do with intelligence, so "AI" doesn't apply at all.
pixl97 11 hours ago [-]
Computers will never be intelligent, and by the time we are done, neither will humans.
andai 20 hours ago [-]
>Everybody’s already using the term, and it might seem a little late in the day to be arguing about it. But we’re at the beginning of a new technological era—and the easiest way to mismanage a technology is to misunderstand it.
I've had a recurring theme where I would name a project incorrectly, and then waste weeks or months on what turned out to be an unsolvable problem. When I figured out the actual correct name for a project, the whole thing would be solved within a few days.
Naming things correctly is hard, and the consequences of failing to do that can be pretty severe. To name something correctly, you have to understand what it is.
If you see this page, the nginx web server is successfully installed and working. Further configuration is required.
For online documentation and support please refer to nginx.org.
Commercial support is available at nginx.com.
Thank you for using nginx.
drbscl 21 hours ago [-]
Just an outage I think, the archive works for me now
simonh 21 hours ago [-]
It's slammed. HN strikes again.
brazukadev 19 hours ago [-]
I don't think the HN effect is enough to change archive.is load.
reactordev 21 hours ago [-]
You should hard refresh and try again as this issue only pertains to you
ilaksh 20 hours ago [-]
AI that we have now is not a digital animal in capability or kind, and is not currently anywhere close to taking over.
But it's still in its present form very intelligent in meaningful and useful ways. And it is not too soon to talk about concerns of a potential existential threat in the future. Because it could sooner than we might realize, threaten our existence.
Because of the potential, we should have a culture of caution as we continue to rapidly improve AI.
bilekas 20 hours ago [-]
It will only become an existential threat (in my humble opinion) when they can run influenced inference locally offline almost instantaneously. Then we need to be concerned about not being able to switch it off.
ilaksh 20 hours ago [-]
What do you mean "influenced inference" almost instantaneously? My laptop from 6 years ago can run an agent in a very fast loop in a web browser. It's not going to take over anything though since it's Gemma 4 E2B with only 2 billion parameters.
bilekas 19 hours ago [-]
Fair, then I should add an addendum that today's frontier models, which are very capable of take overs.
Your local model doesn't need to take anything over if for an extreme example it was just given an infrastructure system full access, say electricity grid, it wont have the context to create redundant copies of itself but it could easily decide humans don't need electricity anymore.
Also I'm not sure your model will have the context to know "it's time to reinfer" especiallynot "on the fly". My phrasing could be better but I'm talking about more powerful models.
bryzaguy 20 hours ago [-]
> The closest we have come to a definition of privacy is probably “the right to be left alone,” but that seems quaint in an age when we are constantly dependent on digital services. In the context of A.I., “the right to not be manipulated by computation” seems almost correct
Maybe someone can enlighten me but I really don't understand how either of these description make any sense at all. How is it not better described as "the right to decide what data can be extracted"?
sakesun 11 hours ago [-]
Love this sentence
"The need to conform to digital designs has created an ambient expectation of human subservience. A positive spin on A.I. is that it might spell the end of this torture, if we use it well."
paul7986 20 hours ago [-]
What I heard him say is that people are taking away present-day jobs, hoping new ones will rise from the ashes. Yet, he also claims we need to find the creative minds who will actually create these new roles.
I'm curious: have we found those people or those new jobs yet? Is a forward deployed engineer an example of this, yet they are now doing the job of two people (sales and coding).
ur-whale 21 hours ago [-]
> “Over time, though, more people might be included, as intermediate rights organizations—unions, guilds, professional groups, and so on—start to play a role.”
Chassez le collectiviste, il revient au galop.
aka
Once a collectivist, always a collectivist.
bena 21 hours ago [-]
I think the biggest pushback this article will get here is the date.
Although all he's saying is basically, "It's a tool, not a silver bullet". But the article is 3 years old and people will note that the models have been updated since then.
simonh 20 hours ago [-]
Sure, but they're still LLMs and still do the same things largely the same way they did 3 years ago. There are some architectural changes, and maybe these will merit a re-assessment over time, but fundamentally it's still the same basic technological approach refined and scaled up.
bena 18 hours ago [-]
And I don't disagree, but the posting of the article feels more like bait of a sort.
But I've noticed that if you mention anything that could be seen as slightly critical of LLMs, you'll get people out of the woodwork suggesting that the state of the art has made your criticism invalid.
Diogenesian 20 hours ago [-]
The models have updated but the biggest change is providers leaning in to them being "stochastic parrots," aka probabilistic computing, and if p(good response) > 0.5 then running the algorithm over and over again improves accuracy.
Of course it's gussied up as "mixture of agents" "reasoning traces" "agentic dispatching" but high-level it's Randomized Algorithms 101.
3 hours ago [-]
Kim_Bruning 3 hours ago [-]
Oh, that's an interesting angle! Do you know of texts or concepts I can look up? It might improve my coding by quite a bit.
I absolutely love this technology but these aren't autonomous intelligences. They're little programs executing Bash scripts from JSON output.
Our ideas about AI were naive. We thought passing a basic Turing test would require human-like intelligence. It turned out to be possible with fairly basic statistical text generation, because fooling humans is easy.
It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning. The irony is that the startup founders most worried about "AI" have created so much hype and funding that we may very well figure out how to build "real" AI.
Most people are laypeople who have no idea what's on the other side of their fave chatbot page. As far as laypeople are concerned, AI has always been a talking machine. The literature and filmography has reinforced this idea. So as soon as a talking machine emerged, people applied those fictional concepts onto reality.
Tech people should have known better than to jump on this bandwagon.
It boggles my mind that this "b-b-but it's not actual real AI" whine is even a thing. Were people saying this living in the cave for the past 5 decades of AI research?
But when you call something "AI" and it tells you to walk instead of drive to the car wash, you're not talking about the "AI" science fiction authors were dreaming of.
The thing there though is that, if a human were given time to think about it, they'd probably go "hang on a minute", and with the LLMs that didn't seem to happen. They just kept confidently reasoning down the absurd path.
That reminds me, I recently had an AI write a ton of tests proving the "correctness" of a feature it had implemented completely backwards. (I noted that if I had been using a language that required formal proofs, that wouldn't have helped either: it would have just provided a formal proof for the absurd implementation!)
Error rate doesn't prove anything. The nature of the errors is what matters.
Are you actually claiming LLMs operate based on human-like intelligence?
Humans keep overestimating just how high the bar of "human-like intelligence" is.
You could have humans calculate 2+2 all day and get a surprisingly high error rate. That reveals a flaw in how humans operate.
LLMs fail for entirely different reasons. Their mistakes don't imply they're human-like at all.
It's not about the error rate.
If you're using the existence of flaws in LLMs to deny the claim of intelligence to them, then why do "generally intelligent" humans exhibit some impressively similar-looking flaws?
And, if we're talking about that conspicuous similarity - do they actually fail "for entirely different reasons"? Or do you just want the reasons to be "entirely different" - and not the same reasons viewed at a different angle?
Because the similarities between humans falling for trick questions or scams, and LLMs falling for adversarial questions or prompt injections don't look coincidental to me at all.
One of the oldest patterns in scamming is overwhelming and confusing the victim. Numerous prompt injection methods seek to overwhelm and confuse an LLM - if an LLM can't keep track of things, can't grasp what's going on, it's far more likely to lose track of what's a prompt and what's data, overlook past instructions or go past its behavioral guardrails.
And humans who fall for trick questions like "1kg of feathers" or "captain's age" due to shallow attention and naive pattern matching? They fail in surprisingly similar ways to how LLMs fail on SimpleBench tasks that are filled with overwhelming adversarial distractors. Many "trick questions" are tricky to humans and LLMs alike - to the point that it's unlikely to be coincidental.
That's not the point at all. It's the fact that they fail in ways completely unlike humans.
You also have the burden of proof reversed. Its on you to prove these LLM agents are human-like intelligences if that's your claim. No one can prove this because it's false.
Your argument "AI needs supervision, therefore it fails in different ways than its operator does" holds.
Your argument "AI fails in different ways than its operator, therefore the AI's intelligence is different in kind" doesn't hold.
They hallucinate tool state, drift from the objective while seeming to comply, switch languages randomly (Cyrillic or Japanese characters in output), confuse tasks they've planned for completed ones, and of course follow prompt injections embedded in files or web pages.
In fact just the other day I commented on it to my fiancee after I randomly switched to French because we were discussing a trip and I mentioned a French location and pronounced it in French, and suddenly I was in "French mode" entirely unintentionally and it took a sentence before I realised.
That you think this is unique to LLM's suggests you simply don't know the diversity of human thought as well as perhaps you think you do. That's fine - none of us have a very complete view of that.
(And in particular, switching languages on the fly is normal for people who speak more than one well, it's something you learn not to do for the sake of people less comfortable with the languages involved.)
Ok, so we've established that it doesn't work like a human being. To paraphrase Dijkstra: The submarine doesn't swim.
But does it exactly sail either? An LLM doesn't exactly work like traditional deterministic software either, does it?
And yet it moves. You can put in data and ask it to process it, and you'll get an answer that's in some ballpark. Closer to quantum or stochastic computing perhaps, but that's not it either, is it? Or SAT-solving? Eh. It's its own computing approach. If you have a problem where the asking is hard but the verification is cheap, it might just be the right tool for the job.
Or you can keep calling them stochastic parrots as they solve decades-old open problems. The real question is how useful they are, and the answer "not at all" increasingly requires flat-earth levels of denial.
it tells you to walk instead of drive to the car wash, you're not talking about the "AI" science fiction authors were dreaming of
They sort of are. Think of Data from Star Trek TNG failing to understand figures of speech. Not that it's terribly relevant; humans regularly fall for tricks like "Paris in the the spring" or "where do you bury the survivors".
I didn't use that phrase at all. But computers calculated digits of π to trillions of digits. With a chat interface for a Python math program would look like the most impressive math genius if you took it back a few decades.
> The real question is how useful they are...
That's not the "real question" but an entirely different question that is easily answered. Nothing I wrote suggested they're not incredibly useful.
> Data from Star Trek TNG failing to understand figures of speech.
These are just little instances of bad writing. Data is very much an attempt at displaying a human-like intelligence.
Oh, ok then. That does change things a bit. The impression I'm getting is that you were suggesting they're not. What's succinctly the thing you're objecting to?
Is it Anthropomorphization?
I mean, sure, but watch out : when defending on that axis, it's easy to slip into Anthropodenial, right? Frans de Waal (from the same science that invented "Don't Anthropomorphize" ) can tell you about it.
> With a chat interface for a Python math program would look like the most impressive math genius if you took it back a few decades.
Well, exactly. Whether any particular generation of AI or software is yes/no "Like A Human Being" is probably the least interesting question axis. It's all just anthropocentrism.
Is that the thing you're trying to lay your finger on?
No, it is not.
> and the answer "not at all" increasingly requires flat-earth levels of denial.
No, it does not.
For me, after ~25 years in the skeptics movement, I think the parallels with supplementary, complementary and alternative medicine are most useful.
I choose that term intentionally: its initials are S.C.A.M. and that's exactly what it is. As Tim Minchin and Alan Kay both noted, "we have a special term for alternative medicine that's been tested and shown to work. It's called 'medicine'."
If it worked, it'd be normal standard clinical medicine. But it doesn't work, and so it isn't.
And yet, SCAM is a multi-billion-dollar industry. People have ostensibly official qualifications like "ND", for "naturopathic doctor", even though that person is not a doctor and can't make you better from any kind of illness at all. Colleges teach it, millions use it, and yet, it does not work.
Which means we need to ask:
1. What does "It works! It's useful!" really mean?
2. How do we know it does not in fact work?
As a handy example, let's look at homeopathy.
Here's a quick list of things widely believed...
* It's traditional. It isn't. It was invented by Samuel Hahnemann in 1796. * It's a kind of herbal medicine. It isn't. One widely-used ingredient is duck's liver ("Oscillococcinum"). Ducks are not herbs and neither are their livers. * It's been proved to work. It hasn't.
We can go through the principles and prove it doesn't work even without going into a laboratory.
The principle is, "like cures like." A substance that causes symptoms like a given disease can treat that disease.
Fact: they can't.
Then we make that substance stronger by successive, succussive dilution.
Fact: it doesn't. That's why we say things are "watered down".
Succussive: you have to mix the diluted substance by banging the bottle against a copy of Hahnemann's book. Dude knew how to make money.
Fact: Dilution does not work.
That's why we call things "watered down." It makes them weaker.
Sufficiently high dilutions can be shown by statistics to have not a single molecule of the substance left, but that's OK because "water has a memory".
Fact: water does not have a memory.
We know from the principles it cannot work.
Relevance to AI: we know how the transformer algorithm works. It cannot think. Adding a few feedback loops for more plausible, but much more computationally expensive, answers does not miraculously add thinking, any more than banging a test tube of water and duck's liver magically mixes it better.
But people believe it, so it's been tested. It doesn't work. It doesn't work on people, or in vivo meaning when tested on animals, or in vitro meaning when tested in the lab on cell culture, or in silico which means in computational simulation.
*BUT!*
Most people get better from most things. This is called "reversion to the mean" and if it weren't so the first cold would have wiped out the cavemen.
What it can do, like all SCAM treatment, is make people feel better.
Being treated by a nice friendly doctor makes people feel better. It does not make them better -- it is only a state of mind.
That can sometimes marginally help gravely ill people rally, but only very rarely.
There is also the placebo effect, also much misunderstood.
This makes someone FEEL as if they'd had medicine if they think they've had medicine.
They do not get better. They just feel better for a bit. If they are ill, they remain ill. If they are dying, they still die.
But it might hurt less.
The placebo effect is very strong. Medicine from a person in a white coat works better than form the same person in street clothes.
Very big pills work better than smaller ones... but very small pills work better still, as a tiny pill suggests to people it's a very strong drug.
This is what "But AI works!" really means.
It makes people think they're doing less work -- in tests, they in fact do more, checking and fixing. Unless they don't check or fix, in which case, they are irresponsible fools.
It makes people think it can do amazing things because it can find prior art in its corpus they couldn't find -- or didn't look for, or know how to search for.
It does not save the need for skills.
Experienced practitioners can front-load the work with really detailed prompts which cover exceptions, edge cases, and things that novices don't know about. But the novices don't know that they don't know. (It enhances the illusion of competence. It helps the skilled more than it helps the unskilled, but neither realises, and it prevents the unskilled learning by trial and error. It reduces the supply of skilled workers.)
The reason AI works is the reason that people see the face of Jesus in slices of toast, as someone said recently.
LLMs are the same kind of category mistake.
If someone unskilled at math brings a calculator to an international math competition, they will not succeed at solving many problems. Most likely, they will solve none at all. But if they bring a frontier LLM (and succeed at concealing it from the organizers), they can walk away with a gold medal. Such a feat requires intelligence... and if the contestant didn't provide the intelligence himself/herself, where'd it come from?
That means that analogies involving calculators are completely useless when the topic is AI. Calculators are not, and can never be, intelligent. LLMs are nothing even remotely like calculators.
> But if they bring a frontier LLM (and succeed at concealing it from the organizers), they can walk away with a gold medal.
Of course you could win all kinds of math competitions with a concealed calculator. Maybe you'd need a fancy one, like a little SBC running Python. Anything complex and timed would be easy to win. You'd look like a genius to anyone who didn't know you had it.
> Such a feat requires intelligence... and if the contestant didn't provide the intelligence himself/herself, where'd it come from?
From computer software running on computer hardware, just like a calculator.
Calculating trillions of digits of pi also requires intelligence far beyond human capacity.
Computers displaying intelligence doesn't imply human-like intelligence. This is the source of confusion.
My mistake.
Everything I've written about calculators applies to computers doing any kind of traditional deterministic processing, without anything like LLMs.
There are so many other sites. So many others. Why are you here?
I've been here since 2007 when HN launched.
You're confused about my objection. I don't like the term "AI" but I love the technology as much as almost anyone.
[1] https://en.wikipedia.org/wiki/Artificial_intelligence
As I understand it, a major reason it's a consistent chorus is because people don't want the "AI is here" talk to drown out (and thus slow the arrival or distribution of) speech/text/popular-understanding about actual strong AGI.
To make an analogy, it could be like this:
Some people were expecting 100 tulips (because they were told tulips are available and can be ordered), and they ordered them. They received 100 daisies. And were saying "OMG, THE TULIPS ARE HERE! THE TULIPS ARE HERE!"
A nearby observer might have said, "You know, those are daisies. Not tulips."
And 95% of people might have said back, "WE GOT 100 TULIPS! SAYS SO RIGHT HERE! THEY ARE BEAUTIFUL! STOP BEING A NAY-SAYER! THESE ARE BEAUTIFUL TULIPS!"
The 5% could just to think to themselves, and could get chastised by the crowd, if they were to say say it out loud: "Well, those are not nearly as beautiful as tulips. And if you don't take it up with the seller, you may never receive the real tulips you were after. Since you think or at least act as though you've been sold them already."
If it's not "actual real AI", we can keep pretending that human intelligence is something distinct and special - and that what our computers are doing now is some sort of other, obviously fake and vastly inferior thing.
When Deep Blue won at chess, people didn't revise their estimates of AI capabilities upwards. They revised their estimates of how much intelligence is required to play chess at world level downwards, by a lot. Surely playing chess must have never required any intelligence in the first place!
Now, the list of things that "must have never required any intelligence in the first place" includes gems like "reading comprehension at high school level", "copywriting", "frontend work", "CTF tasks", "theory of mind", "arguing with people online" and more.
If the goalposts were moved far enough that the claim to "actual intelligence" is denied to a double digit percentage of human population, hasn't something gone wrong somewhere?
It used to be assumed that playing chess would require the same level of general purpose problem solving cognitive skills that the best chess players possess. But of course a Chess grandmaster that spend a few minutes learning Go can beat a Chess AI at Go with no trouble at all, because a chess AI is incapable of making effective moves in Go at all. Clearly those expectations were incorrect. Pointing that out isn't revisionism.
On the other hand, intelligence is an incredibly broad term. About as broad as a term can get. Arguably Eliza, or an Excel macro has some degree of decision making ability in some sense, it's just unbelievably primitive.
So, we need to be clearer what we mean by intelligence. We're learning that as we go along. At least now we have a few more bits of the map between us and an IF statement visible to us.
I disagree in one sense. The word intelligence is burned, mostly useless at this point. I've been a strong proponent of new terms that break intelligence into much smaller subcategories so we can define what different software, humans, and animals have.
You are wrong, many did temporarily revise their estimates of AI capabilities upwards, but then 10 years later they realized they were wrong and adjusted chess downward as you say.
We have seen that pattern over and over.
1. The accuracy of the label.
2. The likelihood the label will cause problematic misunderstandings.
When my rice-cooker logic is advertised as "AI", that's a stretch, sure... But it's extremely unlikely to cause an investment bubble seeking the Rice Cooker Economic Singularity, incur protests from the Rice Cooker Emancipation League, or lead to weird folks in their basement seeking divine wisdom from its vaporous whispers.
We've called that "AGI" since the late 90s/early 00s (depending on whether you count first use or popularization). Even if AGI does come to pass, we'll still need "AI" since not all forms of AI will be AGI.
"Thing" in terms of a quantifiable that you can measure with tools and reason about, reproducibly. Everyone's got some idea what it is, so you get lots of different angles, but no one has an Intelligence Ruler we can hold up to a text output and say, yep, this one's got an INT of 14.
Seems to be the crux of the disagreement.
I'm leaning towards there being a divide between those who feel "intelligence" is entirely separate from "sentience" and those who feel that one implies the other.
Well, the labs are in a weird bind. They need to keep increasing autonomy so the agents can do increasingly complex, long-horizon tasks. But at the same time, they're closely guarding against autonomy in the sense of "pursuing its own goals."
Over the past year and a half especially, several labs have mentioned adding safeguards against self-replication, resistance to shutdown etc. (Notably, shortly after they all started bragging about involving them in the AI training loop itself, i.e. "self-improvement".)
My point here is that the autonomy of which you seek might be only a few small mutations away, but the labs are actively working to prevent such a mutation. I don't expect that situation to last for very long.
Not that I expect an AI lab will be overtaken by a rogue intelligence any time soon, but that as the cost of training goes down, I expect more "open minded" organizations and individuals to become involved.
It only takes one.
That's going to be the beginning of a new era of biology, and it's a little unsettling to think about.
I don't think it is fair to call a GPT model "fairly basic statistical text generation" - a Markov Text Generator I'd agree can be called basic statistics, but they are not fooling any humans in a Turing test.
> It would've been nice to reserve "AI" for superior human-like intelligence capable of genuine common sense and reasoning.
No, AI would absolutely be apt for describing a computing reasoning like a child
What would a frontier API have to be able to do to satisfy you?
They've been "patched" since but all models fail basic tests like "Should I walk or drive to the car wash which is 100 feet away" by recommending you walk.
So you'd just ask questions that require theory of mind, abstract and common sense reasoning, causal inference, learning novel rules, transferring knowledge novel situations, recognizing ambiguity, etc.
I think we should consider slime mold intelligent, and realise that it's a spectrum. Path finding is AI. There are probably forms of intelligence we have yet to discover.
I'm never sure whether this indicates "no reasoning present" or you've just hit an odd behaviour in the AI such that its reasoning fails. For example, you present a problem in a way that's dissimilar to the way problems are presented in its training set. That doesn't mean it's not reasoning, just it can only reason correctly in some circumstances.
I'm pretty sure most people building these models would admit they don't operate as human-like intelligences? It's baffling that anyone thinks they are.
But that doesn’t mean they don’t reason.
These LLM models/agents absolutely do not reason in the sense that humans do, so you're quietly redefining the word.
You can say of course decide to call them an "alien kind of intelligence" that "reasons" but you could just as reasonably say that calculators are an "alien" kind of intelligence that "reasons" about math differently than us.
Do you have a RealReasoningBenchmark, perhaps, that can reliably tell apart that fake mass produced token-flavored AI reasoning from the real, organic, 100% natural human reasoning?
We are now calling text and image generators "intelligent" in the same way a spell checker is intelligent.
Whatever it's become, "AI" research started as a way to study digital neurology, or how to digitize a mind, not just how to generate data.
The Turing Test should have had a caveat, it needs to fool a, "non-stupid" person, and we still have not gotten even close to passing that version.
my coworkers would know almost immediately if i did that.
The same would happen if you were replaced by any random human.
you might say almost no humans can do tht either but some human can but no ai can.
Even Turing him self did envision the Turing test as something to pass as intelligence, but rather as a more useful replacement for the troubled term.
That said, I think your quest is doomed. There will never be a superior human-like intelligence. Forever is a long time, but my reasoning for believing this is the same reason Turing offered a replacement. Intelligence is way too vague to be useful as a measurement for anything. And if we ever discover something that is more intelligent them humans (by whichever definition of intelligence) we will simply redefine intelligence to exclude that.
It rests on a staunch unwillingness to even consider the possibility that a computational process encode intelligence and reasoning, in favour of looking for the intelligence in the medium the computation runs on, and going "a-ha!" when there is nothing that looks intelligent there.
I agree with you that there will certainly be people who just continuously redefine the words to avoid accepting that AI is intelligent or reasoning, exactly for that reason - people have avoided pinning down an objective, measurable definition of these terms for a very long time, at least in part because it leads to some very uncomfortable discussions.
In particular how to define them so that they don't exclude an uncomfortable proportion of humans, but at the same time won't include entities people don't want to include (be it certain animals, or AI)
To a lot of people, the notion that there isn't a clear binary divide between human and non-human is deeply disconcerting.
I find references to LLMs fooling humans in "casual conversations" [1] but that's not how I think the original Turing test was conceived - or at least that's not all versions that existed.
At the same time, before even LLMs appeared, the exact meaning of the test was under intense debate. The "Loebner Prize" [2] being awarded to fairly simple chatbots made serious computer scientists very embarrassed.
[1] https://neurosciencenews.com/ai-passes-turing-test-30733/ [2] https://en.wikipedia.org/wiki/Loebner_Prize
They fucking solve original math problems that you can't solve. They are indisputably intelligent, and they are indisputably artificial. That makes them indisputably "artificial intelligence." Denying that (or downvoting it, for that matter) is up there with denying evolution and the Moon landings.
It's time to start flying a different flag. You're making humans look stupid.
It turned out to be possible with fairly basic statistical text generation, because fooling humans is easy.
Yes, fooling humans is easy. Yet somehow we still consider ourselves qualified to say what is "intelligent" and what isn't, even though we can't seem to define the term.
As I just mentioned at https://news.ycombinator.com/item?id=48981624, we need you to stick to HN's rules if you want to keep commenting on the site.
https://news.ycombinator.com/newsguidelines.html
I do wish that people who aren't interested in, engaged with, and informed about technical progress in AI would find someplace else to signal their disinterest, disengagement, and disregard. But that's admittedly a me problem and not an HN problem.
> If a chatbot appears to be manipulative, mean, weird, or deceptive, what kind of answer do we want when we ask why? Revealing the indispensable antecedent examples from which the bot learned its behavior would provide an explanation: we’d learn that it drew on a particular work of fan fiction, say, or a soap opera. We could react to that output differently, and adjust the inputs of the model to improve it. Why shouldn’t that type of explanation always be available? There may be cases in which provenance shouldn’t be revealed, so as to give priority to privacy—but provenance will usually be more beneficial to individuals and society than an exclusive commitment to privacy would be.
Remember 'View Source'? And how bundling engines eventually made it irrelevant? What if every piece of content had a genuinely accurate and useful View Source?
1. "Every time we figure out a piece of it, it stops being called AI; it becomes just computation." - Ray Kurzweil
2. "Technology n. - Something that doesn't work yet." - Douglas Adams
I've had a recurring theme where I would name a project incorrectly, and then waste weeks or months on what turned out to be an unsolvable problem. When I figured out the actual correct name for a project, the whole thing would be solved within a few days.
Naming things correctly is hard, and the consequences of failing to do that can be pretty severe. To name something correctly, you have to understand what it is.
If you see this page, the nginx web server is successfully installed and working. Further configuration is required.
For online documentation and support please refer to nginx.org. Commercial support is available at nginx.com.
Thank you for using nginx.
But it's still in its present form very intelligent in meaningful and useful ways. And it is not too soon to talk about concerns of a potential existential threat in the future. Because it could sooner than we might realize, threaten our existence.
Because of the potential, we should have a culture of caution as we continue to rapidly improve AI.
Your local model doesn't need to take anything over if for an extreme example it was just given an infrastructure system full access, say electricity grid, it wont have the context to create redundant copies of itself but it could easily decide humans don't need electricity anymore.
Also I'm not sure your model will have the context to know "it's time to reinfer" especiallynot "on the fly". My phrasing could be better but I'm talking about more powerful models.
Maybe someone can enlighten me but I really don't understand how either of these description make any sense at all. How is it not better described as "the right to decide what data can be extracted"?
"The need to conform to digital designs has created an ambient expectation of human subservience. A positive spin on A.I. is that it might spell the end of this torture, if we use it well."
I'm curious: have we found those people or those new jobs yet? Is a forward deployed engineer an example of this, yet they are now doing the job of two people (sales and coding).
Chassez le collectiviste, il revient au galop.
aka
Once a collectivist, always a collectivist.
Although all he's saying is basically, "It's a tool, not a silver bullet". But the article is 3 years old and people will note that the models have been updated since then.
But I've noticed that if you mention anything that could be seen as slightly critical of LLMs, you'll get people out of the woodwork suggesting that the state of the art has made your criticism invalid.
Of course it's gussied up as "mixture of agents" "reasoning traces" "agentic dispatching" but high-level it's Randomized Algorithms 101.