Hacker Newsnew | past | comments | ask | show | jobs | submit | kkoncevicius's commentslogin

What does training data look like for a model like this?

If we could show the current models to someone like Alan Turing, I am sure he would conclude that we have AGI.


You say that they pass the Turing test yet every post on HN complains about the way LLMs write so clearly they haven’t passed it yet because we can still tell it’s a bot.


You can also change the writing style with a prompt quit easily, if the same person would publish millions of articles, we would also recognize them.


The Turing Test is about being able to figure out if “someone” is a computer during a short conversation, not “millions of articles”.

LLMs still live in the uncanny valley and can be sussed out immediately.

For example, I’m extremely annoyed by the fact that offshore developers respond to me almost exclusively using text generated by Claude. You can tell immediately because they use overly descriptive techno word salad that no normal human uses unless they are trying to be ultra specific for a scientific paper - and even then it’s still too much for a real person.


My argument is, you recognize the style because you know it..

5 years ago, you would not have been able to determine that this was the case, and just have assumed it's a know it all character.

Tell your offshore developers to use caveman or so


I don't understand. do you think that people can't figure it out if they speak with llms or not from a few turns?


We don't train current LLMs to mimic the average human's writing style. We train it to be smart, helpful, and knowledgable, too knowledgeable for a human. We can easily train a LLM to pass the Turing test if we wanted to, but then it would just sound dumb or biased.


> We can easily train a LLM to pass the Turing test if we wanted to, but then it would just sound dumb or biased.

Interesting idea. "This is not dumb and biased enough, probably not a human".


There is tons of money trying to get customer support bots to sound human but they fool nobody.


> but it would just sound dumb or biased

We have one of those: Grok.


Yes, and it's indistinguishable from the average X user.


Bold of you to assume they’re actual users.


I think we’re just adapting. LLM felt kind of magical at first, and now we’re all experts in detecting AI slope.


Turing was a helluva smart guy, but he doesn't have the benefit of hindsight.


And after half an hour using it he'd just admit that his test was way too simple as these models are still way too dumb


I think of the Turing test as one of the starting lines, along with image recognition ("a summer break project for a group of grad students" resisted being solved for decades).

It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).


I don’t think he would. The Turing Test as originally formulated is a bit ambiguous but by most non-incentivised interpretations LLMs do not pass it.

In the original test the evaluator knew one was a machine and one a human and could have conversations of arbitrary length.


How are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).


I just asked Fable 5 max to create a 2d game about caterpillar climbing a tree and eating fruits. The game looks good - animations, 8-bit aesthetics, procedural tree branching, but the tree's branches are dead ends. You can't go back once you started climbing a branch. Yes, LLM doesn't have a reliable way to test it's game yet. All the screenshots, and playwright tests will never be enough to test even a simple game. But can we call a machine doing such mistakes a general intelligence? It has no embodied intelligence. No way to experience time the way we do. All it has is text. Yes, they can have images, sound too, but no big models (at least those we are supposed to use for coding) currently are native with video as far as I am concerned. And I am not sure that just video without embodied experience is enough to understand the world the way humans do. Of course we can get incredible results from machines that have a very different experience of the world than we do. But is this a general intelligence? I guess "general" is supposed to mean being able to do everything any human can do (minus the skills requiring a body)?


Current Gemini Flash models can take video input. They're not hyper-specialized coders, but they're better than the competition on many tasks. They seem to be better with spacial reasoning, as well - they are the best choice for OpenSCAD, for example.


Dumb is too simplistic, but they sure don’t pass for being human. That would be the intent of a Turing test… not IQ.


I think it's Karpathy who coined the term “jagged intelligence”. LLMs are both extraordinary smart in domain they have been explicitly trained on (like Math) and positively dumb on things they haven't.


I'm that way too. So are most people.

That's why you shouldn't listen to your pop celebrities for political advice.


Yes there are a few sites with IQ test benchmarks. The frontier models come up around 130 or 140 depending on which model/test.


It also looks like they're saturating the test, with one LLM hitting the maximum possible score. (https://www.trackingai.org/home)

The test wasn't made to accurately measure IQs that high.


That tracks, I wonder what the people who say that the models are too dumb expect to see. Miracles?


Not making mistakes my four years old would not.

They are definitely smart enough to be useful, but dumb enough in their weak spots not to deserve the "general intelligence" qualifier.


Well your generally intelligent four year old can't meet the standard you just made, by definition.


What makes my four years old “generally intelligent” is that she'll make most of these mistake exactly once.


Would you also consider a database of questions and answers smart? LLM are basically lossy text compression databases with a clever query method. Useful for sure but it’s not thinking, it’s recall.

Just look at some training sets to see how the sausage is made: https://huggingface.co/datasets/nickrosh/Evol-Instruct-Code-...


I wonder if it’s a version of Dunning-Kruger effect to call AI models dumb. I haven’t seen a “dumber than me” model since years. Also the smartest people known in the world use them in their fields so I don’t know what is meant by a “too dumb” model.


You need to be smarter (or rather: more knowledgeable in the problem domain) than the model to be able to use it efficiently. Hallucinations are still a problem occasionally but a bigger one is failure of imagination. Even Claude Fable lacks a holistic understanding of many domains it wasn't obviously trained on. The biggest problem with AI (if we assert that LLMs can be the basis of AI) is that these models will make mistakes that exist in an entirely different category of the kind of mistakes humans will make.

As an autistic this is painfully obvious to me but: much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied - not just that, but most rules are also highly contextual and rarely treated literally. E.g. corporate guidelines mostly don't exist to be followed (and following them will often result in punishment) but to be able to shift blame - but you need to know for which ones this is the case and for which ones it isn't. This is further complicated because any AI or AI vendor openly making such distinctions would be rejected - AI would not only need to understand all this nuance but also this additional meta layer.


> much of human interactions operates on rules that are not only unspoken but often unacknowledged or even outright denied.

This sounds interesting on it's own. I would be curious to hear more if you are willing to share.


Which part? I assume you mean the denial?

A simple example would be "work to rule": in many professions work processes are heavily regulated (whether by law or by corporate guidelines) but the unspoken assumption is that you know which rules you should ignore and which ones you actually need to follow - but if you tried to find this out by asking "is this a rule I need to follow or not" you would get the clearly incorrect answer that all rules must be followed; of course if you did follow all the rules (aka "work to rule") you would be disciplined for failing to meet quotas (because you can't be disciplined for following the rules).


Looking at your ladder, made me think where I fit in and I mostly enjoy steps 1 and 2. The implementation and shipping are big chores to me. Once I know, half in my head, that the problem is solvable and I could see a path towards implementation I loose motivation to continue.

So possibly there are other classes of people, not just either enterpreneur or thinkerer. Thou not sure what this combination is, maybe something more like R&D.

And LLMs give a lot of value here. They allow to quickly do steps 2-3 - "hard and tedious" stuff - to confirm that the thing, in fact, became solvable.


The first entries are a bit weird to me. They are mostly about eugenics for people who got their Nobel prices in the early 1900s. The thing is that in those days even the state of California promoted eugenics [1]. And according to Wikipedia, so much so that they might have even played a role in influencing Germany [2].

[1]: https://en.wikipedia.org/wiki/Eugenics_in_California

[2]: https://en.wikipedia.org/wiki/Nazi_eugenics


You've touched on the whole problem with the "cancel culture" of today.

Scientific and political beliefs often do not stand the test of time, but is it fair to blame the people of past eras for their misjudgements using the values of today?

How do we know for sure when we're right about something?

https://www.youtube.com/watch?v=E8V8rtdXnLA


> Scientific and political beliefs often do not stand the test of time, but is it fair to blame the people of past eras for their misjudgements using the values of today?

the problem with this reasoning is that it assumes or tries to imply that people "back then" all held such backwards beliefs, which far too often is not the case. For instance, see how Morse's diatribes on slavery and racism are presented - but the state he was born and operated in already abolished slavery about a decade before his birth.


It then rips on scientists who believed that COVID19 was the result of a lab leak.

The whole page is poor.


This was true some time ago but nowadays I don't get this impression. Seems like whatever style I type in, the LLM is already pre-prompted to respond in "its" "own" "style".


Aha, missed it. But unable to delete the thread now.


I had a similar experience. I wanted to test it by asking it to summarise a scientific OMICs-related paper. It gave a warning about me potentially developing a bio-weapon or something like that. And switched back to Opus 4.8.


We can look at the same numbers in different way:

  Error with 91.3% = 8.7%
  Error with 94.5% = 5.5%

  Error reduction = 8.7% - 5.5% = 3.2%
So the improvement is 3.2% / 8.7% = 36.8%



Ah yes I forgot about that one, nice. Reminds me of https://www.fromjason.xyz/


A lot of art from the middle ages is anonymous. Painting itself is an extension of the artist, containing the intension of the person producing it and hence no name is necessary. This is a theoretical state of quality, where activity is not measured in numbers or on a scale but is seen as expression of a particular unique human being. Then comes the renaissance and painters begin to attach names to their works. Here starts a crucial shift - a turn from quality to quantity. Certain artists are better than others and hence quality itself is now measured (quantified) using a name of the person. After that the name becomes so prevalent that some works begin to be valuable only because a certain name was responsible in producing that work. Think - Picasso. Quantity starts to take over. Then comes film and comics and ads where the painter is expected to have no individuality, and he is praised for having a style and technique that is replaceable. Same is true for corporate software development by the way. Here the name (the intermediate state connecting quality and quantity) starts to disappear and is often replaced by a name of a "golem" - a corporation. Quantity dominates - more and faster is better, and the more "nameless" the better. Ten years ago one might think that this is the limit of dehumanisation and it cannot move any further. But now we have AI - where a work of art (or other kind of work) cannot be associated with any quality (cannot be given a name) in principle. And quantity (more, faster, cheaper) dominates. When you think in these terms, the "techno-optimism" is just a place somewhere in this arrow moving from quality to quantity. Or in other words moving from a qualitative anonymity (my work is an extension of my being) to quantitative anonymity (the work is not associated with any being). Hence, it is not a stable position.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: