Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Sure, but that claim wouldn't be true for humans, right? So it's a nonsequiteur.

The relevant claim would be: all humans can do is move around in their environments, adapt the world around them through action, observe using adaptive sensory motor systems, grow and adapt their brains and bodies in response to novel and changing environments, abstract sensory motor techniques into symbolic concepts, vocalize this using inherited systems of meaning acquired as very young children in adaption within their environments, etc.

In the case of transformers all they can do is, in fact, sample from a compression of historical texts using a weighted probability metric.

If you project both of these into "problems an office worker has"-space, then they can appear simimlar -- but this projection is an incredibly dumb one, and offered as a sales pitch by charlatans looking to pretend that a system which can generate office emails can communicate.



> all they can do is, in fact, sample from a compression of historical texts

To me, results like the Othello paper make any sort of "stochastic parrot" thinking completely untenable.

https://thegradient.pub/othello/


Abstract functions are fully representable by function approximations in the limit n->inf; ie., sampling from a circle becomes a circle as samples -> infinity.

This makes all "studies" whose aim is to approximate a fully representable abstract mathematical domain irrelevant to the question.

This is just more evidence of the naivety, mendacity, and pseudoscientific basis of ML and its research.


...I see...


As you sample all pixels from all photos on a mountain, the pixels don't become the mountain.

The structure of a mountain is not a pattern of pixels. So there is no function for a statistical alg to approximate, no n->infinity which makes the approximation exact.

By sampling from historical pixel patterns in previous images you can generate images in a pixel order that makes sense to a person already acquainted with what they represent. Eg., having seen a mountain (, having perspective, colour vision, depth, counterfactual simulation, imagination, ...).

In all these disagreeably dumb research papers that come out showing "world models" and the like you have the bad mathematicians and bad programmers called "AI researchers" giving a function approximation alg an abstract mathematical domain to approximate.

ie., if the goal is to "learn a circle" and you sample points from a circle, your approximation becomes exact in n->inf, because the target is *ABSTRACT*.

It's so dumb its kinda incomprehensible. It shows what a profound lack of understanding of science is rampent across the discipline.

MNIST, Games, Chess, Circles, Rulesets, etc. are all mathematical objects (shapes, rules). It is trivial to find a mathematical approximation to a mathematical object.

The world is not made out of pixels. Models of pixel patterns are not their targets.


This result is an argument for the conclusion you are reading it as arguing against.


That's because you don't understand what you're reading.


> all they can do is, in fact, sample from a compression of historical texts using a weighted probability metric.

I don't think that's all they can do.

I think they know more than what is explicitly stated in their training sets.

They can generalize knowledge and generalize relationships between the concepts that are in the training sets.

They're currently mediocre at it, but the results we observe from SOTA generative models are not explainable without accepting that they can create an internal model of the world that's more than just a decompression algorithm.

I'm going to step away from LLMs for a moment, but: How are video generator models capable of creating videos with accurate shadows and lighting that is consistent in the entire frame and consistent between frames?

You can't do that simply by taking a weighted average of the sections of videos you've seen in your training set.

You need to create an internal 3D model of the objects in the scene, and their relative positions in space across the length of the video. And no one told the model explicitly how to do that, it learned to do it "on its own".

I think the same principle applies to LLMs.


>You need to create an internal 3D model of the objects in the scene, and their relative positions in space across the length of the video. And no one told the model explicitly how to do that, it learned to do it "on its own".

Compression is understanding. If you have a model which explains shadows you can compress your video data much better. Since you "understand" how shadows work.


> In the case of transformers all they can do is, in fact, sample from a compression of historical texts using a weighted probability metric.

You seem to think LLMs operate independently from humans. That doesn't happen in practice. We prompt LLMs, they don't just sample at random. We teach them new skills, share media and stories with them, work, learn and play together. It's not LLMs alone. They are pulled outside their training distribution by the user. The user brings their own unique life experience into the interaction.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: