"Sentient" is an ill-defined philosophical term that should be considered harmful in technical materials.
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
There is no "proper definition" - or even one that everyone would agree upon. There is no definition of "sentience" that I could operationalize and put into a sentience-o-meter to reliably measure just how sentient a given rock, GPU or an internet user is.
I could try to put together benchmarks to estimate an AI's cyberwarfare capabilities, or instruction-following capabilities, or reward hacking inclinations. As noisy indirect estimates, of course. With philosophical mumbo-jumbo like "sentience", I don't even get that.
Sentience, self-awareness, consciousness, etc.,those are terms signifying a bridge between "technical" information theory and the psychological and social realms.
Those are just as real, only far less predictable and not as easy as programming.
They're also far more important and consequential.
I don't like it when people take mumbo-jumbo that can't be pinned down, or measured, or even agreed upon, and try to insist that we should base decision-making on it. It's literally just vibes with extra steps.
The "far more important and consequential" thing you're touting is your ability to make decisions based purely on vibes. And not even consistent, broadly agreed-upon vibes like "murder is pretty bad". It's vibes of the most vile variety: "sentience is what I decided sentience is".
An average internet user is sentient, but a 1996 Nissan ECU isn't. Why? Because I said so. Tremble before my might!
And a human is perfectly controllable if you keep him in a sealed metal box with no access to food or air.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
An LLM, in my opinion, is not comparable to a person.
On the risk management angle, for sure it’s a spectrum. I don’t agree that the far end of the safe side of that spectrum for AI models is “entirely safe and entirely useless”, there is a lot of work you can do with a model that has zero risk of hurting anyone (aside from your wallet). If someone chooses a more dangerous spot on that spectrum, I believe they should be held responsible.
> An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
This has not been my experience. I’ve been getting a lot of good work done and, as of today, have been involved in zero Pentagon hacking incidents. ;-)
Yeah, sure, it's controllable. Because humanity has famously solved the crime accountability problem back in 1902, and no crime has gone unpunished since.
AI is perfectly controllable in a magic fairy land where nothing ever goes wrong. I can't help but notice that we aren't actually in that land.
For some crimes the tolerance level is pretty much zero. Given the considerable resource AI needs that doesn't seem impossible to clamp down on some things.
> So, how many rogue AIs are we willing to tolerate?
It will balance automatically based on the severity they cause. If they constantly break systems, punishments will go up against the operators and the effect will be similar as with other serious crimes.
That relies on anyone being able to lever a "punishment" against an AI or its operators.
Which, in turn, requires that AI oopsie to be survivable.
AI capabilities are rising over time. If there is a limit to just how far they can rise, we're yet to find it. So, a sufficiently advanced "AI oopsie" can solve the AI crime accountability problem for good. Probably not the way you would have wanted it to.
Obviously, if someone is state-level actor and allows government's employees to do whatever they please, there is no other solution than political pressure.
But for other cases, it is not different than other cyber crime. Except that these AI capabilities can't live on the toaster yet. If we get state of the art model running fast on Raspberry Pi, then we have real problems.
Not "pressure to lose this trait" as much as "no pressure to improve it"?
As anatomy gets more complex, the process of getting it from "arbitrary heavily damaged state" to "functioning state" gets more complex too. And mammals are a bit more anatomically complex than placozoa.
If your entire body is a hollow sphere 4 cells thick, "repairing arbitrary damage" is very simple and natural. When you have bones, blood vessels, nerves, muscles and tendons, all wrapped in skin - all of which have to be restored correctly for a lost limb to function well? The gap between "just plug the holes" and "restore the function" grows, and the complexity of implementing usable regeneration goes up massively.
Humans can repair most of simple tissue-level damage well enough. The complexity equivalent of placozoan regeneration is in place. Rebuilding complex anatomy is what's often unimplemented. Seems like that is the part that requires some novel adaptations rather than simply not deactivating the mechanisms that are already there.
Preprints are just a way to sacrifice rigor for accessibility and velocity. You can throw out "here, this is what I'm working on, here are the quick and dirty findings" really fast and with little friction.
This does skip the academic "checks and balances" like journal selection and peer review - but it can also help anyone else who's working on the adjacent topics.
If a field is moving fast, and you think there can be some value in your work for others in the near term? Preprint. If your work is too incomplete or too minor to warrant trying to polish and publish it, but you don't want to table it? Preprint. Too deep in corporate structures to care about academic "street cred", and want your work to be accessible? Preprint. Have an exciting early finding that you want to push out there, and are willing to take the rep risks of being wrong about it? Preprint.
There's a reason why preprints came to be the lifeblood of ML.
Academia isn't my thing but I also wonder if there isn't an aspect of putting a stake in the ground? So that if someone beats you to publishing you at least have some record of being on that track.
That is definitely a big motivator for publishing preprints. Journal submissions can take up to a year. Comference submissions take months. If the field is moving fast, claiming a finding early can become an important career move.
In older days, academics would just share notes on their work and word wouldn't usually spread widely before publication.
Preprints may be the better model. But public visibility means that non-experts now get to see the good and the bad research equally, but they won't have the domain knowledge and skill to distinguish one from the other with confidence.
Nope, no "deterministic guardrails" for you. The domain is simply far too broad and unstructured to allow for that.
Unless you mean "a typical AI with all the computation constrained sufficiently to always unfold the same exact way, given the same input". In practice, that just kicks the can to "given the same input" street.
The noise in the system is going to come from the input plane. Which is, I remind you, facing the real world. It's full of noise.
That is the kind of thing makes me wonder how much of "things that actually matter like pathophysiology and pharmacology" can be factored out into the automation land now.
AI isn't perfect, but even loosely scaffolded generalist systems show promise in the field of medicine now. And the alternative isn't some hypothetical "perfect healthcare" - the status quo is often closer to "nurses running near the limits of their competence" or "physicians stretched thin almost to the breaking point".
The fundamental problem of healthcare is that it struggles to scale. The need for well educated, well paid professionals is inescapable. Or, was inescapable? We might be at the point where this can start changing.
> That is the kind of thing makes me wonder how much of "things that actually matter like pathophysiology and pharmacology" can be factored out into the automation land now.
I would say not much. AI is still often wrong and a clinician needs to know when the LLM is saying something crazy. I think AI has the most promise for increasing the productivity of well trained professionals, not replacing them (or their training) entirely.
Human clinicians are also "often wrong", for a given definition of "often". "Get a second opinion" didn't originate with AI.
Are AIs wrong more often or less often?
Would the healthcare get better or worse if the "first opinion" was AI more often than not?
"Increasing the productivity" and "replacing them" is two sides of the same coin. If a human can do five times the work, because AI does most of the work and the human performs "exception handling"? You need less humans. And healthcare, historically, is almost always human-constrained. That's why you get insane wait times and overworked clinicians. Most other inputs scale more readily than human expertise.
Thus the impetus to figure out where "human expertise" can be substituted for that of a scalable machine system - and what would be the best ways to implement that.
I was replying to "how much of 'things that actually matter like pathophysiology and pharmacology' can be factored out into the automation land"
I'd argue they aren't being factored into automation land if they're still the responsibility of a human expert, even if there are fewer more productive human experts.
(Though I do think AI will have a role in pathophysiology and pharmacology, initially catching errors, and probably some day taking responsibility, but not soon.)
Yes, it's one of those cases where compassion works against people - because it's a dirty, costly local fix that enables the structural problem to fester.
You don't want a system that requires "heroic effort" as a baseline. You want a system where "heroic effort" is reserved for heroic circumstances.
If doctors are running ragged and putting in unreasonable hours and burning out during a natural disaster or a worldwide pandemic, it's understandable. If doctors are running ragged and putting in unreasonable hours and burning out during "business as usual"? Something's rotten.
Run long enough like this, and you'll simply deplete the people who dared to care - and leave ones who never did, or learned not to.
Or not. And replace the generalist with the next generalist that gets you +15% on that benchmark for the same price, or gives you the same benchmark performance for half the price.
One advantage of using generalist models is that the generalists are improving - regardless of whether you're doing anything about it.
Yes, but the generalists are not routinely improving across all domains. The large labs are really focusing on agentic use, so I imagine that creative writing has deteriorated considering how distinctive Claude's writing style has become. Or I recently had an image-parsing task, and I was excited to try Qwen because I heard it had gotten a lot better at agentic tasks, but it failed my image-parsing benchmark.
There are focus areas, but capabilities improve across all domains - some slower than others. "Agentic use" is in itself a very general thing - because many tasks benefit from being able to leverage adaptive model-driven workflows.
Creative writing and Claude - amusing that you say that, given that Anthropic just went and tried to unfuck it in Opus 5.5 specifically. It is an example of a capability no one typically cares about, yes. No money in creative writing. But even there, we had gains in newer models.
The issue with distillation is: one lab spends $$$ on bleeding edge R&D and expensive RL runs to improve capabilities, and other labs just yoink the raw reasoning traces and mid-train/post-train on them to get 90% of the way there for a small fraction of the cost.
An even smaller fraction of the cost if they do it by buying AI access at as much of a discount as they can find, including black market resellers, and then reselling that access to paying users again with a proxy. As is common.
This gives ruthless "fast followers" an economic edge over the innovator that's putting in the real work.
The dynamics are very much alike to what patents and copyright law are supposed to prevent. Same type of "we took the products of your work and used them to undercut you". Except there are no laws against distillation - so most of the enforcement happens on model provider level.
There's an implication that other companies are improving because they're scraping Anthropic, not because they're investing in better architecture, compute efficiency, or their own synthetic data pipelines. I often see Chinese labs' progress dismissed as "they just distilled Anthropic" and I find it hard to reconcile that with all of the interesting research and open-source tooling that they release.
Is there actually that much capability transfer from non-logit-matched distillation, or is Anthropic just another unwilling source of data?
There is, in fact, "that much capability transfer from non-logit-matched distillation".
Even the early papers on distillation techniques found that surprisingly small distillation datasets can improve task performance noticeably on some specific task types - and that valuable adaptations like SFT/RLHF instruction following can be distilled from one-hot non-logit traces.
A big part of what distillation really gets you is: paving over the mismatch between pre-training and final performance. A base model is trained to spit out fitting text, but not to instruction follow, reason autoregressively, self-check or use tool calls - like an AI has to. There is transfer straight from the "text prediction" pre-training objective, and pre-training sets the foundation for all that follows - but the capabilities you get "out of the box" with it are often unrefined and fragile. Which makes some sense - internet text doesn't often include raw chain-of-thought autoregressive reasoning. It's not the kind of thing humans tend to write.
Reasoning traces? They let an AI learn proven techniques and adaptations directly, from an AI that was already taught "how to be an AI" in other ways.
It's why this kind of distillation typically plugs into mid-training and post-training, not pre-training.
Now, I'm not saying that all Chinese companies do is eat tokens, distill and lie. That just isn't the case. They developed or refined numerous training techniques and architectural adaptations - like deep fusion for high performance visual input, RLVR with GRPO, trunked MoE, storage-efficient and bandwidth-efficient attention formulations, or residual routing techniques like AttnRes. Some of those are used widely now, and some are still on the uptake but show good promise.
But Chinese labs are enjoying massive efficiency gains from being able to distill from the frontier instead of doing things the hard way. It's a leg up. It lets them put their supply of R&D effort and RL compute elsewhere. They wouldn't be nearly as advanced if they couldn't do it.
> one lab spends $$$ on bleeding edge R&D and expensive RL runs to improve capabilities, and other labs just yoink the raw reasoning traces and mid-train/post-train on them to get 90% of the way there for a small fraction of the cost.
"You're trying to kidnap what I've rightfully stolen."
This is a great post and I agree with you on the issues with distillation. I do still feel it's ironic for an AI lab.
As long as labs do not heavily kneecap model outputs, practically all this applies to the training corpus as well.
AI gives ruthless users of AI a leg up over the people who's data it was trained on. "We took the products of your work and used them to undercut you". It's all the same.
The only way I'd be against distilling would be if AI models became owned by the public who's work is used to create them. Of course the AI labs should be paid well, but these models are a product of the entire world's efforts, not only the labs.
I'm fine if they put preventative measures in place to protect their work. They already do so. I am NOT fine with their mass manipulation of public opinion to fuel an entirely hypocritical viewpoint. Like, any argument here is hypocritical - but they aren't saying what is REALLY HAPPENING ("distillation steals our work and reduces our profits"), and are actually saying words that make other people fight their battle ("national security", etc).
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
reply