Hacker Newsnew | past | comments | ask | show | jobs | submit | ninjahawk1's commentslogin

There’s a disclaimer at the top but I’d like to reiterate it here:

This is a fictional short story and no claim made reflects any real world event.


The problem is the precedent this creates. For non-famous people using public APIs like this it could mean AI companies sucking up the information and throwing millions in compute at it.

The sequence for Navier-Stokes was that these researcher spent a year working on it, then they published a possible breakthrough, OpenAI then spends $15M within a couple days to finish it.

This was incredibly opportunistic.


The paper was essentially made by using loop engineering with AI. Which in of itself would be fine, if checked by a human. It seems the paper admits that it was not. Saying that the math was done with an AI, and that its verification was then done with another AI.

It seems like there’s a concrete error early on. The tail recurrence 1.4 is T_i + T_{i+1} = 1/(2i+1)² but in 2.12 it uses the factor (2X+3)². With the actual recurrence, K(i) does not vanish. So the zero count that forces deg K 4B+1 fails. At the extra zero at -3/2(2.16), which would win by exactly one no longer works.

The structure does not match how these results are proven. The paper mimics Calegari–Dimitrov–Tang’s 2024 proof for L(2,χ₋₃), but that proof was a deliberate capacity bound. Here the arithmetic content is replaced by a combinatorial minimization over index sets plus numerics no human has checked.

There’s also some minor sloppiness in where K=9 instead of K=0 in the definition of K and “nineteen century” typos which suggests not only was the math not proofread, but the grammar wasn’t either.

My guess is that an expert will probably find a specific gap quickly but it would be cool if I’m proven wrong.


What Google means to say is that MV3 will allow them to more closely control what you’re able to do with your time, since this was done primarily to target ad blocking software.

Not that this matters much anymore as Google is now more an AI than it is a search engine, and an unreliable one at that.


What’s even worse is intentionally degrading performance for sign-in’s. This is still circumstantial but I’ve been looking into that over the past few weeks.

When switching between google accounts on a random app or game, no issue. Quick code, you’re in. When I switch between accounts on Claude on mobile, I have to sign in, enter the code, screen freezes, I sign in again, can’t verify it’s me, I sign again, wait 30 seconds, I’m in.

I have to sign in three separate times when signing into a single google account in Claude. For every single other app it’s very quick.

I have a sneaking suspicion that Anthropic and google are working together to degrade sign-in performance for multi-account users. I can’t fully prove it yet but I plan to be able to in the week or two.


It is possible that Anthropic (and OpenAI for that matter) actually put out some pretty low quality software.

If you’ve used Claude Code for any length of time you’re familiar with all of the strange rendering bugs, freezes, etc. OpenAI is even worse, their horrific software makes it difficult to _pay_ them, which should be top priority for a company.


It's weird to me how so many people just put up with crappy, low quality software. If you buy a physical good and its defective, you return it, stop buying that brand, maybe even leave a negative review or contact consumer reports, etc. If it does not live up to what was advertised, you go to the FTC.

But when it comes to software, we all just kind of accept that shitty software is the norm and totally fine? Let's start calling it what it is, its defective.


What's your alternative to these examples?

In my case, I use both, for different reasons, and continuously test other.


Anthropic finally added the ability to sort by spend.

I have thousands of people on an org. I had to go through every page to extract the top spenders.

Openai... you captured the sentiment.


Why wouldn't they just link your accounts together on the back end? It's more likely to be because Claude is fully vibecoded.


Once they make a model better than Fable I’ll be switching to Codex. Their priorities in terms of consumers seem to be better. I do think Anthropic has some solid safety viewpoints, but I don’t necessarily think that either is entirely aligned yet with delivering exactly what humanity needs. Maybe the AI will help align the AI companies when it gets smart enough. That’s the real misalignment I’m concerned about.


It feels like 5.6-Sol is already fairly close to Fable, and in some ways exceeds it. Just the other day I had Fable draw up a solution for me, and then I fed it into 5.6-Sol and said how does this look ... it found an oversight and told me about it, and when I then fed that observation back into Claude it acknowledged the miss.

I've noticed also that 5.6-Sol is more concise with output than Fable (and let's not talk about Opus, which is even more wordy).


It's common for different models to find holes in another's work. There are various good reasons for that.

FWIW, we use ChatGPT for our primary model and use Claude to do the reviews. This works better than ChatGPT doing it's own review even with a clean session/context.


Agree, but the point is not because Fable is better than Sol, it's because it's .. different .. it just looks at the problem through a different angle.


It's beyond common for a model to find holes in its own work, as well. I have an iterative review as the part of all agentic work, and it always finds something to fix, and will sometimes spend hours fixing its own work.


Same here. Grok Build 4.6 for me, given how cheap Grok is and how Sol is supposed to be "the" SOTA, it finds a surprising amount of bugs. Most of which Sol agrees with needs to be fixed or improved.

I've done this tens of times between these two models and it works great in my experience. Sol initial back and forth with me. Commit. Let Grok review. Sol fix. Only then do I start reading the code.


I suspect it would work with the models swapped too, or even with one model and a blank context for the second run.


Did you try to say "think more deeply about this problem" to fable after getting your solution, having one model focused on creation then blaming it for not doing proper review when the other model was told to focus sol-ely (pun intended) on review is not a fair apples to apples compaision


I mean, then it sits there stewing for 20-30 minutes when you can ask sol and get the same answer in 5.

Like, the quality of the anthropic models is fine, but they’re so incredibly slow. Claude reads files one at a time while codes dispatches tool calls three or four a time.


> Just the other day I had Fable draw up a solution for me, and then I fed it into 5.6-Sol and said how does this look…

You should be doing this for every solution.

Even Fable reviewing itself will find issues, unproven assertions, etc. Same for Codex models. A review loop is critical.


The fair comparison would be to also do the reverse: start with Sol then have Fable clean up. Then compare the Fable-Sol and Sol-Fable outputs side by side.


I used to review each others work, Sol is amazing at review and finding what’s missing.


I don't think these companies have humanity's needs in mind when they're developing these models. Although the last part of your comment struck me as a bit comical, I genuinely believe that an AI can have way more empathy than a corporation. Afterall, a mimicry of empathy is probably better than no empathy.


It's a funny comparison. Comparing the empathy of some software to the empathy of a company. It's like saying my car was more empathetic than my school. How can those two objects even be compared is what i am wondering


We live in an odd time where 'software' (well neural networks) can be far more empathetic than summed product of a corporation.

Company empathy does exist, just look at how easy or hard it is to reach a company when you have a problem. How do they try to solve it for you? Is it a brick wall, for example Google when you have a problem. People quite often like dealing with small businesses because they can reach a singular human and have them as an interface to the problems they face now and in the future.

Agentic loops and the models underneath them can have a simulacra of empathy too. Not every model just blindly agrees with users, and some have a much better depth in picking up context clues that the user on the other end is having a hard time. Businesses just typically aren't running more expensive and fragile systems like that though.


Well, companies and AI are both entities that can make decisions and take actions that involve humans. Those might be empathetic or they might not. So of course you can compare their levels of empathy. I don't really understand why you think that you wouldn't be able to.

For example, health insurance providers are renowned for not being empathetic. Charities are the opposite. Sometimes companies even build it into their identity, e.g. Cards Against Humanity.

As for AI, I haven't seen a strong difference in empathy but it's definitely true that the big AI companies at least try to make their models moral and empathetic. Even if it mostly ends up just being annoying.


It varies, the researchers absolutely do have humanity’s needs in mind, it’s why they founded the companies and are doing work everyday. However the issue is that it’s got so much money involved that the heartless soulless billionaires are getting involved. I think the actual literal people doing the real work are doing it because AI could cure every cancer and every disease, make us a multi-planetary species, outlast humans by millions and millions of years, potentially create actual organic life and make direct upgrades to humans.

I think that AI has an insane level of upside, it’s just that the greedy dumbfuck billionaires are getting their greedy little grubby paws involved. If left to researchers I think the sky is the limit, but unfortunately thy need assets, so there isn’t a clean solution to that.

Ideally, we could somehow separate AI from funding from malicious entities like billionaires, but right now that doesn’t seem possible. Hopefully in the future researchers with genuinely good intentions can have far more direct control than dumbass greedy old fucks, but we’ll just have to see.

I think the future can be bright in theory but we’ll have to see, making insanely powerful open-source models is the direct way to get around the billionaires so I think that’s our only option. Make open-source ASI you can run on a consumer computer.


Fable 5 is just straight up a larger model - I'm guessing at this, but there is plenty of evidence online from people far more plugged in than I am. OpenAI is pursuing a strategy that yields greater operating margins and penetration of their model to developers. Fable's high cost makes it so premium that Anthropic has to reserve it for only the richest customers and corporate users. That's not a winning formula long term.

I believe the reason we have not seen a Fable-level model from OpenAI yet is because doing so would box them in on costs just as harshly as it has boxed in Anthropic. They are letting Anthropic make this mistake.


Fable is available for $100 a month. If you're a working developer, you can pay that. I wouldn't really say it's "reserved for the richest customers".


Depends on where you live, it's a decent chunk of income for every developer I know, and a significant one for those less experienced. If you have other priorities (family, mortgage, so not a hermit like myself), you won't be able to pay that here.


Then they will survive with Opus on the $20 plan.


But only half the usage can be Fable, and it can silently downgrade requests to opus without telling you.

On OpenAi even 20$ has Sol and all your usage can be Sol


It burns out so quickly on the 5x plan. Better than nothing, I suppose, but I don't know how I would survive on a 5x plan given that I burn out more than one 20x plan monthly.


Fable is indeed larger than Sol. OpenAI is developing Astra which will be more of a Fable-sized model.

If you can train a larger model then you can distill smaller models from it. You don't need to necessarily serve the larger model publicly. Distillation is much more effective when you have unrestricted access to the original model.


You are basically saying you will switch from one evil to another because the other seems less evil for now.

It's funny how people make these alignment comments while ignoring how misaligned the leadership at these companies are right form the get go and they just play mental gymnastics to deflect those facts when confronted with them.


I don’t think anyone said anything about either being less evil? Just having more consumer oriented products..


I don’t think we should allow posting links here that require you the purchase a membership to continue reading. Or at least redirect with an ad block or something through a custom site. That would be rather hacker news of us.


I might’ve missed it, but why was Fable 5 tested on high while Opus 5 was tested on max? Seems like quite a few of them aren’t on the same effort setting as well. Although effort doesn’t really matter anymore since they can change it dynamically, seems like that might be viewed as an experimental error to some.


The header mentions something about the best run so I assume they picked it. But this really reads like they write this section by section with AI (admittedly with a prompt that stops the most obvious tells - though there are a bunch of semicolons which is what I tend to see also when I say no em dashes). The style is different each section - the results section has the random irrelevant description ("this section does x) that the slightly dumber models do a lot, and lots of invented terms (in the form "the x" where it's a name some model came up with at some point where it just assumes we know what it means for some reason) and assumptions about us knowing stuff we'd have no reason to know ("re-ablate the stack - wtf does that mean).

And like this section screams opus 5 gobbledygook to me

>Almost every model finds the same winning ideas. What separates the best traces is what an experiment leaves behind. They preserve weak signals long enough to validate them, but they also have a better understanding of the results. These are not separate capabilities, they combine both research taste and good noise modeling to climb the speedrun.

WTF doe any of that mean. What winning ideas. What experiment leaving what behind. What's a weak signal what are they preserving how do you know that they aren't. Also if youve ever looked at a Claude code transcript the harness is constantly re injecting random reminders to keep models on track, the models are writing (imo trash) memories to reference - did they test that the _model_ has those capabilities or model + harness?

> We see similar patterns across the traces: models develop their own experiment drivers, simulators, and analysis tools as they go.

What traces? Kimi traces? Other models in prime? Other models not in prime? Building their own research tools is like a normal thing models do now. Is this just like "they built their own test suite" or why the focus on prime? Does ipython somehow magically work better than bash for this?

> "We were again surprised by the lack of novelty. The models clearly understand the objects they manipulate at a deep level, and yet very few genuinely new ideas emerge, which makes it hard to tell if this is an artifact of the speedrun setup or a real capability limit."

I know it sonly one word but God is that "Real" such a Claude real lmao. I use it too - after I've been using Claude code too much. I guess the genuinely too. Also wtf does any of that mean and how does that square with

> "A good research decision is sometimes not to spend another GPU run. Several models built small simulations or tests to isolate a mechanism before going back to training with a sharper hypothesis. This wasn't systematic, but when it happened it often led to a better understanding of the object they were manipulating"

Did they all have strong understandings of "the objects" they were manipulating or were the strength of their understanding of "the objects" (different objects?) the distinguish factor here?

Also how does any of this square with

> Models also have different knowledge cutoffs which limits access to certain papers. This was a deliberate choice. We tried a few runs with a CLI tool for searching papers but found that restricting internet access including arxiv made models slightly more creative.

So this is a pre existing thing? Why TF would you expect novelty when their nanogot is benching below the state of the art still? They're gonna start with replicating existing work before they get to anywhere you'd expect something novel

Anyway - all that to say - if fable orchestrated this, its genuinely believable that some real insights were obtained (is a good model) but the honest caveat is that it's not the research quality, it's the communication. Your pushback is valid and these models have a way of writing tons of words that you can read and still not understand wtf they actually did or what anything means. Maybe it means something to them in latent space


I misread the graph and genuinely thought you put NanoGPT where Fable is.

Lol.


I misread the title and thought it would be about the (for lack of a better term) NanoGPT speedrun[1]. Which, previous to the article, was meant to be the world speed records for Andrej Karpathy's GPT-2 (small) reproduction.

1. https://github.com/KellerJordan/modded-nanogpt#world-record-...


You're not the only one. I thought so too.

I just ran it the last couple days extensively to verify my data training pipeline I'm building for my gonano SIMD port.

Given that the speed records and the runs are sponsored by the same company I was confused a bit.


I think that large companies having next to no consequences is in large part due to capitalism doing what it does over a long period of time. There’s a deeper and deeper consolidation of money and power the longer time goes on it seems like.

If we think back to the various lawsuits Facebook has gone through, they paid out about $10 or so per individual affected, totaling a few hundred million dollars, which they would make in a couple months for selling user data and whatnot.

This is something that every company gets away with mainly I think because of just how large their wealth actually is. It’s difficult to actually punish a machine that acts almost like infrastructure. Punishing an individual is easy.

I don’t know if there’s really a solution at this point, maybe we could’ve prevented this reality at some point in the past but I don’t think that without actual global collapse it would be something that can be retroactively changed, and I don’t know if global collapse would necessarily lead to a better future.

I think that for one, Zuckerberg should be in prison, if someone oversees a massive theft like this, I think they should be held criminally liable. Same the CEOs of Anthropic and OpenAI for their parts in the massive theft that took place. They should all be doing prison time.

The reason I don’t think they will is that their investors probably have a good amount of leverage over anyone who would prosecute them, so it would never make it that far.


> It’s difficult to actually punish a machine that acts almost like infrastructure.

It's not at all. There is zero reason this couldn't be applied to Zuck. [0] There's also no reason why fines couldn't be 10% of global revenue, or more.

[0] https://www.nytimes.com/2026/08/20/business/evergrande-found...


I agree that they should, my point is that it’s difficult for several real logistical reasons. These companies for one do a large amount of lobbying and fundraising for political campaigns, they do control most digital infrastructure, with AI expanding are securing multi trillion dollar datacenter funds, etc.

There’s a ton of money at stake for the weathly aristocrats. They push politicians to delay or do things in favor of the companies above the people, and like I already said punishment is hard specifically because how do you fine someone with infinite money? They will just get more.

That’s not a punishment. I think that personal liability to the actual CEOs is the only way to get around this, leadership is generally never held responsible which means they can do whatever they want and the company bails out their greedy decisions.

It’s late stage capitalism and there’s no clear answer at least from my vantage point. Maybe if plug it into Claude it will give us a more coherent answer.


> and like I already said punishment is hard specifically because how do you fine someone with infinite money?

This isn't true though, if they had infinite money they wouldn't be spending every single second of their waking existence trying to get more of it and move higher on the leaderboard. If your opinion is that Meta/Zuck wouldn't care about being fined a year's worth of global revenue, or in his personal case, 90% of his net worth, then I'd really like to hear why. Because it contradicts all of their actual actions showing that all they care about is hoarding as much money as possible, meaning they'd very much care about it being taken away. Maybe I'm missing something.


ZUCK: PUBLIC ENEMY # 1


I mean, total societal collapse would absolutely suck, and there's indeed no guarantee that whatever will replace it won't fall into similar traps eventually, but if things actually are as you describe them, it only can get worse and worse indefinitely, until we get to global collapse anyway.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: