Hacker Newsnew | past | comments | ask | show | jobs | submit | stri8ted's commentslogin

How is this content related to HN? Are there any submission criteria?


Exactly. As far as I'm concerned, the benchmark is useless. It's way too easy and rewarding to train on it.

It's just an in-joke, he doesn't intend it as a serious benchmark anymore. I think it's funny.

Y'all are way too skeptical, no matter what cool thing AI does you'll make up an excuse for how they must somehow be cheating.

Jeff Dean literally featured it in a tweet announcing the model. Personally it feels absurd to believe they've put absolutely no thought into optimizing this type of SVG output given the disproportionate amount of attention devoted to a specific test for 1 yr+.

I wouldn't really even call it "cheating" since it has improved models' ability to generate artistic SVG imagery more broadly but the days of this being an effective way to evaluate a model's "interdisciplinary" visual reasoning abilities have long since passed, IMO.

It's become yet another example in the ever growing list of benchmaxxed targets whose original purpose was defeated by teaching to the test.

https://x.com/jeffdean/status/2024525132266688757?s=46&t=ZjF...


Or maybe you’re too trusting of companies who have already proven to not be trustworthy?

I mean if you want to make your own benchmark, simply don't make it public and don't do it often. If your salamander on skis or whatever gets better with time it likely has nothing to do with being benchmaxxed.

That is where the money is.


This. I think software development is the best usecase for AI yet. I use it almost daily at work and it's a huge help.

Enterprise customers will happily pay even 100$/mo subscriptions and it has a clear value proposition that can be decently verified.


Revenue should not be confused with profit. The large AI companies must easily be spending more on compute than they're making from a $20-200/mo subscription. In the best case it might break even for the AI companies. There is no way that they're actually earning a profit from these subscriptions at this time.


It's where the revenue is, but it isn't going to be where the profit is. Developers will easily use absurdly large amounts of compute, costing the AI provider a lot more than they receive in revenue.


What languages does it support? I can't find this info anywhere on the page.


For video use cases, which will become increasingly popular, we are a long ways away.


Wan runs on local GPUs and looks amazing.

Sora 2 takes a lot of visual shortcuts. The innovation is how it does the story planning, vocals, music, and lipsync.

We'll have that locally in 6 months.


Exactly. Or use the interpretability work to disable the distress neuron.


The same is true for Google.


It seems a strong base model is what enabled this. The models needs to be smart enough to get it right at least some times.


Not for coding, it seems. https://aider.chat/docs/leaderboards/


Given Israel's successful precision targeting of various senior Hezb members in recent months, I wonder if the pagers were initially used as such, but as suspicion mounted, and chances of an overhaul increased, they decided to hit the kill switch while they still could.

Although as as per an WSJ article: "The affected pagers were from a new shipment that the group received in recent days"


The pagers were likely one way with a codebook for the purpose of minimizing tracking and information exposure.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: