"extremely severe" - aren't we laying it on a bit thick here? The whole impact seems to be that users who have telemetry turned off got this feature ~2 days later, when the bug was noticed.
a "feature" that every user has asked for - which is the equivalent of changing claude.md to agents.md - and the release isn't even smooth because the telemetry isn't wired up quite right.
for an organization that is being used as a model for new agentic software development practices.... and every software exec on earth is trying to reshape their organizations after - its a pretty stupid bug for a feature that should've been straightforward in the first place + took forever for them to get around to.
its just kind of emblamatic of the rough edges that exist EVEN FOR SIMPLE THINGS whenever human judgement is totally removed the equation.
Or it’s as simple as “hey, they don’t want telemetry? Cool, we’ll turn off all the extra phoning home” and their feature flag code uses the same paths as telemetry because it’s the same endpoint, etc.
Take a look at https://github.com/leanprover/comparator which was used to verify the result. It's of course not impossible that they're hitting some bug, but way harder than one would intuitively think. For starters, they'd have to hit two bugs in two independently written Lean kernels.
In what world is peer review a higher standard than formal verification in Lean?
Not in the world mathematicians have been living in for the past decades at least. Nearly all big theorems that have been formalized so far had been published beforehand, and it was usually regarded as a step up in rigor. Wrong results get published in peer reviewed journals all the time.
I find it really strange the way “peer-reviewed” is used by the general public as some gold standard of truth. As a former academic who has been there, the process is extremely arbitrary and variable. Are people aware that the “peer” refers not to a community or a committee, but literally to one random guy or maybe a couple with zero accreditation? And that the journal editor can do whatever they want with this peer’s “review” including completely ignoring it?
Admittedly, peer review originally referred to the fact that the journal being published was read ("reviewed") by your peers, and thus there would be the opportunity for them to provide rebuttals after publication. The process the term now refers to in academia only emerged in the 50's and 60's (Nature for example wasn't "peer reviewed" until 1967). It was literally a marketing ploy to make journals sound more prestigious which got legitimized when certain grant agencies and regulatory bodies started including it in their requirements. All the best science is published in journals that describe themselves as peer reviewed, the government says science not published in such journals isn't up to snuff, and taken literally it seems useful. It's no wonder the public thinks it's an important and long established part of the scientific process.
"I find it really strange the way “peer-reviewed” is used by the general public as some gold standard of truth."
OK but this member of the general public has at least subscribed to New Scientist since 1987, nine O levels, two A levels, two AS levels and a HND in Civ Eng. All pretty mediocre but I have a fair idea on how sciencing is supposed to work and how it ... actually works. Obviously, I ended up in IT.
I should also point out that maths "peer reviewed" is a bit special. For example Mr Wiles went through quite a maelstrom before his proof of some dodgy marginalia was accepted as "true".
I feel like there's a pretty huge difference between using inputs and outputs as part of a general training corpus, and looking at a specific users workspace after hearing rumours and yoinking their ideas to beat them to the point.
Unless OpenAI finished a whole new training run on the latest data in the last few days, the possible allegation seems to be the latter.
They have been collaborating on this solution for a year, and Astra was trained in February this year so it’s entirely possible the direction of their research was in the training corpus.
The nice thing is, once all of these proofs are formalized in a machine-checkable language, it should be relatively straightforward to translate the corpus between different languages, if someone finds something with a nicer syntax.
Uplift on "random dictionary words" (excluding 'stop words' & proper nouns) was ~20%; uplift on "random words from my local epub library" (same exclusion rules) was ~40%.
The random words from my local epub library (leans toward postmodern fiction) were definitely more evocative than the dictionary words when I eyeballed them.
I randomised each turn but kept the story prompt request the same across control, dictionary, personal library.
I must stress that I'm not claiming scientific method or certainty here - just sharing an approach that seemed to work well enough for me, and seemed like a reasonable conclusion: introduce noise, get more interesting output.
I haven't done the math but I think you'd need a much larger sample size than 1k per category to prove the uplift!
I think you're vastly overestimating the average persons ability to use Blender if you can do that in an afternoon; just figuring out how to place a colored cube and the camera probably takes an afternoon if you pick up Blender for the first time.
Yeah, I've bounced off Blender twice now. And I've written a (basic) 3D modeller.
I think part of the problem is that pretty much all the tutorial material for Blender seems to be in video form, which is easily my least effective way to learn, even leaving aside the "I've only got one screen" issue.
And knowing these little tricks to get what you want with image generation models also takes time. Not to mention you need some knowledge on some other software just to make the underlying layout.
reply