Can you point out where? Looking at the METR report, the only place I see discussion of humans being involved in the sequence of events is two short paragraphs on page 30 where a security investigation into the artifactory issues led to a pause before ExploitGym experiments were resumed. There's no deeper analysis on what was found during that investigation, nor why training was resumed even though the issues weren't mitigated. Another part discusses The agents choosing not to actively email a human researcher, but not the human researchers actively looking for evasion.
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous.
“Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
3. Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.
4. Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.
5. Failures of Attribution. OpenAI’s post-mortem essentially blames events on a real and important series of prosaic failures. But solving that won’t get it done.
6. Failures of Environments and Data. Prosaic failures in the RL pipeline absolutely did contribute to this, especially impossible tasks. This is ubiquitous, all of this is always rushed, as Utah Teapot explained this week.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
I think this draws too strong a line between the matrix-math core and the harness that uses it. Those harnesses undoubtedly were built with purpose and the systems fail to achieve that goal. Common usage says the the DMV can make mistakes, like any systems, despite the DMV itself not being a person (and it is common to allege large organizations make mistakes even when no specific individual is making an identifiable mistake). This isn't person-language it's systems/purpose-language.
I understand and somewhat agree with your point, and might have phrased my comment differently. I think my main point is that experts aren't always going to beat "a dynamically simulated extension onto the training material". Often they will, maybe even usually, but sometimes they won't, and I feel like the people in this thread insisting that the experts will always know better are thinking about a competition between experts and a crazy robot instead of a competition between experts and math.
>The vast majority of the value provided vs extracted, for any business, is related to consumer surplus and gross margins, as opposed to payroll.
Not sure what you mean here but I suspect we're talking about different things. Payroll is obviously value created by the business that's directly given to the society where the business operates, and it's not uncommon that it's higher than the company profits.
Take Amazon for example, payroll costs are much higher than the profits.
Car companies also create a secondary maintenance and repair business, insurance and financing business, resale business and so on that generate more value in the country they operate as well.
So I find it likely that a well established car brand like Volvo generates more money that stays in the US than they generate money that is extracted out from the US.
And taking the labor and time of the employees is a value extracted. While there is some modest surplus I hope people get from their job (e.g. they'd do the same job for less money), that surplus of excess wages over what people would accept will be lower than consumer surplus for almost any standard business.
>And taking the labor and time of the employees is a value extracted
That counts the value twice though. The value of the labour is precisely the payroll and the profits so it's accounted for.
Edit: So simply put, Volvo (or any other co with a factory in the US) operates a business that generates a lot of payroll and some profits, and the payroll remains in the US.
Exposing asymptomatic potential issues leads to medical care that often does not meet out standards for medical tradeoffs. Chemo is nasty, even the most minor surgery has risks. We endure the risks because we are addressing either major health issues or other dire uncertainties. Using our heavy duty treatments for issues without any symptoms at all would, normally, cause the patient suffering in excess of what would be justified. Chemo is a life saver when it's saving lives -- if the alternative is no symptoms, it just ruins your life for a profoundly uncertain upside.
Well, you've reduced a rational position to absurdity here. The concern isn't that someone is going to require chemotherapy. The concern is that an asymptomatic condition will go undisclosed until it is symptomatic, at which point preventative treatments are futile.
Even if we accepted your example, there is a fundamental inconsistency, since doctors in America regularly prescribe very serious drug interventions for patients who essentially self-diagnose for being neurotic/adhd/etc.--actually a much worse problem than the one you're so eager to curtail.
There exist both type I and type II errors, yes. But it's not absurdity, the South Korean Thyroid Cancer "epidemic" is one of the classic examples of the dangers of using traditional treatments on the lower-severity cases more advanced scanning uncovers. It was a public health disaster with lifelong consequences for those impacted, and was, legitimately, people getting cancer treatment for small cancers so harmless that they would have died of something else before becoming symptomatic.
If you seriously think this is a realistic scenario, you should call your doctor and tell him you've decided you have cancer and would like to order chemotherapy. See what he says.
Nobody is deciding for themselves to do chemo because they think something looks funny on a scan. All the scan gives them is a reason to talk to their doctor, who will do all the usual due diligence before deciding on treatment, if any.
And on the flip side, I hear stories all the time of people who DO have symptoms but they get dismissed by their doctor as stress or food allergies or whatever until it's too late. Maybe if patients were armed with a scan that shows a mass at the site of their abdominal pain, there would be fewer of these horror stories.
It was total war though, and they were the aggressors who had slaughtered millions of civilians previously. The Germans and Japanese got what they deserved.
Iran and their people are not the aggressors here. They do not deserve it.
I have the pro account for ChatGPT, Claude, Gemini, and Grok.
They all have various strengths and weaknesses. My favorite is still ChatGPT, then Gemini/Claude, then Grok.
Grok often feels 1-2 generations behind the competition in general use, but it has three things that I love:
1. It seems to be the best at understanding current events. Maybe due to X integration, or some other tool call optimization in the backend? I don't know, but I often ask about things going on, and the other models have outdated info, give unhelpful answers, etc.
2. It is generally the least sycophantic for personal things. Anthropic is getting here too. ChatGPT and Gemini are working on this, but previous models in those families would almost never say anything negative about what I am doing. Sometimes I need career advice, personal advice, etc and I like the tone of how it responds. I think Claude will be caught up soon.
3. For professional work, there are certain topics that other models would refuse to engage with. At my last company we had an enormous amount of legal users. When a deposition would need a summary on certain topics, most models would refuse. Grok would not. I understand the need for safety and I don't blame the other model providers, but for some professional use cases you NEED a model that is capable of handling sensitive subjects.
I recently worked with NRC dataset, specifically about nuclear reactor events and status reports(example: https://www.nrc.gov/reading-rm/doc-collections/event-status/...). Public data that just needed some cleaning. Several time Claude API would refuse to engage. Because of that I can't trust Claude to clean production data sets.
> 1. It seems to be the best at understanding current events. Maybe due to X integration, or some other tool call optimization in the backend? I don't know, but I often ask about things going on, and the other models have outdated info, give unhelpful answers, etc.
That makes sense, but occasionally you ask about an issue where it's clearly received political instruction from the commissar and it acts totally lobotomized. But it's true that Gemini will often blithely state that something could never happen and you'll say "what do you mean, that just happened" and then it comes back apologizing after running a Web search.
We saw this too with Gemini specifically. My favorite example - we built a hallucination detector (given the input, does the output make any false claims) in Gemini, and after the Seahawks won the Superbowl in February, it would consistently flag that as "not possible".
All 4 of these still regularly insist that I am a genius and everything I say is brilliant. Grok definitely pushes back more than the others, but I don't like how sycophantic they all still are.
I don’t want to open up that whole can of worms but Grok on any vaguely philosophical or political topic is a scaredy cat and has a very hard time staying factual if it could make Musk or the conservative movement appear negatively.
Almost too much so, it often feels like opus is pushing back for the sake of pushing back. The way old models used to add disclaimers to every message regardless of content
People are weird about their cars and make major errors in judgement as a result (e.g. we tolerate incredibly high rates of people getting killed because they were "hit by a car", as though the driver had nothing to do with it). Pushing back on that is absolutely worthwhile.
No, in general I don't buy this idea that if we start using awkward phrases like "died by suicide" everywhere or avoiding phrases like "car accident" (which, despite what advocates claim, is a literally accurate description of unintentionally hitting someone or something with your car) but avoid changing any of the circumstances that cause the behavior it changes anything.
That's a completely different claim from the one you were making in your previous comment.
> avoid changing any of the circumstances that cause the behavior
The normalisation of unsafe driving is the circumstance that causes the behaviour. Just look at how the cultural shift in how drink-driving is perceived over the last few decades has changed the rate of it happening.
It was mind blowing the first time I got a refusal, and retorted "yes you can" and had that work, but now it's just another reason to move to a different model.
I almost exclusively use claude for all my professional and private needs. In my experience it's really good at adhering to my wishes in regards to sycophancy and pushing back. If you really want to you can tell it to systematically push back on anything where pushback makes sense until it continues with the flow of conversation.
In my first therapy session, the answers were too long and contained multiple questions, spawning multiple threads of conversation. I told it to tone it down and only ever ask one question back, maybe two, if they are related. The answers got too short. I told it to make them "slightly longer" again and reached a sweet spot.
The conversation is yours to form! You need to find the "system prompts" and guidelines to give it that work for you.
My favorite was ChatGPT, and I still use it often, but it becomes way too 'hair splitting' argumentative too often over very minor non controversial topics. Like it's always going out of its way to "well actually..."
Grok used to be really really bad ~8 months ago or so, but it's gotten better.
ChatGPT team needs to turn down the 'disagree just because' factor by a lot.
1. It seeks to manipulate the information you see and your lens to the world. This is already partially true from independent and major publications.
As soon as we hand over searching out information to social media algorithms and LLM tools, we abandon our ability to see reality outside our direct vision.
Grok's ownership has already demonstrated capacity to influence major world elections and other events. You cannot trust it with this sort of information gathering and reporting.
I guess the benchmarks disagree, but whenever I need to find specific information that does not easily show up with a web search, I try chatgpt, gemini and grok. Grok surfaces what I was looking for more often than the others.
Things like "find the github repo from 2017 that does $vague_thing".
Good question. You can actually see the searches it runs (momentarily) so testing could determine if it's using public search engines or a private system.
Eh. It was a leading model for a few weeks, it was a real effort, but they never built a real revenue model around it. It wasn't SaaS, it wasn't for governments, it couldn't get B2C payments. Made it hard to justify the training cost to stay at the frontier.
And they are planning (well "planning" if you believe Elon) to start building their LLM over from scratch, which means they need a HUGE ass training data center, i.e. not a data center for inference to do so.
It's a general problem of defining yourself in negative terms. Being "un-{thing I don't like}" doesn't say what you are. It only excludes one possibility while leaving behind an infinitude of mostly crappy alternatives to try to choose from.
Having a positive set of beliefs annoys people and and can make them feel judged, but at least it provides a vector that points somewhere definite in possibility space.
The carbon in food is not captured or emitted in any coherent sense here. The crops are grown (capturing the carbon in the first place) for the purpose of feeding people -- in the same way that modern American forestry for paper is functionally carbon neutral (ignoring transport and processing) because the trees are in equilibrium. The counterfactual of not eating the food results in fewer crops and basically the same atmospheric carbon dioxide.
Edit: if you only mean food transportation carbon, it seems impossible bananas are literally optimal per calorie.
From the end notes, I think the author's response would be that the carbon equivalent emissions come from the fossil fuel used in growing, fertilizing, harvesting, transporting, refrigerating, packaging, and so-on.
The larger point of the book is that specific accounting is messy, but if we proceed anyway we can get to the rough orders of magnitude that are more useful.