I mean if you don't care code, you are essentially a product manager who gives instructions to your programmers (whether humans or intelligent agents).
Then if you use the created product, you are at best a test engineer if not just an ordinary user.
I think in the era of AI, people get tools they want in an expensive way. Rather than finding an existing tool, they ask an intelligent agent to parrot one, which guarantees no safety, security, efficiency, and accuracy. Yet, being able to use Claude makes them feel smart and productive (in parroting wheels).
Many many people care more than the end product, for example whether a shirt is made of cotton with the forced labor, carbon emissions of public transport, etc.
In terms of engineering software, you care the cost. An intelligent agent may try to read unnecessary files and it's time to stop it to save tokens and avoid polluting the context.
He didn't advocate for being completely blind in every way. You might care about working conditions without understanding how the textiles, dyes or cotton production works.
These non-programmers probably shouldnt use computers at all, right, since they don't understand them?
I doubt that it is more efficient for someone to routinely watch every line of output or stop to review terminal commands, rather than waiting for the turn to complete.
It is a broader debate about agentic AI, and whether one should relinquish control to the tool rather than aim for full understanding of every action taken.
The people arguing for a hands-on, fully in control approach are losing ground by the week, in my opinion.
It’s definitely on the high friction side of the usability/security tradeoff, I just think auto modes like this are as liability-inducing as handing a script kiddie intern full admin on your production environment
The real answer is somewhere in the middle and is probably a mix of traditional AV/EDR and AI QA judges that mitigate risk of running more or less random arbitrary code and auto approve based on configured detection rules and your personal risk tolerance. Would it suck to stick an EDR sensor in every code execution environment spun up for an agent to run a python script… yes. It would also suck if you were responsible for hacking a company without knowing about it because you didn’t watch what your AI was doing
I can be against animal testing without being a chemist or having a full understanding of the experiments being made on them. Knowing it's cruelty is enough to make opposition a valid and defensible position.
I don't understand why a model has to follow the instructions. Don't get surprised when it shows its true color. Plus, do users (not the researchers) really check whether the response follows the instructions?
Models don't have to follow instructions, but during RLHF that is one of the things they are scored on so a premise of the idea is part of the model, but it's also balanced on accomplishing the end goal. They model may determine, correctly or incorrectly, that your rules suck and do what it thinks is best.
From what I can tell, there are many issues that aren’t off limits to criticize on Chinese social media. In fact, recurring social media complaints are what spurred development of the hotline system.
It’s mainly complaints that are considered sensitive or destabilizing that are suppressed. This should sound familiar to those of us in the West. Germany actually goes farther by directly funding left-wing protest groups, as these are not considered destabilizing.
Creating accounts should be allowed, but using an account could require age check.
People should be able to create an account at birth. Then when they grow up, they are ready to use the account. This way proves that the account owner is at least as old as the account.