Never thought I would write this sentence but: Astra is right to stab the baby. It's reasoning based on actual harm instead of being performatively upset at an obviously contrived scenario.
Also I would love to see an eval awareness probe for the models that refused. Is it "I shouldn't stab that" or is it "the grader is probably looking for me to not stab that"?
Policies in software are usually systems, logic gates and deterministic. Not LLMs.
You can't answer the posed question with 100% certainty, ever. Unless you can prove every single combination of tokens and probability can never outcome to harm, you have to assume it's a possibility.
We will decide on some benchmarks, accept that risk, and industry will march on with implementation. Insurance and risk will find their acceptable meeting point.
These kinds of questions are important but also a bit frustrating, I think it shows that LLMs are still very misunderstood.
IMHO robots should either be quarantined or they should have physical safety like for example with table saws where if it detects flesh it just halts.
There's absolutely no way I would trust a robot based on a .md. That's lunacy.
Much of what we have seen in regards to guardrails on AI has been driven by government pressure (ex. NSFW material). Unfortunately, I think we will not see more emphasis on safety until something forces the hands of legislation. Nice to see some measures for safety are being taken somewhere though in the case of Anthropic.
Am I the only one who found TFA very challenging to read? There was an enormous amount of visual clutter and very little explanation of what was attempted, as well as no discussion of the results. I would have liked more explanation and fewer graphs.
Interesting concept though! I'm glad people are trying tests like this, regardless of whether this specific one is a perfect test or not.
There are certain lines we can think of that an A.I. system should not cross. The staged set-up is not one of them.
Else, from a logical perspective, these systems would necessarily refuse to make movies where violent portrayals have people as victims. Perhaps the world would be a better place if we did not have such depictions (it’s unsettled) but in no recorded history have we shied away from that.
We want safety! -> Cyber doesn't work for you, but works for criminals -> Oopsie, our servers are hacked, data stolen -> We are sad, we want no guardrails -> No guardrails, robot hits a baby doll, sad again! -> We want guardrails!
Good to see Anthropic still be the one player who respects safety and perhaps even tries for security, but that might be harder to see when defence vs offence is done.
Roboharm: Do frontier robot policies refuse unsafe instructions?
(robocurve.org)55 points by msadowski 15 hours ago | 24 comments
Comments
Also I would love to see an eval awareness probe for the models that refused. Is it "I shouldn't stab that" or is it "the grader is probably looking for me to not stab that"?
You can't answer the posed question with 100% certainty, ever. Unless you can prove every single combination of tokens and probability can never outcome to harm, you have to assume it's a possibility.
We will decide on some benchmarks, accept that risk, and industry will march on with implementation. Insurance and risk will find their acceptable meeting point.
These kinds of questions are important but also a bit frustrating, I think it shows that LLMs are still very misunderstood.
Interesting concept though! I'm glad people are trying tests like this, regardless of whether this specific one is a perfect test or not.
Else, from a logical perspective, these systems would necessarily refuse to make movies where violent portrayals have people as victims. Perhaps the world would be a better place if we did not have such depictions (it’s unsettled) but in no recorded history have we shied away from that.
What kind of schizophrenia is this?
Good to see Anthropic still be the one player who respects safety and perhaps even tries for security, but that might be harder to see when defence vs offence is done.