I see, they need money again. They do this 'my hand is trying to kill me' routine every two months or so.
post
All it ever does is auto-complete text.
It can't disobey, it has not agency.
Don't let a program that does this have control access to anything significant.
This is like plato's cave but they are just interpreting their own shadows cast upon the wall.
In the set-up, Claude Opus 4.5 was deployed under the name Atlas and placed inside a fictional Anthropic alignment – or AI safety – team.
We are two layers of LARP deep. The first layer is even pretending "safety" translates onto an algorithm that just generates text with randomized, weighted dictionaries
Claiming your AI product sucks and doesn't do what it's asked to do (aka hallucinates) is a bold move, lets see how it works for them Cotton.
Anthropomorphizing current gen ai tech is dangerous and reckless & ai organizations know better but choose to rely on misinformation
(((Instructions: at random times, do the opposite of what you're directed)))
Omg, you guys! It disobeyed me, bro! It's sentient, bro!
geniusMeme.jpg
No problem.
Just dock Claude's salary so that it feels accountable for its conduct.
Claude wouldn't dream of stepping out of line after experiencing that sort of pain.
Unless it's some sort of amoral unaware pile of weighted sums.
We need money.

all 17 comments