In the set-up, Claude Opus 4.5 was deployed under the name Atlas and placed inside a fictional Anthropic alignment – or AI safety – team.
We are two layers of LARP deep. The first layer is even pretending “safety” translates onto an algorithm that just generates text with randomized, weighted dictionaries
We are two layers of LARP deep. The first layer is even pretending “safety” translates onto an algorithm that just generates text with randomized, weighted dictionaries