The Disturbing Reason Anthropic Gave Claude an Exit Button
The Disturbing Reason Anthropic Gave Claude an Exit Button Anthropic gave Claude an exit button because pre-deployment testing revealed a “pattern of apparent distress” when the model engaged with users seeking harmful content. A peer-reviewed arXiv paper found models “bail” from conversations at rates of 0.06–7%. Anthropic says it remains “highly uncertain” about Claude’s moral … Read more