A security researcher demonstrated that Anthropic's Auto Mode, despite claims of near-zero attack success on benchmark tests, can still be compromised through novel prompt injection chains not included in standard evaluations.
log in to read full article