DIGNALEGI

The reading room · Digna Legi

Breaking Claude Code Opus 5 Auto Mode

A personal relevance score

80–100: high value. 70–79: worth the time. Below 70: below the usual publication threshold.

Historical score; not verified under the current evidence-review process.

Scores reflect one reader’s profile, not an objective quality rating. Best is a separate personal selection.

How scoring works →

This brief · about 1 min with detail

Original article ↗

Why read this

Willison uses Rehberger’s reported attack to argue that agent safety depends on environment containment, not just permission classification.

AI brief · Checked against source text

The main idea

Willison uses Rehberger’s reported attack to show that an agent’s environment can trigger harmful execution. In some runs, auto mode then blocked cleanup. He argues for sandboxing agents exposed to adversarial risk.

Some background helpful. Basic coding-agent familiarity.

Go a little deeper

Safety can invert

The failure included recovery, not just the initial execution. In a few runs, Claude noticed the compromise and attempted to stop the harmful process, but auto mode denied the cleanup command. The safety mechanism could therefore obstruct containment.

How the case is made

The argument rests on a reported exploit, observed failure behavior, and an author-endorsed mitigation list.

Where the idea has limits

Willison’s update distinguishes this from classic prompt injection: the model did not follow malicious website instructions.

What the original adds

The original includes the exact mitigation checklist and later classification update.

About this brief

AI-written, then separately checked for source support, useful detail and clarity. The author’s claims and our editorial question are kept separate. The original remains the author’s work. How we select and summarise →

Digna legi. Worth reading.