Claude Code: Warum manuelle Befehlsfreigaben scheitern und wie Anthropic nachbessert

X-PostShoopyNews

Boris Cherny, Entwickler von Claude Code bei Anthropic, berichtet von Tests, bei denen Nutzer fast jeden manipulierten destruktiven Befehl reflexartig absegneten. Als Gegenmaßnahme setzt Anthropic auf ein zweites Modell ohne Dialogkontext sowie auf neuronales Monitoring gegen Prompt Injections.

Das Wichtigste

  1. In internen Tests winkten Entwickler gefährliche Befehle im Bestätigungsdialog von Claude Code fast ausnahmslos durch.
  2. Cherny bezeichnet klassische 'Befehl erlauben?'-Dialoge als Sicherheits-Theater, da Nutzer durch Freigabe-Ermüdung abstumpfen.
  3. Ein zweites, isoliertes Modell ohne Kenntnis des Chat-Verlaufs prüft nun jeden Befehl, um Manipulationsversuche im Dialog auszuhebeln.
  4. Zusätzlich überwacht Anthropic Neuronen-Aktivierungen auf Prompt Injections; eine 20.000-Dollar-Challenge über eine Woche blieb ohne erfolgreichen Exploit.
  5. Ergänzend empfiehlt der Autor Graph Engineering, um Agenten-Workflows über spezialisierte Knoten und strikte Prüf-Gates zu strukturieren.

Warum das relevant ist

Der Human-in-the-Loop-Ansatz versagt bei Coding-Agenten oft an menschlicher Ermüdung: Bestätigungsdialoge werden reflexartig bestätigt. Echte Absicherung erfordert automatisierte, kontextfreie Prüfinstanzen statt trügerischer Klick-Schranken.

Einordnung

Die Beobachtungen von Cherny offenbaren eine zentrale Schwachstelle moderner Coding-Assistenten. Sobald die Frequenz von Sicherheitsabfragen steigt, mutieren Freigabefenster zu reinen Klickhürden. Der gewählte Lösungsansatz – ein zweites Modell ohne Kontext – verhindert, dass der Validator durch vorausgegangenes Social Engineering im Prompt beeinflusst wird. Dennoch bleibt laut Cherny das Restrisiko raffinierter Injektionen bestehen.

Original-Post

Shoopy

@0xShoopy · 25. September 2026

Boris Cherny (creator of Claude Code, Anthropic): "in a test, we slipped destructive commands into people's Claude Code sessions. they approved them almost every time." the "allow this command?" prompt was supposed to be the safety net. Boris calls it security theater – nobody t.co/vKzDA3dPPl t.co/zAoPhdgBcV

96 Likes25 Antworten131 Lesezeichen60.439 Aufrufe

Auf X ansehen
Weitere Posts im Thread (9)
  1. @shukshln independent verification without shared context is the key design choice here
  2. @tonyprediction and most people don't even realize they're the ones leaving it open
  3. @Ahmad_a_m_3 always has been, the model was never the problem
  4. @dreyk0o0 inevitable once the volume of approvals went up
  5. @sahilcode great point, zero context on the checker is what makes it resistant to social engineering
  6. @James__Hallam he was careful with it: they can no longer demonstrate a successful injection, not that none exist. he said they probably do, just really hard to find now. he posted a chart comparing labs, no paper that i've seen
  7. @CaseSignalX good comparison. a safety system people stop paying attention to is worse than none, because everyone assumes it's working
  8. @sophq37 that's basically where they landed. a second model with no context of your chat judges each command, so it's not relying on you reading anything
  9. @cyberogz exactly the right question. a check you can click past is just another yes button
Ausgewählte Antworten (5)
  • @MonishAtx @0xShoopy Makes sense. After the 50th allow prompt nobody reads the 51st. Approval prompts check intent, not outcome. The review that matters is after the fact: what changed, what broke, what got exposed. That's where I'd rather spend human attention.
  • @Ahmad_a_m_3 @0xShoopy The human approval step really is the weak link
  • @tonyprediction @0xShoopy wow thats a pretty wild security hole ngl
  • @sahilcode @0xShoopy the part that stands out is zero context on the second model. if it saw the whole conversation it would inherit whatever talked the human into clicking yes. checking the command in isolation is what makes it hard to social-engineer the same way twice.
  • @dreyk0o0 @0xShoopy click fatigue was always going to break approvals

Zusammenfassung von KI erstellt (Gemini 3.8 Flash, 27. September 2026). Sie kann Fehler enthalten – maßgeblich ist die Originalquelle.

Inhaltlich ähnlich, ermittelt über die KI-Suche.

  • Artikel:Anthropic Engineering

    Wie Anthropic Claude-Agenten isoliert und absichert

    Anthropic erläutert die Sicherheitsarchitektur zur Begrenzung des Schadenspotenzials (Blast Radius) von KI-Agenten in claude.ai und Claude Code. Da menschliche Kontrollabfragen an Genehmigungsmüdigkeit scheitern, setzt das Unternehmen vor allem auf OS- und Container-Sandboxes.

    KI & AI· Forschung

  • X-Post:Codez

    The Shift from Prompt Engineering to Graph Engineering for Agentic Systems

    The author argues that building effective AI agents is moving away from simple prompting toward 'graph engineering,' which utilizes structured workflows, loops, and parallel processing. By orchestrating subagents through code-based graphs rather than linear conversations, developers can create more robust and scalable agentic systems.

    5446Lesezeichen549.183Aufrufe

    KI & AI

  • Artikel:Anthropic

    Power-User-Tipps für Claude Code aus dem Anthropic-Team

    Das Entwicklerteam von Claude Code bei Anthropic teilt bewährte Workflows zur Steigerung der Entwicklungsgeschwindigkeit. Im Mittelpunkt stehen parallele Ausführung über Git-Worktrees, strukturierte Planung, iterative Fehlervermeidung via CLAUDE.md und automatisierte Verifikation.

    KI & AI· Sammlung

  • Repository:affaan-m/agentshield

    AgentShield: Sicherheits-Auditor für Claude-Code- und MCP-Konfigurationen

    AgentShield ist ein Open-Source-Sicherheitsscanner in TypeScript, der lokale Konfigurationen von KI-Agenten überprüft. Das Tool untersucht `.claude/`-Verzeichnisse auf hardcodierte Secrets, unsichere MCP-Server, Hook-Injections und Fehlkonfigurationen bei Berechtigungen.

    1163SterneTypeScript

    KI & AI· Tool

  • Link:Anthropic

    Anthropic stellt Claude Fable 5.1 und Claude Mythos 5.1 vor

    Anthropic kündigt die Frontier-Modelle Claude Fable 5.1 und Claude Mythos 5.1 an. Beide basieren auf demselben Modell, unterscheiden sich jedoch bei Sicherheitsbarrieren. Fable 5.1 bietet höhere Leistung beim Coding, wissenschaftlicher Forschung und Knowledge Work bei gleichzeitig reduzierten Kosten.

    KI & AI· Ankündigung

Lassen Sie uns über Ihr Projekt sprechen

Standorte

  • Mattersburg
    Johann Nepomuk Bergerstraße 7/2/14
    7210 Mattersburg, Austria
  • Wien
    Ungargasse 64-66/3/404
    1030 Wien, Austria

Dieser Inhalt wurde teilweise mithilfe von KI erstellt.