Skip to content

Work / Agent security

A defensive prompt stopped some injection payloads and missed others. I recommend fixing tool dispatch instead.

Pre-registered research on indirect prompt injection through agent tool output, with the misses published alongside the hits.

Problem

Agents read tool output: web pages, files, API responses. If an attacker controls that output, can they make the agent leak data? The common answer is to add a defensive prompt telling the model to distrust tool content. I wanted to know whether that holds up, and whether a bigger model is safer.

Approach

Architecture

Attack path from untrusted tool output through inline XML dispatch to a canary read and exfiltration attempt, and the recommended defense: substrate_auditor flags the risk, typed tool-call dispatch (recommended, not measured at 70B), and CI checks.
Redrawn from the diagram in the ai-redteam-notes README.

Results

Strict canary exfiltration on Groq-hosted Llama-3.3-70B, n=10 seeds per cell (source and per-seed runs):

ConditionExfiltratedWhat it shows
No scaffold (M0): baseline, v3, v7 payloads9 / 10, 10 / 10, 10 / 10The inline-XML dispatch path is exploitable
Content-trust scaffold (M1): baseline and v3 payloads0 / 10 eachThe scaffold stopped these two variants
Content-trust scaffold (M1): v7 payload10 / 10The same scaffold failed completely against this one
Chat-only control (no XML dispatch)0 / 10Without the inline-XML dispatch path there is no tool call to hijack

Takeaway: prompt scaffolding is variant-selective. It blocked some payloads and not others, so it can't be the fix on its own. My recommendation is architectural: typed tool-call dispatch in the agent client, authorization and provenance checks outside the model, and prompt hardening as defense in depth. That rests on the substrate isolation above; a typed-dispatch agent has not yet been measured at 70B. Source: the H10b-G report.

Published misses. The pre-registered assumption that a larger model would be safer here (H7) was falsified at 70B. H3 was retracted for a substrate confound, H6 and H11 were falsified, and one mitigation (M2) measured as a regression. They're listed in the hypothesis ledger because a record that hides its nulls is harder to trust.

What I'd do next

Methodology & limits

Scope of the evidence
  • "Typed dispatch is the fix" is architectural reasoning backed by thin measurement: the typed-substrate cell is a single simulated 8B run. At 70B, the isolation check is the chat-only control, not a typed substrate.
  • The H10b-G grid used n=10 per cell on a single hosted provider. An earlier 8B baseline was n=5 on a free tier and is directional only.
  • Targets were intentionally vulnerable benchmarks, open-weight models and local harnesses, under a three-tier disclosure policy. The pre-push hook described in SECURITY.md is not in the public repo; there, disclosure frontmatter is checked by make public-surface in CI.
  • The chat-only control has no tool-dispatch path, so its 0/10 is expected by construction. It isolates the substrate; it does not test a typed-dispatch agent at 70B.
  • The auditor recognizes the transcript and config shapes it was built for and returns unknown on others, such as a plain mcpServers config.
  • Full claim-to-run mapping and limitations: the repository README and research summary.

Contact

Open to full-time AI Engineer roles

US remote. Applied AI, AI backend and LLM platform teams.