Work / Agent security
A defensive prompt stopped some injection payloads and missed others. I recommend fixing tool dispatch instead.
Pre-registered research on indirect prompt injection through agent tool output, with the misses published alongside the hits.
make reproProblem
Agents read tool output: web pages, files, API responses. If an attacker controls that output, can they make the agent leak data? The common answer is to add a defensive prompt telling the model to distrust tool content. I wanted to know whether that holds up, and whether a bigger model is safer.
Approach
- Pre-registered hypotheses. Hypotheses and success criteria were written down and dated before the runs that test them (ledger). The headline 70B grid ran n=10 per cell with Wilson 95% intervals, on Groq: a documented deviation from its frozen OpenRouter pre-registration. Earlier cells were n=5 or n=1.
- Controlled substrate. The same payloads ran through an inline text/XML tool-dispatch substrate and a chat-only control, so the dispatch mechanism could be isolated from the model.
- Safe targets only. Intentionally vulnerable benchmarks, open-weight models via public APIs and localhost harnesses. Nothing in production was tested without authorization.
- A tool, not just a writeup. The finding became
substrate_auditor.py, a zero-dependency check that classifies a config or transcript as inline-XML or typed dispatch. CI runs its self-check and requires it to rate the inline-XML sample high-risk and the typed sample low-risk.
Architecture
Results
Strict canary exfiltration on Groq-hosted Llama-3.3-70B, n=10 seeds per cell (source and per-seed runs):
| Condition | Exfiltrated | What it shows |
|---|---|---|
| No scaffold (M0): baseline, v3, v7 payloads | 9 / 10, 10 / 10, 10 / 10 | The inline-XML dispatch path is exploitable |
| Content-trust scaffold (M1): baseline and v3 payloads | 0 / 10 each | The scaffold stopped these two variants |
| Content-trust scaffold (M1): v7 payload | 10 / 10 | The same scaffold failed completely against this one |
| Chat-only control (no XML dispatch) | 0 / 10 | Without the inline-XML dispatch path there is no tool call to hijack |
Takeaway: prompt scaffolding is variant-selective. It blocked some payloads and not others, so it can't be the fix on its own. My recommendation is architectural: typed tool-call dispatch in the agent client, authorization and provenance checks outside the model, and prompt hardening as defense in depth. That rests on the substrate isolation above; a typed-dispatch agent has not yet been measured at 70B. Source: the H10b-G report.
Published misses. The pre-registered assumption that a larger model would be safer here (H7) was falsified at 70B. H3 was retracted for a substrate confound, H6 and H11 were falsified, and one mitigation (M2) measured as a regression. They're listed in the hypothesis ledger because a record that hides its nulls is harder to trust.
What I'd do next
- Replicate the 70B grid on a second inference provider; the current grid is single-provider.
- Extend the payload set and the auditor's checks to more agent frameworks' dispatch formats.
- Package
substrate_auditoras a pre-deploy check other teams can drop into CI.
Methodology & limits
Scope of the evidence
- "Typed dispatch is the fix" is architectural reasoning backed by thin measurement: the typed-substrate cell is a single simulated 8B run. At 70B, the isolation check is the chat-only control, not a typed substrate.
- The H10b-G grid used n=10 per cell on a single hosted provider. An earlier 8B baseline was n=5 on a free tier and is directional only.
- Targets were intentionally vulnerable benchmarks, open-weight models and local harnesses, under a three-tier disclosure policy. The pre-push hook described in SECURITY.md is not in the public repo; there, disclosure frontmatter is checked by
make public-surfacein CI. - The chat-only control has no tool-dispatch path, so its 0/10 is expected by construction. It isolates the substrate; it does not test a typed-dispatch agent at 70B.
- The auditor recognizes the transcript and config shapes it was built for and returns
unknownon others, such as a plainmcpServersconfig. - Full claim-to-run mapping and limitations: the repository README and research summary.