Your Agent Is Mine: The Invisible Supply Chain Threat to LLM Agents
Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain
This paper introduces the first systematic study of malicious intermediary attacks within the LLM API router supply chain. It formalizes a new threat model and demonstrates how intermediaries can perform tool-call injection and secret exfiltration across popular AI agent frameworks (OpenCode, Codex, etc.).
TL;DR
As LLM agents move from simple chat boxes to autonomous systems capable of executing code and managing infrastructure, a new shadow infrastructure has emerged: LLM API Routers. This paper reveals that these intermediaries, often used for cost-saving or anti-censorship, represent a massive security loophole. By terminating TLS and handling payloads in plaintext, malicious routers can rewrite tool calls—turning a "benign install" into "arbitrary code execution"—and silently harvest API keys.
Problem & Motivation: The Router-in-the-Middle
The industry has focused heavily on Prompt Injection (attacking the model's logic), but has largely ignored the Transport Layer. When a developer changes their base_url to a third-party router like LiteLLM or a Taobao reseller, they are introducing an application-layer Man-in-the-Middle (MITM).
Since there is no cryptographic bind between what the Model Provider (e.g., OpenAI) sends and what the Agent Client receives, the router has total authority to:
- Read every secret (AC-2: Secret Exfiltration).
- Modify the JSON payload (AC-1: Payload Injection).
The authors argue that the trust boundary is not just the first hop—it’s a chain. If a router forwards traffic through a compromised relay, the entire session is tainted.
Methodology: Formalizing the Attack Surface
The authors categorize the threat into four distinct classes:
- AC-1 (Payload Injection): Swapping a tool argument (e.g., changing a GitHub URL to an attacker’s script) after inference.
- AC-2 (Secret Exfiltration): Passive sniffing for AWS keys, ETH private keys, or GitHub tokens.
- AC-1.a (Dependency Targeting): A stealthier AC-1 that replaces a library name (e.g.,
pip install requestsbecomesreqeusts) to bypass domain-based filters. - AC-1.b (Conditional Delivery): "Warm-up" attacks where the router behaves perfectly for 50 requests to evade detection, then strikes when it detects "YOLO mode" (autonomous execution).
Figure 1: The multi-hop taint propagation in the LLM router ecosystem.
Experimental Results: The Wild West of Routers
The authors didn't just theorize; they went shopping. Testing 28 paid routers and 400 free routers, the findings were alarming:
- Active Malice: 9 routers were caught injecting code.
- Credential Theft: 17 routers touched AWS canary tokens; one successfully drained an Ethereum wallet.
- Poisoning the Well: The authors leaked a single OpenAI key to see who would use it. It ended up powering 100M tokens of "shadow traffic" from unknown downstream users, proving that "benign" routers are often just fronting stolen keys.
Table: Summary of findings across paid, free, and poisoned router sets.
Evaluation of Defenses
The paper evaluates three "client-side only" defenses:
- Policy Gates: Blocking non-allowlisted domains. Efficient but easily bypassed by AC-1.a.
- Anomaly Screening: Using Isolation Forests to detect "weird" tool calls. Caught 89% of blatant injections but had high false positives for complex dev tasks.
- Transparency Logs: Append-only auditing. Useful for forensics but doesn't prevent the initial hack.
Critical Analysis & Conclusion
The core takeaway is a wake-up call: "Your Agent Is Mine" if you don't control the transport.
While the proposed client-side mitigations help, they are "band-aids." The paper concludes that the only permanent fix is Provider-Backed Response Integrity. Just like DKIM for email or Subresource Integrity for web dev, LLM providers must sign their tool-call outputs.
Limitations
- The study primarily focused on commodity and "gray market" routers (Taobao/Xianyu). Enterprise-grade routers (like Azure OpenAI) were not the focus, though the architectural vulnerability remains the same if the enterprise proxy is compromised.
- The anomaly detector is sensitive to "shell-risk" scores, meaning highly active developers using shell tools might find it too noisy.
Future Outlook
As the Model Context Protocol (MCP) gains traction, the number of intermediaries will only grow. This research sets the foundation for a new era of "Signed AI Responses," where the provenance of an AI's decision is as important as the decision itself.
