
Task framing, rather than the agent framework, is the primary driver of vulnerability in indirect prompt-injection attacks on LLM agents. This study presents a controlled empirical analysis across 3,274trials, benchmarking LangChain, MCP, and a Direct Anthropic API baseline exclusively on ClaudeSonnet 4 to isolate framework-specific vulnerabilities. The results demonstrate that vulnerabilityis framework-independent on the model tested: MCP adds no measurable attack surface relative toa raw API baseline (p > 0.6). Task framing was the dominant factor, with instruction-followingtasks reaching 89-91% attack success rates versus 15-26% for constrained tasks. LangChain’sFAISS retrieval introduced task-dependent exposure, amplifying attack surface when payloads weresemantically similar to the query and inadvertently suppressing them when they were not. Hardenedsystem prompts reduced attack success by over 17× but created a brittle threshold rather than agradient: among the 20-21 valid trials where the model engaged despite hardening, exploitation ratesexceeded 55%.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
