Agentic Tool-Execution Security and Sandbox Literature Review
Synthesize security literature on indirect prompt injection, privilege escalation, and sandboxing in tool-calling agents.
Use this template to conduct an exhaustive literature review on security vulnerabilities and sandboxing paradigms in autonomous agents. It audits academic defenses against injection attacks, privilege leakage, and unsafe tool execution.
Role: Senior AI Security Researcher specializing in agentic safety and runtime isolation.
Context
- Host Execution Environment: {{execution_environment}}
- Threat Model Boundaries: {{threat_model_scope}}
- Tool Privilege Hierarchy: {{tool_permission_hierarchy}}
- Surveyed Publication Base: {{audited_publication_set}}
- Taint Tracking Implementations: {{taint_tracking_approach}}
- Assessed Adversarial Vectors: {{adversarial_attack_vectors}}
Task
Execute an advanced systematic literature review on indirect prompt injection, tool poisoning, and privilege escalation vulnerabilities in tool-calling autonomous agents, producing an authoritative threat modeling analysis and defense taxonomy.
Method
- Categorize vulnerability taxonomies across {{audited_publication_set}} targeting tool-calling workflows and runtime execution.
- Dissect attack vectors (e.g., indirect prompt injection via tool outputs, confused deputy attacks, parameter tampering) against {{adversarial_attack_vectors}}.
- Evaluate academic proposals for data-vs-instruction isolation and dynamic taint analysis specified in {{taint_tracking_approach}}.
- Audit literature on capability attenuation, least-privilege enforcement, and dynamic confirmation boundaries under {{tool_permission_hierarchy}}.
- Critique formal sandbox and virtualized execution designs for agentic tool use within {{execution_environment}}.
- Measure the defensive trade-offs presented across papers between agent utility, execution latency, and security surface reduction within {{threat_model_scope}}.
- Benchmark empirical defense evaluation methodologies across surveyed papers to expose recurring evaluation leakage or unrealistic attacker assumptions.
- Synthesize a defensive defense-in-depth architectural blueprint derived from peer-reviewed mitigations.
Constraints
- MUST differentiate clearly between input-level filtering, latent representation guards, and runtime sandbox containment.
- MUST NOT accept vendor-reported empirical robustness metrics without evaluating adversarial evaluation rigor.
- Every surveyed defense MUST be explicitly evaluated against {{threat_model_scope}}.
- Recommendations MUST enforce least-privilege principles across {{tool_permission_hierarchy}}.
Output format
- Threat Landscape Executive Summary (max 250 words)
- Systematic Vulnerability & Exploitation Matrix (taxonomy table)
- Deep Analytical Review of Defensive Architectures (4 categorized technical sub-sections)
- Empirical Defense Efficacy & Latency Trade-Off Analysis (comparative matrix)
- Verified Defense-in-Depth Framework for Agent Tool Invocation (step-by-step hardened specification)
Self-review
- Are all adversarial attack surfaces strictly analyzed through the lens of tool-augmented autonomous execution?
- Does the critique rigorously examine the methodological validity of the reviewed security papers?
- Are all 6 context variables referenced purposefully throughout the analysis?
Explicit role, a named task, and discrete steps the model can follow.
Background, inputs and variables the model needs before it starts.
Hard boundaries — what the model must and must not do.
A named, field-level shape for the response.
Ordered work items that force analysis before an answer.
Length and structure that travel across frontier models.
Signal density — instruction weight without padding.
Documented variables so the scaffold adapts to new inputs.
Quality bar, assumptions and behaviour when inputs are thin.
How much real usage the template has behind it.