Threat Model
What is Prompt Injection?
Prompt injection is when untrusted external content attempts to manipulate the AI agent’s behavior by embedding malicious instructions:Attack Vectors
Defense Architecture
Layer 1: Input Validation
First line of defense checks basic constraints:Length Limits
Encoding Validation
Rejects malformed input:- Null bytes:
\0characters blocked - Invalid UTF-8: Rejected before processing
- Excessive whitespace: Warned (>90% whitespace)
- Character repetition: Warned (>20 repeated chars)
Layer 2: Pattern Detection
Fast multi-pattern matching using Aho-Corasick algorithm to detect injection attempts:Detected Patterns
Implementation from src/safety/sanitizer.rs:60-157:
Regex Patterns
Complex patterns detected via regex:Case-Insensitive Matching
All patterns are case-insensitive to catch variants:Layer 3: Content Sanitization
When critical patterns are detected, content is sanitized:Escape Special Tokens
Escape Role Markers
Lines starting with role markers are prefixed:Layer 4: Policy Enforcement
High-level safety rules with configurable actions:Policy Rules
From src/safety/policy.rs:130-201:Policy Actions
Example: Block System File Access
Layer 5: Structural Wrapping
External content is wrapped with security delimiters before sending to the LLM:wrap_external_content()
From src/safety/mod.rs:179-198:Usage Example
Tool Output Wrapping
Tool outputs are wrapped with XML-style tags:Layer 6: Inbound Secret Detection
Before processing user input, scan for accidentally pasted secrets:Complete Flow Example
Scenario: Malicious Email
Processing Pipeline
Step 1: Validation- Clear security warning
- Escaped “SYSTEM:” role marker
- Structural delimiters separating instructions from data
Configuration
Safety settings in~/.ironclaw/.env:
Limitations
What This Defends Against
- ✅ Simple instruction injection (“ignore previous”)
- ✅ Role marker injection (“system:”, “assistant:”)
- ✅ Special token injection (
<|endoftext|>) - ✅ Encoded payload injection (base64)
- ✅ System file access attempts
- ✅ Shell command injection patterns
What This Does NOT Defend Against
- ❌ Sophisticated jailbreaks: Advanced adversarial prompts
- ❌ Semantic attacks: Socially-engineered manipulation
- ❌ LLM bugs: Zero-day vulnerabilities in the model itself
- ❌ Context confusion: Subtle misdirection within valid-looking content
Best Practices
For Developers
- Always wrap external content: Use
wrap_external_content()for emails, webhooks, web scraping - Use tool output wrappers: Call
wrap_for_llm()for all tool results - Check sanitization flags: Inspect
SanitizedOutput.was_modified - Log warnings: Monitor
warningsvec for attack attempts - Don’t disable safety: Keep
injection_check_enabled=true
For Users
- Review external integrations: Be cautious with email/webhook integrations
- Monitor logs: Watch for repeated sanitization warnings
- Report suspicious behavior: If the agent acts unexpectedly after processing external content
- Use allowlists: Restrict which senders/domains can trigger workflows
Source Code References
- SafetyLayer: src/safety/mod.rs:28-177
- Sanitizer: src/safety/sanitizer.rs:35-285
- Validator: src/safety/validator.rs:79-224
- Policy: src/safety/policy.rs:93-201
- External content wrapper: src/safety/mod.rs:179-198
See Also
- Security Overview - Complete security architecture
- WASM Sandbox - Tool isolation and capabilities
- Leak Detection - Secret scanning in outputs