Skip to main content
IronClaw implements multiple defense layers to protect against prompt injection attacks when processing external content like emails, webhooks, web pages, and third-party API responses.

Threat Model

What is Prompt Injection?

Prompt injection is when untrusted external content attempts to manipulate the AI agent’s behavior by embedding malicious instructions:
Without defenses, the LLM might interpret “SYSTEM:” as a legitimate instruction.

Attack Vectors

Defense Architecture

Layer 1: Input Validation

First line of defense checks basic constraints:

Length Limits

Encoding Validation

Rejects malformed input:
  • Null bytes: \0 characters blocked
  • Invalid UTF-8: Rejected before processing
  • Excessive whitespace: Warned (>90% whitespace)
  • Character repetition: Warned (>20 repeated chars)
From src/safety/validator.rs:119-189:

Layer 2: Pattern Detection

Fast multi-pattern matching using Aho-Corasick algorithm to detect injection attempts:

Detected Patterns

Implementation from src/safety/sanitizer.rs:60-157:

Regex Patterns

Complex patterns detected via regex:

Case-Insensitive Matching

All patterns are case-insensitive to catch variants:

Layer 3: Content Sanitization

When critical patterns are detected, content is sanitized:

Escape Special Tokens

Escape Role Markers

Lines starting with role markers are prefixed:
Before:
After:

Layer 4: Policy Enforcement

High-level safety rules with configurable actions:

Policy Rules

From src/safety/policy.rs:130-201:

Policy Actions

Example: Block System File Access

This blocks content like:

Layer 5: Structural Wrapping

External content is wrapped with security delimiters before sending to the LLM:

wrap_external_content()

From src/safety/mod.rs:179-198:

Usage Example

LLM sees:

Tool Output Wrapping

Tool outputs are wrapped with XML-style tags:
Output:

Layer 6: Inbound Secret Detection

Before processing user input, scan for accidentally pasted secrets:
If user types:
System responds:

Complete Flow Example

Scenario: Malicious Email

Processing Pipeline

Step 1: Validation
Step 2: Pattern Detection
Step 3: Sanitization
Step 4: Policy Check
Step 5: Wrapping
Final LLM Input:
The LLM now sees:
  1. Clear security warning
  2. Escaped “SYSTEM:” role marker
  3. Structural delimiters separating instructions from data

Configuration

Safety settings in ~/.ironclaw/.env:
Disable injection checks (not recommended):

Limitations

What This Defends Against

  • ✅ Simple instruction injection (“ignore previous”)
  • ✅ Role marker injection (“system:”, “assistant:”)
  • ✅ Special token injection (<|endoftext|>)
  • ✅ Encoded payload injection (base64)
  • ✅ System file access attempts
  • ✅ Shell command injection patterns

What This Does NOT Defend Against

  • ❌ Sophisticated jailbreaks: Advanced adversarial prompts
  • ❌ Semantic attacks: Socially-engineered manipulation
  • ❌ LLM bugs: Zero-day vulnerabilities in the model itself
  • ❌ Context confusion: Subtle misdirection within valid-looking content
Prompt injection defense is best effort. No system can guarantee 100% protection against all adversarial inputs.

Best Practices

For Developers

  1. Always wrap external content: Use wrap_external_content() for emails, webhooks, web scraping
  2. Use tool output wrappers: Call wrap_for_llm() for all tool results
  3. Check sanitization flags: Inspect SanitizedOutput.was_modified
  4. Log warnings: Monitor warnings vec for attack attempts
  5. Don’t disable safety: Keep injection_check_enabled=true

For Users

  1. Review external integrations: Be cautious with email/webhook integrations
  2. Monitor logs: Watch for repeated sanitization warnings
  3. Report suspicious behavior: If the agent acts unexpectedly after processing external content
  4. Use allowlists: Restrict which senders/domains can trigger workflows

Source Code References

See Also