> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/nearai/ironclaw/llms.txt
> Use this file to discover all available pages before exploring further.

# System Architecture

> Understand IronClaw core components and how they work together

## Overview

IronClaw is built on a modular architecture that separates concerns while maintaining tight security boundaries. The system orchestrates agent reasoning, tool execution, multi-channel communication, and persistent memory through a set of core components.

## Architecture Diagram

```mermaid theme={null}
graph TB
    subgraph Channels
        REPL[REPL Channel]
        HTTP[HTTP Channel]
        WASM[WASM Channels<br/>Telegram, Slack]
        WEB[Web Gateway<br/>SSE + WebSocket]
    end
    
    subgraph Agent Core
        AL[Agent Loop]
        Router[Intent Router]
        Scheduler[Job Scheduler]
        Worker[Worker Pool]
        Routine[Routines Engine<br/>Cron, Event, Webhook]
    end
    
    subgraph Execution
        Local[Local Workers<br/>In-Process]
        Orch[Orchestrator]
        Docker[Docker Sandbox<br/>Containers]
    end
    
    subgraph Tools
        Registry[Tool Registry]
        Builtin[Built-in Tools]
        MCP[MCP Servers]
        WasmTools[WASM Tools]
    end
    
    subgraph Storage
        Workspace[Workspace<br/>Hybrid Search]
        DB[PostgreSQL<br/>Jobs, History]
    end
    
    REPL --> AL
    HTTP --> AL
    WASM --> AL
    WEB --> AL
    
    AL --> Router
    Router --> Scheduler
    Scheduler --> Worker
    AL --> Routine
    
    Worker --> Local
    Worker --> Orch
    Orch --> Docker
    
    Local --> Registry
    Docker --> Registry
    
    Registry --> Builtin
    Registry --> MCP
    Registry --> WasmTools
    
    Worker --> Workspace
    Worker --> DB
```

## Core Components

### Agent Loop

The central orchestrator that coordinates all system activity.

<CodeGroup>
  ```rust src/agent/agent_loop.rs theme={null}
  pub struct Agent {
      config: AgentConfig,
      deps: AgentDeps,
      channels: Arc<ChannelManager>,
      context_manager: Arc<ContextManager>,
      scheduler: Arc<Scheduler>,
      router: Router,
      session_manager: Arc<SessionManager>,
  }
  ```
</CodeGroup>

**Responsibilities:**

* Route incoming messages from channels
* Classify user intent (command vs query vs task)
* Delegate to appropriate handlers
* Coordinate session and thread management
* Trigger background systems (heartbeat, routines, self-repair)

### Router

Classifies incoming messages to determine handling strategy.

```rust theme={null}
pub enum MessageIntent {
    Command,      // System commands (/help, /status)
    Query,        // Information requests
    Task,         // Work requiring tools/planning
    Conversation, // Chat, general interaction
}
```

**Intent Classification:**

* **Commands**: Direct system operations (`/quit`, `/undo`, `/status`)
* **Queries**: Information retrieval from memory or knowledge
* **Tasks**: Complex work requiring planning and tool execution
* **Conversation**: General chat and interaction

### Scheduler

Manages parallel job execution with priorities and resource limits.

<CodeGroup>
  ```rust src/agent/scheduler.rs theme={null}
  pub struct Scheduler {
      config: AgentConfig,
      context_manager: Arc<ContextManager>,
      jobs: Arc<RwLock<HashMap<Uuid, ScheduledJob>>>,
      subtasks: Arc<RwLock<HashMap<Uuid, ScheduledSubtask>>>,
  }
  ```
</CodeGroup>

**Key Features:**

* Parallel job execution (configurable limit)
* Per-job worker isolation
* Subtask spawning for parallel tool execution
* Automatic cleanup on completion
* Job state tracking (pending, in\_progress, completed, failed, stuck)

### Worker

Executes individual jobs with LLM reasoning and tool calls.

```rust theme={null}
pub struct Worker {
    job_id: Uuid,
    deps: WorkerDeps,
}
```

**Execution Flow:**

1. **Planning** (optional): Generate action plan with LLM
2. **Tool Selection**: Choose tools based on context
3. **Parallel Execution**: Run independent tools concurrently
4. **Result Processing**: Sanitize output, update context
5. **Iteration**: Loop until job complete or max iterations
6. **Completion**: Mark job as completed/failed/stuck

<Tip>
  Workers support both **planning mode** (generate upfront plan) and **direct selection** (iterative tool selection). Planning mode is more efficient for complex multi-step tasks.
</Tip>

### Routines Engine

Background automation for scheduled and reactive tasks.

```rust theme={null}
pub enum Trigger {
    Cron(String),              // Schedule: "0 9 * * MON"
    Event { pattern: String }, // Regex: "deploy.*failed"
    Webhook { path: String },  // HTTP: /webhook/deploy
}
```

**Routine Types:**

* **Cron**: Time-based schedules (daily reports, periodic checks)
* **Event**: Message pattern matching (alert on errors)
* **Webhook**: HTTP endpoint triggers (CI/CD integration)

**Use Cases:**

* Daily standup summaries
* Alert monitoring and triage
* Periodic health checks
* Automated reporting

### Orchestrator

Manages Docker sandbox containers for isolated code execution.

<CodeGroup>
  ```rust src/orchestrator/mod.rs theme={null}
  pub struct ContainerJobManager {
      // Container lifecycle management
      // Per-job authentication tokens
      // LLM proxy for worker containers
      // Credential injection boundary
  }
  ```
</CodeGroup>

**Security Model:**

* Per-job bearer tokens (ephemeral, in-memory only)
* Network-isolated containers
* Resource limits (CPU, memory, timeout)
* Credential injection at orchestrator boundary
* No direct database access from containers

**Worker/Orchestrator Pattern:**

```text theme={null}
┌─────────────────────────────────────────────────────────┐
│                     Orchestrator                        │
│  - HTTP API (:50051)                                    │
│  - Token validation                                     │
│  - LLM proxy                                            │
│  - Credential injection                                 │
└────────────────────┬────────────────────────────────────┘
                     │ HTTP + Bearer Token
                     ▼
┌─────────────────────────────────────────────────────────┐
│              Docker Container (Worker)                  │
│  - Isolated filesystem                                  │
│  - Limited tools (shell, file ops)                      │
│  - No direct secrets access                             │
└─────────────────────────────────────────────────────────┘
```

<Warning>
  Worker containers have **no direct access** to secrets. All credentials are injected by the orchestrator at request time after token validation.
</Warning>

## Data Flow

### Message Processing

```mermaid theme={null}
sequenceDiagram
    participant C as Channel
    participant AL as Agent Loop
    participant R as Router
    participant S as Scheduler
    participant W as Worker
    participant T as Tools
    
    C->>AL: IncomingMessage
    AL->>R: Classify intent
    R-->>AL: MessageIntent
    AL->>S: Create job
    S->>W: Spawn worker
    W->>T: Select & execute tools
    T-->>W: ToolOutput
    W->>W: LLM reasoning
    W-->>S: Job complete
    S-->>AL: Result
    AL->>C: OutgoingResponse
```

### Job Lifecycle

```mermaid theme={null}
stateDiagram-v2
    [*] --> Pending
    Pending --> InProgress: Scheduler.schedule()
    InProgress --> Completed: Success
    InProgress --> Failed: Error
    InProgress --> Stuck: Timeout/MaxIterations
    InProgress --> Cancelled: User cancel
    Stuck --> InProgress: Self-repair
    Stuck --> Failed: Repair failed
    Completed --> [*]
    Failed --> [*]
    Cancelled --> [*]
```

## Self-Repair System

Automatic detection and recovery of stuck operations.

**Detection:**

* Jobs stuck in `InProgress` beyond threshold
* Tools with high failure rates
* Unresponsive worker processes

**Recovery Strategies:**

<AccordionGroup>
  <Accordion title="Stuck Job Recovery">
    1. Detect job stuck > threshold (default 5min)
    2. Analyze context and last action
    3. Attempt recovery:
       * Retry failed tool
       * Restart worker with fresh context
       * Escalate to manual intervention
    4. Notify user of outcome
  </Accordion>

  <Accordion title="Broken Tool Recovery">
    1. Track tool failure rates
    2. Identify consistently failing tools
    3. Recovery options:
       * Clear tool cache
       * Rebuild WASM tool
       * Disable tool temporarily
       * Suggest alternative tools
  </Accordion>
</AccordionGroup>

## Context Management

Each job maintains isolated context for safe parallel execution.

```rust theme={null}
pub struct JobContext {
    pub id: Uuid,
    pub title: String,
    pub description: String,
    pub user_id: String,
    pub state: JobState,
    pub created_at: DateTime<Utc>,
    pub metadata: serde_json::Value,
}
```

**Context Isolation:**

* Each job has independent memory
* No shared mutable state between jobs
* Tool execution scoped to job context
* LLM history isolated per job

**Context Compaction:**

When conversation history grows large:

1. **Detect**: Monitor token count per thread
2. **Summarize**: Use LLM to summarize old turns
3. **Preserve**: Keep recent turns intact
4. **Replace**: Swap old turns with summary
5. **Continue**: Resume conversation with more tokens

<Info>
  Compaction triggers automatically at 75% of max context window. Recent turns (last 10) are always preserved.
</Info>

## Session Management

Multi-threaded conversations with undo/redo support.

**Features:**

* Multiple concurrent threads per user
* Turn-based checkpointing
* Undo/redo with state restoration
* Session persistence to database
* Automatic pruning of stale sessions

**Turn Structure:**

```rust theme={null}
pub struct Turn {
    pub user_input: String,
    pub assistant_response: String,
    pub tool_calls: Vec<ToolCall>,
    pub state: TurnState,
    pub created_at: DateTime<Utc>,
}
```

## Performance Characteristics

### Parallel Tool Execution

Tools with no dependencies execute concurrently:

```rust theme={null}
// Sequential (slow)
let result1 = execute_tool("api_call_1").await;
let result2 = execute_tool("api_call_2").await;
let result3 = execute_tool("api_call_3").await;
// Total: ~600ms

// Parallel (fast)
let results = execute_tools_parallel([
    "api_call_1",
    "api_call_2", 
    "api_call_3"
]).await;
// Total: ~200ms
```

<Tip>
  The worker automatically detects independent tool calls and executes them in parallel using a `JoinSet`.
</Tip>

### Resource Limits

| Resource | Default Limit | Configurable |
| - | - | - |
| Max parallel jobs | 10 | Yes (`max_parallel_jobs`) |
| Job timeout | 30 minutes | Yes (`job_timeout`) |
| Max iterations | 50 | Yes (per-job metadata) |
| Stuck threshold | 5 minutes | Yes (`stuck_threshold`) |
| Session idle timeout | 1 hour | Yes (`session_idle_timeout`) |

## Next Steps

<CardGroup cols={2}>
  <Card title="Security Model" icon="shield" href="/concepts/security">
    Learn about defense-in-depth security layers
  </Card>

  <Card title="Channel System" icon="satellite-dish" href="/concepts/channels">
    Multi-channel communication architecture
  </Card>

  <Card title="Tool System" icon="wrench" href="/concepts/tools">
    Extensible tool system and WASM sandbox
  </Card>

  <Card title="Workspace & Memory" icon="brain" href="/concepts/workspace">
    Persistent memory and hybrid search
  </Card>
</CardGroup>
