> ## Documentation Index
> Fetch the complete documentation index at: https://mintlify.com/nearai/ironclaw/llms.txt
> Use this file to discover all available pages before exploring further.

# LLM Module

> Language model integration with multiple providers and resilience features

## Overview

The LLM module provides integration with multiple language model providers:

* **NEAR AI** (default): Session token or API key auth via Chat Completions API
* **OpenAI**: Direct API access with your own key
* **Anthropic**: Claude models via direct API
* **Ollama**: Local model inference
* **OpenAI-compatible**: Any endpoint that speaks the OpenAI API
* **Tinfoil**: Private inference endpoints

Includes resilience features: retry logic, failover, circuit breakers, response caching, and smart routing.

## Core Types

### LlmProvider Trait

The main trait that all LLM providers must implement.

```rust theme={null}
#[async_trait]
pub trait LlmProvider: Send + Sync {
    fn model_name(&self) -> &str;
    fn cost_per_token(&self) -> (Decimal, Decimal);
    
    async fn complete(&self, request: CompletionRequest) -> Result<CompletionResponse, LlmError>;
    
    async fn complete_with_tools(
        &self,
        request: ToolCompletionRequest,
    ) -> Result<ToolCompletionResponse, LlmError>;
    
    async fn list_models(&self) -> Result<Vec<String>, LlmError> {
        Ok(Vec::new())
    }
    
    async fn model_metadata(&self) -> Result<ModelMetadata, LlmError> {
        Ok(ModelMetadata {
            id: self.model_name().to_string(),
            context_length: None,
        })
    }
}
```

#### Methods

<ResponseField name="model_name" type="fn(&self) -> &str">
  Return the model identifier
</ResponseField>

<ResponseField name="cost_per_token" type="fn(&self) -> (Decimal, Decimal)">
  Return (input\_cost, output\_cost) per token in USD
</ResponseField>

<ResponseField name="complete" type="async fn(&self, request: CompletionRequest) -> Result<CompletionResponse>">
  Generate a completion for a conversation
</ResponseField>

<ResponseField name="complete_with_tools" type="async fn(&self, request: ToolCompletionRequest) -> Result<ToolCompletionResponse>">
  Generate a completion with tool calling support
</ResponseField>

<ResponseField name="list_models" type="async fn(&self) -> Result<Vec<String>>">
  List available models from the provider (optional)
</ResponseField>

<ResponseField name="model_metadata" type="async fn(&self) -> Result<ModelMetadata>">
  Fetch metadata about the model (context length, etc.)
</ResponseField>

### ChatMessage

A message in a conversation.

<ParamField path="role" type="Role">
  Message role (System, User, Assistant, Tool)
</ParamField>

<ParamField path="content" type="String">
  Message content/text
</ParamField>

<ParamField path="tool_call_id" type="Option<String>">
  Tool call ID if this is a tool result message
</ParamField>

<ParamField path="name" type="Option<String>">
  Tool name for tool result messages
</ParamField>

<ParamField path="tool_calls" type="Option<Vec<ToolCall>>">
  Tool calls made by the assistant
</ParamField>

#### Constructors

<ResponseField name="system" type="fn(content: impl Into<String>) -> Self">
  Create a system message
</ResponseField>

<ResponseField name="user" type="fn(content: impl Into<String>) -> Self">
  Create a user message
</ResponseField>

<ResponseField name="assistant" type="fn(content: impl Into<String>) -> Self">
  Create an assistant message
</ResponseField>

<ResponseField name="assistant_with_tool_calls" type="fn(content: Option<String>, tool_calls: Vec<ToolCall>) -> Self">
  Create an assistant message with tool calls
</ResponseField>

<ResponseField name="tool_result" type="fn(tool_call_id: impl Into<String>, name: impl Into<String>, content: impl Into<String>) -> Self">
  Create a tool result message
</ResponseField>

#### Example

```rust theme={null}
use ironclaw::llm::{ChatMessage, Role};

let messages = vec![
    ChatMessage::system("You are a helpful assistant."),
    ChatMessage::user("What is the capital of France?"),
    ChatMessage::assistant("The capital of France is Paris."),
];
```

### CompletionRequest

Request for a chat completion.

<ParamField path="messages" type="Vec<ChatMessage>">
  Conversation history
</ParamField>

<ParamField path="model" type="Option<String>">
  Optional per-request model override
</ParamField>

<ParamField path="max_tokens" type="Option<u32>">
  Maximum tokens to generate
</ParamField>

<ParamField path="temperature" type="Option<f32>">
  Sampling temperature (0.0 to 2.0)
</ParamField>

<ParamField path="stop_sequences" type="Option<Vec<String>>">
  Sequences that stop generation
</ParamField>

<ParamField path="metadata" type="HashMap<String, String>">
  Opaque metadata passed to provider
</ParamField>

#### Methods

<ResponseField name="new" type="fn(messages: Vec<ChatMessage>) -> Self">
  Create a new completion request
</ResponseField>

<ResponseField name="with_model" type="fn(self, model: impl Into<String>) -> Self">
  Set model override
</ResponseField>

<ResponseField name="with_max_tokens" type="fn(self, max_tokens: u32) -> Self">
  Set max tokens
</ResponseField>

<ResponseField name="with_temperature" type="fn(self, temperature: f32) -> Self">
  Set temperature
</ResponseField>

#### Example

```rust theme={null}
use ironclaw::llm::{CompletionRequest, ChatMessage};

let request = CompletionRequest::new(messages)
    .with_max_tokens(1000)
    .with_temperature(0.7);

let response = llm.complete(request).await?;
println!("Response: {}", response.content);
```

### CompletionResponse

Response from a chat completion.

<ParamField path="content" type="String">
  Generated text content
</ParamField>

<ParamField path="input_tokens" type="u32">
  Number of tokens in the prompt
</ParamField>

<ParamField path="output_tokens" type="u32">
  Number of tokens generated
</ParamField>

<ParamField path="finish_reason" type="FinishReason">
  Why generation stopped
</ParamField>

#### FinishReason

<ResponseField name="Stop" type="()">
  Model naturally completed the response
</ResponseField>

<ResponseField name="Length" type="()">
  Hit max\_tokens limit
</ResponseField>

<ResponseField name="ToolUse" type="()">
  Model wants to use tools
</ResponseField>

<ResponseField name="ContentFilter" type="()">
  Filtered by content policy
</ResponseField>

<ResponseField name="Unknown" type="()">
  Unknown reason
</ResponseField>

### ToolCompletionRequest

Request for a completion with tool use support.

<ParamField path="messages" type="Vec<ChatMessage>">
  Conversation history
</ParamField>

<ParamField path="tools" type="Vec<ToolDefinition>">
  Available tools for the model
</ParamField>

<ParamField path="model" type="Option<String>">
  Optional model override
</ParamField>

<ParamField path="max_tokens" type="Option<u32>">
  Maximum tokens to generate
</ParamField>

<ParamField path="temperature" type="Option<f32>">
  Sampling temperature
</ParamField>

<ParamField path="tool_choice" type="Option<String>">
  How to handle tools: "auto", "required", or "none"
</ParamField>

<ParamField path="metadata" type="HashMap<String, String>">
  Opaque metadata
</ParamField>

#### Methods

<ResponseField name="new" type="fn(messages: Vec<ChatMessage>, tools: Vec<ToolDefinition>) -> Self">
  Create a new tool completion request
</ResponseField>

<ResponseField name="with_tool_choice" type="fn(self, choice: impl Into<String>) -> Self">
  Set tool choice mode ("auto", "required", "none")
</ResponseField>

#### Example

```rust theme={null}
use ironclaw::llm::{ToolCompletionRequest, ToolDefinition};

let tools = vec![
    ToolDefinition {
        name: "get_weather".to_string(),
        description: "Get current weather".to_string(),
        parameters: serde_json::json!({
            "type": "object",
            "properties": {
                "location": { "type": "string" }
            },
            "required": ["location"]
        }),
    }
];

let request = ToolCompletionRequest::new(messages, tools)
    .with_tool_choice("auto");

let response = llm.complete_with_tools(request).await?;
for tool_call in response.tool_calls {
    println!("Tool: {} - Args: {}", tool_call.name, tool_call.arguments);
}
```

### ToolCompletionResponse

Response from a tool-enabled completion.

<ParamField path="content" type="Option<String>">
  Text content (may be empty if tool calls present)
</ParamField>

<ParamField path="tool_calls" type="Vec<ToolCall>">
  Tool calls requested by the model
</ParamField>

<ParamField path="input_tokens" type="u32">
  Prompt tokens
</ParamField>

<ParamField path="output_tokens" type="u32">
  Generated tokens
</ParamField>

<ParamField path="finish_reason" type="FinishReason">
  Why generation stopped
</ParamField>

### ToolCall

A tool call requested by the LLM.

<ParamField path="id" type="String">
  Unique call identifier
</ParamField>

<ParamField path="name" type="String">
  Tool name
</ParamField>

<ParamField path="arguments" type="serde_json::Value">
  Tool arguments as JSON
</ParamField>

### ToolDefinition

Definition of a tool for the LLM.

<ParamField path="name" type="String">
  Tool name
</ParamField>

<ParamField path="description" type="String">
  What the tool does
</ParamField>

<ParamField path="parameters" type="serde_json::Value">
  JSON Schema for parameters
</ParamField>

### ToolResult

Result of tool execution to send back to the LLM.

<ParamField path="tool_call_id" type="String">
  ID of the tool call this is a result for
</ParamField>

<ParamField path="name" type="String">
  Tool name
</ParamField>

<ParamField path="content" type="String">
  Result content
</ParamField>

<ParamField path="is_error" type="bool">
  Whether this result is an error
</ParamField>

## Provider Creation

### create\_llm\_provider

<ResponseField name="create_llm_provider" type="fn(config: &LlmConfig, session: Arc<SessionManager>) -> Result<Arc<dyn LlmProvider>>">
  Create an LLM provider based on configuration
</ResponseField>

```rust theme={null}
use ironclaw::llm::{create_llm_provider, SessionManager};
use ironclaw::config::LlmConfig;

let session_mgr = Arc::new(SessionManager::new(session_config));
let llm = create_llm_provider(&llm_config, session_mgr)?;
```

### build\_provider\_chain

<ResponseField name="build_provider_chain" type="fn(config: &LlmConfig, session: Arc<SessionManager>) -> Result<(Arc<dyn LlmProvider>, Option<Arc<dyn LlmProvider>>)>">
  Build the full provider chain with retry, failover, circuit breaker, cache, etc. Returns (main\_provider, cheap\_provider).
</ResponseField>

```rust theme={null}
use ironclaw::llm::build_provider_chain;

let (llm, cheap_llm) = build_provider_chain(&config, session_mgr)?;
```

The provider chain applies decorators in this order:

1. Raw provider (from config)
2. RetryProvider (exponential backoff)
3. SmartRoutingProvider (cheap/primary split)
4. FailoverProvider (fallback model)
5. CircuitBreakerProvider (fast-fail when degraded)
6. CachedProvider (response cache)

## Resilience Features

### RetryProvider

Retries failed requests with exponential backoff.

<ResponseField name="new" type="fn(provider: Arc<dyn LlmProvider>, config: RetryConfig) -> Self">
  Wrap a provider with retry logic
</ResponseField>

#### RetryConfig

<ParamField path="max_retries" type="u32" default="3">
  Maximum number of retry attempts
</ParamField>

```rust theme={null}
use ironclaw::llm::{RetryProvider, RetryConfig};

let retry_llm = Arc::new(RetryProvider::new(
    base_llm,
    RetryConfig { max_retries: 3 }
));
```

### FailoverProvider

Fails over to backup providers when primary fails.

<ResponseField name="new" type="fn(providers: Vec<Arc<dyn LlmProvider>>) -> Result<Self>">
  Create a failover provider with ordered list of providers
</ResponseField>

<ResponseField name="with_cooldown" type="fn(providers: Vec<Arc<dyn LlmProvider>>, config: CooldownConfig) -> Result<Self>">
  Create failover with cooldown (temporary disable after failures)
</ResponseField>

#### CooldownConfig

<ParamField path="cooldown_duration" type="Duration" default="300s">
  How long to disable a provider after failures
</ParamField>

<ParamField path="failure_threshold" type="usize" default="3">
  Number of failures before cooldown
</ParamField>

```rust theme={null}
use ironclaw::llm::{FailoverProvider, CooldownConfig};
use std::time::Duration;

let failover = Arc::new(FailoverProvider::with_cooldown(
    vec![primary_llm, fallback_llm],
    CooldownConfig {
        cooldown_duration: Duration::from_secs(300),
        failure_threshold: 3,
    }
)?);
```

### CircuitBreakerProvider

Fast-fails when backend is degraded (circuit breaker pattern).

<ResponseField name="new" type="fn(provider: Arc<dyn LlmProvider>, config: CircuitBreakerConfig) -> Self">
  Wrap a provider with circuit breaker
</ResponseField>

#### CircuitBreakerConfig

<ParamField path="failure_threshold" type="usize" default="5">
  Failures before opening circuit
</ParamField>

<ParamField path="recovery_timeout" type="Duration" default="30s">
  How long to wait before testing recovery
</ParamField>

<ParamField path="half_open_max_calls" type="usize" default="3">
  Test calls allowed in half-open state
</ParamField>

```rust theme={null}
use ironclaw::llm::{CircuitBreakerProvider, CircuitBreakerConfig};

let cb_llm = Arc::new(CircuitBreakerProvider::new(
    base_llm,
    CircuitBreakerConfig {
        failure_threshold: 5,
        recovery_timeout: Duration::from_secs(30),
        ..Default::default()
    }
));
```

### CachedProvider

Caches responses to avoid redundant API calls.

<ResponseField name="new" type="fn(provider: Arc<dyn LlmProvider>, config: ResponseCacheConfig) -> Self">
  Wrap a provider with response caching
</ResponseField>

#### ResponseCacheConfig

<ParamField path="ttl" type="Duration" default="3600s">
  Cache entry time-to-live
</ParamField>

<ParamField path="max_entries" type="usize" default="1000">
  Maximum cache entries
</ParamField>

```rust theme={null}
use ironclaw::llm::{CachedProvider, ResponseCacheConfig};

let cached = Arc::new(CachedProvider::new(
    base_llm,
    ResponseCacheConfig {
        ttl: Duration::from_secs(3600),
        max_entries: 1000,
    }
));
```

### SmartRoutingProvider

Routes simple requests to cheap model, complex to primary.

<ResponseField name="new" type="fn(primary: Arc<dyn LlmProvider>, cheap: Arc<dyn LlmProvider>, config: SmartRoutingConfig) -> Self">
  Create smart routing between two providers
</ResponseField>

#### SmartRoutingConfig

<ParamField path="cascade_enabled" type="bool" default="true">
  Retry on primary if cheap model fails
</ParamField>

<ParamField path="complexity_threshold" type="f32" default="0.6">
  Complexity score above which to use primary (0.0 to 1.0)
</ParamField>

```rust theme={null}
use ironclaw::llm::{SmartRoutingProvider, SmartRoutingConfig};

let smart = Arc::new(SmartRoutingProvider::new(
    expensive_llm,
    cheap_llm,
    SmartRoutingConfig {
        cascade_enabled: true,
        complexity_threshold: 0.6,
    }
));
```

## Session Management

### SessionManager

Manages NEAR AI authentication sessions.

<ResponseField name="new" type="fn(config: SessionConfig) -> Self">
  Create a new session manager
</ResponseField>

<ResponseField name="get_or_create_session" type="async fn(&self) -> Result<String>">
  Get current session token or create new session
</ResponseField>

<ResponseField name="refresh_session" type="async fn(&self) -> Result<()>">
  Refresh the current session
</ResponseField>

#### SessionConfig

<ParamField path="session_path" type="PathBuf">
  Where to store session data
</ParamField>

<ParamField path="auth_base_url" type="String">
  NEAR AI auth endpoint
</ParamField>

## Error Handling

### LlmError

<ResponseField name="AuthFailed" type="{ provider: String }">
  Authentication failed
</ResponseField>

<ResponseField name="RequestFailed" type="{ provider: String, reason: String }">
  API request failed
</ResponseField>

<ResponseField name="RateLimited" type="{ retry_after: Option<Duration> }">
  Rate limit exceeded
</ResponseField>

<ResponseField name="InvalidResponse" type="String">
  Response parsing failed
</ResponseField>

<ResponseField name="ContextLengthExceeded" type="{ requested: usize, max: usize }">
  Request exceeds model context window
</ResponseField>

<ResponseField name="CircuitOpen" type="String">
  Circuit breaker is open
</ResponseField>

## Cost Tracking

### TokenUsage

Tracks token usage and costs.

<ParamField path="input_tokens" type="u32">
  Tokens in prompts
</ParamField>

<ParamField path="output_tokens" type="u32">
  Tokens in completions
</ParamField>

<ParamField path="total_cost" type="Decimal">
  Total cost in USD
</ParamField>

```rust theme={null}
use ironclaw::llm::TokenUsage;

let usage = TokenUsage {
    input_tokens: response.input_tokens,
    output_tokens: response.output_tokens,
    total_cost: calculate_cost(&response, &llm),
};
```

## Related Modules

<Card title="Agent Module" icon="robot" href="/api/agent">
  Agent reasoning and tool selection using LLMs
</Card>

<Card title="Workspace Module" icon="folder" href="/api/workspace">
  Workspace system prompts and memory context
</Card>
