Skip to main content

Overview

The LLM module provides integration with multiple language model providers:
  • NEAR AI (default): Session token or API key auth via Chat Completions API
  • OpenAI: Direct API access with your own key
  • Anthropic: Claude models via direct API
  • Ollama: Local model inference
  • OpenAI-compatible: Any endpoint that speaks the OpenAI API
  • Tinfoil: Private inference endpoints
Includes resilience features: retry logic, failover, circuit breakers, response caching, and smart routing.

Core Types

LlmProvider Trait

The main trait that all LLM providers must implement.

Methods

fn(&self) -> &str
Return the model identifier
fn(&self) -> (Decimal, Decimal)
Return (input_cost, output_cost) per token in USD
async fn(&self, request: CompletionRequest) -> Result<CompletionResponse>
Generate a completion for a conversation
async fn(&self, request: ToolCompletionRequest) -> Result<ToolCompletionResponse>
Generate a completion with tool calling support
async fn(&self) -> Result<Vec<String>>
List available models from the provider (optional)
async fn(&self) -> Result<ModelMetadata>
Fetch metadata about the model (context length, etc.)

ChatMessage

A message in a conversation.
Role
Message role (System, User, Assistant, Tool)
String
Message content/text
Option<String>
Tool call ID if this is a tool result message
Option<String>
Tool name for tool result messages
Option<Vec<ToolCall>>
Tool calls made by the assistant

Constructors

fn(content: impl Into<String>) -> Self
Create a system message
fn(content: impl Into<String>) -> Self
Create a user message
fn(content: impl Into<String>) -> Self
Create an assistant message
fn(content: Option<String>, tool_calls: Vec<ToolCall>) -> Self
Create an assistant message with tool calls
fn(tool_call_id: impl Into<String>, name: impl Into<String>, content: impl Into<String>) -> Self
Create a tool result message

Example

CompletionRequest

Request for a chat completion.
Vec<ChatMessage>
Conversation history
Option<String>
Optional per-request model override
Option<u32>
Maximum tokens to generate
Option<f32>
Sampling temperature (0.0 to 2.0)
Option<Vec<String>>
Sequences that stop generation
HashMap<String, String>
Opaque metadata passed to provider

Methods

fn(messages: Vec<ChatMessage>) -> Self
Create a new completion request
fn(self, model: impl Into<String>) -> Self
Set model override
fn(self, max_tokens: u32) -> Self
Set max tokens
fn(self, temperature: f32) -> Self
Set temperature

Example

CompletionResponse

Response from a chat completion.
String
Generated text content
u32
Number of tokens in the prompt
u32
Number of tokens generated
FinishReason
Why generation stopped

FinishReason

()
Model naturally completed the response
()
Hit max_tokens limit
()
Model wants to use tools
()
Filtered by content policy
()
Unknown reason

ToolCompletionRequest

Request for a completion with tool use support.
Vec<ChatMessage>
Conversation history
Vec<ToolDefinition>
Available tools for the model
Option<String>
Optional model override
Option<u32>
Maximum tokens to generate
Option<f32>
Sampling temperature
Option<String>
How to handle tools: “auto”, “required”, or “none”
HashMap<String, String>
Opaque metadata

Methods

fn(messages: Vec<ChatMessage>, tools: Vec<ToolDefinition>) -> Self
Create a new tool completion request
fn(self, choice: impl Into<String>) -> Self
Set tool choice mode (“auto”, “required”, “none”)

Example

ToolCompletionResponse

Response from a tool-enabled completion.
Option<String>
Text content (may be empty if tool calls present)
Vec<ToolCall>
Tool calls requested by the model
u32
Prompt tokens
u32
Generated tokens
FinishReason
Why generation stopped

ToolCall

A tool call requested by the LLM.
String
Unique call identifier
String
Tool name
serde_json::Value
Tool arguments as JSON

ToolDefinition

Definition of a tool for the LLM.
String
Tool name
String
What the tool does
serde_json::Value
JSON Schema for parameters

ToolResult

Result of tool execution to send back to the LLM.
String
ID of the tool call this is a result for
String
Tool name
String
Result content
bool
Whether this result is an error

Provider Creation

create_llm_provider

fn(config: &LlmConfig, session: Arc<SessionManager>) -> Result<Arc<dyn LlmProvider>>
Create an LLM provider based on configuration

build_provider_chain

fn(config: &LlmConfig, session: Arc<SessionManager>) -> Result<(Arc<dyn LlmProvider>, Option<Arc<dyn LlmProvider>>)>
Build the full provider chain with retry, failover, circuit breaker, cache, etc. Returns (main_provider, cheap_provider).
The provider chain applies decorators in this order:
  1. Raw provider (from config)
  2. RetryProvider (exponential backoff)
  3. SmartRoutingProvider (cheap/primary split)
  4. FailoverProvider (fallback model)
  5. CircuitBreakerProvider (fast-fail when degraded)
  6. CachedProvider (response cache)

Resilience Features

RetryProvider

Retries failed requests with exponential backoff.
fn(provider: Arc<dyn LlmProvider>, config: RetryConfig) -> Self
Wrap a provider with retry logic

RetryConfig

u32
default:"3"
Maximum number of retry attempts

FailoverProvider

Fails over to backup providers when primary fails.
fn(providers: Vec<Arc<dyn LlmProvider>>) -> Result<Self>
Create a failover provider with ordered list of providers
fn(providers: Vec<Arc<dyn LlmProvider>>, config: CooldownConfig) -> Result<Self>
Create failover with cooldown (temporary disable after failures)

CooldownConfig

Duration
default:"300s"
How long to disable a provider after failures
usize
default:"3"
Number of failures before cooldown

CircuitBreakerProvider

Fast-fails when backend is degraded (circuit breaker pattern).
fn(provider: Arc<dyn LlmProvider>, config: CircuitBreakerConfig) -> Self
Wrap a provider with circuit breaker

CircuitBreakerConfig

usize
default:"5"
Failures before opening circuit
Duration
default:"30s"
How long to wait before testing recovery
usize
default:"3"
Test calls allowed in half-open state

CachedProvider

Caches responses to avoid redundant API calls.
fn(provider: Arc<dyn LlmProvider>, config: ResponseCacheConfig) -> Self
Wrap a provider with response caching

ResponseCacheConfig

Duration
default:"3600s"
Cache entry time-to-live
usize
default:"1000"
Maximum cache entries

SmartRoutingProvider

Routes simple requests to cheap model, complex to primary.
fn(primary: Arc<dyn LlmProvider>, cheap: Arc<dyn LlmProvider>, config: SmartRoutingConfig) -> Self
Create smart routing between two providers

SmartRoutingConfig

bool
default:"true"
Retry on primary if cheap model fails
f32
default:"0.6"
Complexity score above which to use primary (0.0 to 1.0)

Session Management

SessionManager

Manages NEAR AI authentication sessions.
fn(config: SessionConfig) -> Self
Create a new session manager
async fn(&self) -> Result<String>
Get current session token or create new session
async fn(&self) -> Result<()>
Refresh the current session

SessionConfig

PathBuf
Where to store session data
String
NEAR AI auth endpoint

Error Handling

LlmError

{ provider: String }
Authentication failed
{ provider: String, reason: String }
API request failed
{ retry_after: Option<Duration> }
Rate limit exceeded
String
Response parsing failed
{ requested: usize, max: usize }
Request exceeds model context window
String
Circuit breaker is open

Cost Tracking

TokenUsage

Tracks token usage and costs.
u32
Tokens in prompts
u32
Tokens in completions
Decimal
Total cost in USD

Agent Module

Agent reasoning and tool selection using LLMs

Workspace Module

Workspace system prompts and memory context