Overview
The LLM module provides integration with multiple language model providers:- NEAR AI (default): Session token or API key auth via Chat Completions API
- OpenAI: Direct API access with your own key
- Anthropic: Claude models via direct API
- Ollama: Local model inference
- OpenAI-compatible: Any endpoint that speaks the OpenAI API
- Tinfoil: Private inference endpoints
Core Types
LlmProvider Trait
The main trait that all LLM providers must implement.Methods
fn(&self) -> &str
Return the model identifier
fn(&self) -> (Decimal, Decimal)
Return (input_cost, output_cost) per token in USD
async fn(&self, request: CompletionRequest) -> Result<CompletionResponse>
Generate a completion for a conversation
async fn(&self, request: ToolCompletionRequest) -> Result<ToolCompletionResponse>
Generate a completion with tool calling support
async fn(&self) -> Result<Vec<String>>
List available models from the provider (optional)
async fn(&self) -> Result<ModelMetadata>
Fetch metadata about the model (context length, etc.)
ChatMessage
A message in a conversation.Role
Message role (System, User, Assistant, Tool)
String
Message content/text
Option<String>
Tool call ID if this is a tool result message
Option<String>
Tool name for tool result messages
Option<Vec<ToolCall>>
Tool calls made by the assistant
Constructors
fn(content: impl Into<String>) -> Self
Create a system message
fn(content: impl Into<String>) -> Self
Create a user message
fn(content: impl Into<String>) -> Self
Create an assistant message
fn(content: Option<String>, tool_calls: Vec<ToolCall>) -> Self
Create an assistant message with tool calls
fn(tool_call_id: impl Into<String>, name: impl Into<String>, content: impl Into<String>) -> Self
Create a tool result message
Example
CompletionRequest
Request for a chat completion.Vec<ChatMessage>
Conversation history
Option<String>
Optional per-request model override
Option<u32>
Maximum tokens to generate
Option<f32>
Sampling temperature (0.0 to 2.0)
Option<Vec<String>>
Sequences that stop generation
HashMap<String, String>
Opaque metadata passed to provider
Methods
fn(messages: Vec<ChatMessage>) -> Self
Create a new completion request
fn(self, model: impl Into<String>) -> Self
Set model override
fn(self, max_tokens: u32) -> Self
Set max tokens
fn(self, temperature: f32) -> Self
Set temperature
Example
CompletionResponse
Response from a chat completion.String
Generated text content
u32
Number of tokens in the prompt
u32
Number of tokens generated
FinishReason
Why generation stopped
FinishReason
()
Model naturally completed the response
()
Hit max_tokens limit
()
Model wants to use tools
()
Filtered by content policy
()
Unknown reason
ToolCompletionRequest
Request for a completion with tool use support.Vec<ChatMessage>
Conversation history
Vec<ToolDefinition>
Available tools for the model
Option<String>
Optional model override
Option<u32>
Maximum tokens to generate
Option<f32>
Sampling temperature
Option<String>
How to handle tools: “auto”, “required”, or “none”
HashMap<String, String>
Opaque metadata
Methods
fn(messages: Vec<ChatMessage>, tools: Vec<ToolDefinition>) -> Self
Create a new tool completion request
fn(self, choice: impl Into<String>) -> Self
Set tool choice mode (“auto”, “required”, “none”)
Example
ToolCompletionResponse
Response from a tool-enabled completion.Option<String>
Text content (may be empty if tool calls present)
Vec<ToolCall>
Tool calls requested by the model
u32
Prompt tokens
u32
Generated tokens
FinishReason
Why generation stopped
ToolCall
A tool call requested by the LLM.String
Unique call identifier
String
Tool name
serde_json::Value
Tool arguments as JSON
ToolDefinition
Definition of a tool for the LLM.String
Tool name
String
What the tool does
serde_json::Value
JSON Schema for parameters
ToolResult
Result of tool execution to send back to the LLM.String
ID of the tool call this is a result for
String
Tool name
String
Result content
bool
Whether this result is an error
Provider Creation
create_llm_provider
fn(config: &LlmConfig, session: Arc<SessionManager>) -> Result<Arc<dyn LlmProvider>>
Create an LLM provider based on configuration
build_provider_chain
fn(config: &LlmConfig, session: Arc<SessionManager>) -> Result<(Arc<dyn LlmProvider>, Option<Arc<dyn LlmProvider>>)>
Build the full provider chain with retry, failover, circuit breaker, cache, etc. Returns (main_provider, cheap_provider).
- Raw provider (from config)
- RetryProvider (exponential backoff)
- SmartRoutingProvider (cheap/primary split)
- FailoverProvider (fallback model)
- CircuitBreakerProvider (fast-fail when degraded)
- CachedProvider (response cache)
Resilience Features
RetryProvider
Retries failed requests with exponential backoff.fn(provider: Arc<dyn LlmProvider>, config: RetryConfig) -> Self
Wrap a provider with retry logic
RetryConfig
u32
default:"3"
Maximum number of retry attempts
FailoverProvider
Fails over to backup providers when primary fails.fn(providers: Vec<Arc<dyn LlmProvider>>) -> Result<Self>
Create a failover provider with ordered list of providers
fn(providers: Vec<Arc<dyn LlmProvider>>, config: CooldownConfig) -> Result<Self>
Create failover with cooldown (temporary disable after failures)
CooldownConfig
Duration
default:"300s"
How long to disable a provider after failures
usize
default:"3"
Number of failures before cooldown
CircuitBreakerProvider
Fast-fails when backend is degraded (circuit breaker pattern).fn(provider: Arc<dyn LlmProvider>, config: CircuitBreakerConfig) -> Self
Wrap a provider with circuit breaker
CircuitBreakerConfig
usize
default:"5"
Failures before opening circuit
Duration
default:"30s"
How long to wait before testing recovery
usize
default:"3"
Test calls allowed in half-open state
CachedProvider
Caches responses to avoid redundant API calls.fn(provider: Arc<dyn LlmProvider>, config: ResponseCacheConfig) -> Self
Wrap a provider with response caching
ResponseCacheConfig
Duration
default:"3600s"
Cache entry time-to-live
usize
default:"1000"
Maximum cache entries
SmartRoutingProvider
Routes simple requests to cheap model, complex to primary.fn(primary: Arc<dyn LlmProvider>, cheap: Arc<dyn LlmProvider>, config: SmartRoutingConfig) -> Self
Create smart routing between two providers
SmartRoutingConfig
bool
default:"true"
Retry on primary if cheap model fails
f32
default:"0.6"
Complexity score above which to use primary (0.0 to 1.0)
Session Management
SessionManager
Manages NEAR AI authentication sessions.fn(config: SessionConfig) -> Self
Create a new session manager
async fn(&self) -> Result<String>
Get current session token or create new session
async fn(&self) -> Result<()>
Refresh the current session
SessionConfig
PathBuf
Where to store session data
String
NEAR AI auth endpoint
Error Handling
LlmError
{ provider: String }
Authentication failed
{ provider: String, reason: String }
API request failed
{ retry_after: Option<Duration> }
Rate limit exceeded
String
Response parsing failed
{ requested: usize, max: usize }
Request exceeds model context window
String
Circuit breaker is open
Cost Tracking
TokenUsage
Tracks token usage and costs.u32
Tokens in prompts
u32
Tokens in completions
Decimal
Total cost in USD
Related Modules
Agent Module
Agent reasoning and tool selection using LLMs
Workspace Module
Workspace system prompts and memory context