18 KiB
Core Architecture
This document describes the agent-side core. For the tool-execution service, its 9P namespace, registry, metadata, sandbox, process lifecycle, and cancellation path, see architecture-toolsrv.md.
This document covers the internal Go packages that implement Ollie's agent runtime. The code lives under cmd/olliesrv/internal/; it is not a separately imported public library. olliesrv composes the agent, backend, session, filesystem, and tool-client packages.
Package Map
flowchart TB
subgraph Agent["cmd/olliesrv/internal/agent/"]
AGENT["agent.go — Agent struct, identity, events"]
LOOP["loop.go — main agent loop, streaming"]
TURN["turn.go — turn orchestration, Submit"]
DISPATCH["dispatch.go — tool execution, batching"]
HISTORY["history.go — message history, usage"]
COMPACT["compact.go — compaction, cold summarization"]
CACHE["cache.go — tool result caching"]
RETRY["retry.go — error tracking, retry logic"]
STATE["state.go, chat.go, peer.go, subagent.go"]
RUNTIME["runtime.go — Runtime, preamble, tool dispatch"]
CONFIG["agent_config.go — profile configuration"]
PROMPT["prompt_resolver.go — prompt resolution"]
end
subgraph Backend["cmd/olliesrv/internal/backend/"]
BACKEND["backend.go — Backend interface and streaming"]
PROVIDERS["openai, anthropic, ollama, gemini, copilot, kiro"]
POOL["pool.go — backend pool"]
end
subgraph ToolClient["cmd/olliesrv/internal/toolclient/"]
CLIENT["9P toolsrv client"]
SPAWN["spawn.go — local/remote toolsrv lifecycle"]
end
subgraph Session["cmd/olliesrv/internal/session/"]
SESS["session.go, registry.go — sessions and agents"]
PERSIST["persist.go — session persistence"]
end
subgraph Fs["cmd/olliesrv/internal/fs/"]
ROOT["newroot.go — NewRoot, RestoreSessions"]
SPEC["spec.go — 9P namespace declaration"]
SUPPORT["support.go — namespace support"]
end
subgraph Toolsrv["cmd/toolsrv/internal/"]
SERVER["server — 9P tool server and proc handlers"]
REGISTRY["registry — per-agent loaded tools"]
EXEC["exec — tool process execution"]
SANDBOX["sandbox — native Landlock policy"]
end
Agent --> Backend
Agent --> ToolClient
Agent --> Session
ToolClient --> Toolsrv
Fs --> Session
Toolsrv --> REGISTRY
Toolsrv --> EXEC
EXEC --> SANDBOX
The Agent Struct (cmd/olliesrv/internal/agent/agent.go)
Agent owns one agent identity inside a session. Its runtime is swappable: /agent loads a new profile, creates a fresh toolsrv connection, rebuilds the runtime, and clears history. The session remains the stable host.
type Agent struct {
history *History
runtime *Runtime
name string
profile string
agentsDir string
currentAction atomic.Pointer[actionHandle]
fifo Fifo
Feed Feed
output EventHandler
sessionID string
id string
parentID string
cwd string
chatLog []byte
// synchronization, persistence, and observation state
}
func NewAgent(cfg AgentParams) *Agent
All output flows through the agent's EventHandler as Event{Role, Name, Content, ResponseID, OutputFormat} values. Events include user, assistant, reasoning, call, tool, usage, state, limitretry, retry, maxsteps, error, and info. Session-level events use a separate topic-based pubsub bus for 9P observers.
Agent Construction (cmd/olliesrv/internal/agent/agent.go, runtime.go, agent_config.go)
AgentConfig
The full construction config:
| Field | Purpose |
|---|---|
History |
Pre-existing session history to restore (nil = new) |
Runtime |
Pre-built *Runtime (tools, hooks, prompt, params) |
Profile |
Name of the active agent configuration profile |
ID |
Immutable agent ID (uname) |
Name |
Mutable display name |
AgentsDir |
Path to agent JSON directory |
SessionsDir |
Path to persisted session files |
SessionID |
Unique session identifier |
AgentID |
Immutable agent principal |
CWD |
Working directory for tools and prompt |
NewToolServer |
Factory for fresh ToolsrvConn on agent switch |
NewBackend |
Factory for backend by name |
Log |
Logger instance |
MaxSteps |
Override agent config's maxSteps |
ReadPlanStep |
Function to read next unchecked plan step |
ListHandlers |
Additional list handlers |
Runtime
Runtime contains the active backend, authenticated 9P toolsrv connection, rendered preamble, tool metadata, executor closure, generation parameters, max-step limit, compaction model, resolved user prompt, and startup messages. BuildRuntime lists loaded tools, resolves agent and user prompts, injects dispatch flags into tool schemas, and routes calls through ToolsrvConn.CallTool(ctx, ...).
type Runtime struct {
Backend backend.Backend
ToolServer *toolclient.ToolsrvConn
Preamble *Preamble
Tools []backend.Tool
ToolMeta map[string]protocol.ToolInfo
ToolRevision uint64
Exec toolExecutor
GenParams backend.GenerationParams
MaxSteps int
CompactionModel string
UserPrompt string
Messages []string
}
BuildRuntime
BuildRuntime(cfg, runner, cwd, env) wires everything together:
- Lists the agent’s currently loaded tools from the runner and records the agent-scoped registry revision
- Resolves the system prompt by executing the config's
promptcommands - Extracts generation parameters from the config
- Builds the tool executor closure (routes calls through the runner)
- Extracts parallelism and tier classifiers
- Applies tool/executor restrictions if configured
- Returns a
*Runtimestruct ready forNew
Agent construction
NewAgent(AgentParams) creates an agent with identity, session, profile, history, runtime, working directory, persistence callbacks, event handler state, chat-log state, and session environment. The configuration type on disk is AgentConfig; construction dependencies are carried separately by AgentParams.
The Agent Loop (cmd/olliesrv/internal/agent/)
The agent package is organized by concern:
| File | Responsibility |
|---|---|
loop.go |
Main loop: stream LLM, execute tools, update history |
turn.go |
Submit entry point, turn orchestration |
dispatch.go |
Tool execution, batching, conflict detection |
history.go |
Message history, usage tracking |
compact.go |
Context compaction, cold summarization |
cache.go |
Tool result caching with file staleness |
retry.go |
Error tracking, transient retry logic |
state.go |
State management, signals, WaitChange |
chat.go |
Chat log, streaming output |
peer.go |
Peer agent management |
subagent.go |
Sub-agent depth tracking |
The run() function in loop.go is the core execution engine. It loops until the model stops calling tools.
Per-Iteration Flow
flowchart TB
START["Start turn"]
GATE["Context gate:\nstrip cold material\nif over budget"]
BUILD["Build message history:\nsystem prompt + task state + hot tail"]
STREAM["Stream from backend\n(with retry logic)"]
CHECK{"Tool calls\nin response?"}
EXEC["Execute tool calls\n(serial or parallel)"]
UPDATE["Update state:\nappend messages + results"]
INFER["Infer TaskState\nupdates"]
SAFETY["Safety checks:\nerror limits, stall,\nreplan, max steps"]
COMPACT["Auto-compact\nif needed"]
DONE["Turn complete"]
START --> GATE --> BUILD --> STREAM --> CHECK
CHECK -- yes --> EXEC --> UPDATE --> INFER --> SAFETY --> COMPACT --> BUILD
CHECK -- no --> DONE
Streaming & Retry
The streaming loop handles three failure modes:
| Failure | Detection | Recovery |
|---|---|---|
| Rate limit (429) | *RateLimitError |
Retry using provider Retry-After when available, otherwise exponential backoff |
| Transient (5xx/network) | *TransientError |
Retry with exponential backoff |
| Stream drop | Stream closes without completion | Retry |
| Context overflow | *ContextOverflowError |
Compact and retry once |
| Tool unsupported | *ToolUnsupportedError |
Return through normal turn error handling |
On the first error, the turnError hook fires. If it exits 0 (handled), the loop returns immediately — the hook is responsible for recovery (e.g., switching models via freeloader).
Tool Execution
Tool calls are dispatched with parallelism awareness. Tool schemas are
loaded progressively: the default profile starts with client_9p, and semantic
tool hints tell the agent to issue tool_load <name> through its agent ctl
file.
After a tool round, the agent compares the toolsrv agent-scoped registry
revision with Runtime.ToolRevision. If the revision changed, it lists the
loaded tools, rebuilds backend schemas and metadata, and updates the rendered
tool preamble before the next model request.
- Classify each call from the tool metadata and dispatch flags.
- Batch read-safe calls when the metadata permits parallel execution.
- Execute calls through the runtime executor, with per-turn caching for eligible file reads.
- Hard-limit model-facing tool output to 32 KiB.
- Assign each result a hot, warm, or cold memory tier for context management.
Tool calls use the current turn context, so interrupting a turn cancels in-flight tool execution.
Safety Mechanisms
| Mechanism | Threshold | Action |
|---|---|---|
| Consecutive errors (soft) | 5 rounds | Inject nudge: "try a different approach" |
| Consecutive errors (hard) | 10 rounds | Abort the turn |
| Replan gate | 8 rounds without PLAN: block | Inject replan nudge |
| Stall detection | 5 rounds with same LastAction | Inject stall nudge |
| Max steps | Configurable per-agent | Soft exit with nudge |
| Plan re-injection | Every 10 tool rounds | Re-inject task state |
TaskState Inference
After each tool round, inferTaskStateUpdate updates the structured overlay:
LastAction← most recent tool call + truncated resultPlanStep← next unchecked item from the plan file (or heuristic from assistant text)
Session & Context (cmd/olliesrv/internal/agent/history.go, compact.go)
Message History
History holds a flat []backend.Message slice. No rolling window — grows until compaction.
Checkpoint: History.Checkpoint(ts TaskState) forks the session — returns a new History that inherits the given TaskState but starts with a clean message history. Used for narrow-context sub-agents that know what to do without inheriting parent message noise.
Token tracking per session:
TotalInputTokens,TotalOutputTokens,TotalRequestsTotalCachedInputTokens,TotalCacheCreationTokens- Per-turn accumulators reset at turn start
LastTurnCostUSD,SessionCostUSD
Compaction
Triggered when estimated tokens reach 50% of model context length.
Three-zone strategy:
| Zone | Content | Size |
|---|---|---|
| Cold | Structured TaskState JSON | 1 message |
| Warm | One-line decision index | ~10 messages |
| Hot | Verbatim recent messages | 8 messages |
Process:
- Flatten tool call/result pairs to plain text
- Send full history + compaction prompt to the LLM
- Parse structured JSON response into TaskState
- Rebuild history as cold + warm + hot zones
- Write pre-compaction snapshot to
{id}.compaction.jsonlfor recovery
Result Tiers
Tool results are classified at execution time:
| Tier | Behavior | Examples |
|---|---|---|
| Hot | Stays verbatim in messages | file_write, file_edit, reasoning_think |
| Warm | Summarized on next compaction | (custom tools) |
| Cold | Immediately collapsed to one-line summary | file_read, memory_recall, lsp_* |
- The tool's
tiermetadata (cold,warm, orhot) hotwhen no tier metadata is present
Persistence
Sessions are saved as PersistedSession JSON:
{
"id": "1716300000000-abc123",
"agent": "coding",
"messages": [...],
"taskState": {...}
}
Saved after every tool round (saveSession callback) and on clean exit.
Hooks and lifecycle callbacks
Hooks are shell commands executed with:
- stdin: JSON payload with context (session_id, cwd, model, error, etc.)
- env:
OLLIE_*variables derived from the payload map - timeout: 60 seconds (configurable in tests)
- process group:
Setpgid=truefor clean kill on timeout
Exit Code Semantics
| Code | Meaning |
|---|---|
| 0 | Success. Stdout appended to context. |
| 2 | Block. Stops the action (e.g., blocks a tool call, prevents compaction). |
| Other | Non-blocking warning. Execution continues. |
Hook Execution Order
Multiple commands per hook run sequentially. If any exits 2, execution stops immediately and returns Blocked=true.
Commands (cmd/olliesrv/internal/agent/commands.go)
HandleCommand forwards slash-command text to the agent ctl file over 9P. The current command set is defined by the filesystem control handler; examples include profile/backend/model changes, compaction, context and usage inspection, session lifecycle, tool/skill listing, prompt injection, queue management, and working-directory changes. Inject queues an interruption for the running turn; InjectRewrite replaces the pending interruption.
Agent Switching
When /agent <name> is issued:
- Load the new agent config JSON
- Create a fresh runner via
newToolServer() - Call
BuildRuntimewith the new config - If the config specifies a backend/model, switch those too
- Replace the entire
Runtimepointer on the agent struct - Clear the session (fresh context for the new persona)
Backend System (cmd/olliesrv/internal/backend/)
Shared Types
| Type | Purpose |
|---|---|
Message |
Conversation turn (role, content, tool_calls, tool_call_id) |
Tool |
Function declaration (name, description, JSON Schema params) |
ToolCall |
Model's request to invoke a function |
StreamEvent |
Incremental streaming delta (content, reasoning, tool_calls, done, usage) |
Usage |
Token counts + optional cost |
GenerationParams |
Unified sampling parameters |
Error Types
| Error | HTTP Code | Loop Behavior |
|---|---|---|
RateLimitError |
429 | Retry with backoff |
TransientError |
5xx, network | Retry with backoff |
ContextOverflowError |
400 (specific) | Compact and retry |
ToolUnsupportedError |
400/422 (specific) | Propagate to turnError hook |
streamRequest
Shared HTTP streaming infrastructure used by all backends:
- Execute the HTTP request
- Classify response status into typed errors
- Spawn a goroutine that parses SSE/NDJSON into
StreamEventchannel - Return the channel for the loop to consume
Backend Implementations
OllamaBackend (/api/chat):
- Local inference, no API key
- Context length from
/api/show(cached) - Model list from
/api/tags - NDJSON streaming
OpenAIBackend (/v1/chat/completions):
- Works with OpenAI, OpenRouter, any compatible API
- Configurable base URL, name, extra headers
- Context length from model list endpoint (cached)
- SSE streaming (
data: {...}lines) - Handles both
tool_callsand legacyfunction_call
AnthropicBackend (/v1/messages):
- Extended thinking (reasoning budget via
ThinkingBudget) - Prompt caching (cache_control blocks)
- 200k context window (hardcoded)
- SSE streaming with content_block_delta events
GeminiBackend (Gemini API):
- Google Gemini models
CopilotBackend (GitHub Copilot API):
- Copilot Chat integration
CodeWhispererBackend (AWS CodeWhisperer API):
- AWS CodeWhisperer integration
Config System (cmd/olliesrv/internal/agent/agent_config.go)
Agent Config Schema
The on-disk AgentConfig currently contains prompt, userPrompts, backend, model, tools, autoLoad, maxSteps, compactionModel, systemPrompt, and embedded backend GenerationParams. Tool and prompt resolution are performed while building the runtime.
Prompt Type
Dual-mode unmarshaling:
- JSON string: treated as literal text (with
!command, file path, or env expansion) - JSON array: each element is a shell command; stdout concatenated
The array form is preferred — it enables composable prompt assembly from template fragments.
HookCmds Type
Dual-mode: a single string "cmd" or an array ["cmd1", "cmd2"]. Multiple commands run sequentially.
Prompt Resolution
Prompt assembly and resolution are documented in architecture-prompting.md. The agent runtime consumes the resulting preamble and per-turn user prompt.
Cost Tracking (cmd/olliesrv/internal/agent/cost.go)
When the backend doesn't report cost directly (only OpenRouter does), cost is computed from a built-in pricing table:
- Matches model name prefix (case-insensitive)
- Covers Anthropic Claude, OpenAI GPT/o-series, Google Gemini
- Cached input tokens charged at discount (10% Anthropic, 50% others)
- Cache creation tokens charged at 125% of input rate
- Unknown/local models return zero cost
Sandbox
Sandbox policy is owned and enforced by toolsrv. See architecture-toolsrv.md.
Event System
The agent emits events through its EventHandler. Session-level observers use session.PublishEvent, SubscribeEvents, and topic wildcards. The filesystem layer converts session events into chat and blocking event files.
Concurrency Model
- One turn at a time:
submitMuserializesSubmit()calls - Interrupt:
currentActionholds aCancelCauseFunc;Interrupt()calls it - Queue drain: after each turn completes, queued prompts are drained sequentially
- State observation: agent state and chat changes signal 9P waiters through change channels and condition variables
- Tool parallelism: read-safe tools fan out within a turn when metadata permits; result caching and
/tmp/ollie/{sessionID}locks coordinate execution - Peer messaging: writing to
peer/{name}triggers an asyncSubmit()on the target agent; delivery is fire-and-forget within the session
Session ID Format
<unix-nanoseconds>-<6-hex-chars>
Lexicographically sortable by creation time. Unique even within the same nanosecond due to the random suffix.