24 KiB
ollie-core: Detailed Architecture
This document covers the internal design of the Go library — the packages that implement the agent engine. It is the dependency imported by all consumers (olliesrv, frontends, tests).
Package Map
flowchart TB
subgraph Agent["agent/ package"]
AGENT["agent.go\nAgent struct, New()"]
LOOP["loop.go\nrun() — agent loop"]
HISTORY["history.go\nHistory, TaskState, compaction"]
TURN["turn.go\nper-turn state machine"]
HOOKS["hooks.go\nlifecycle callbacks"]
CMDS["commands.go\n/command dispatch"]
BUILD["build_runtime.go\nBuildRuntime()"]
PROMPT["prompt_resolver.go\nsystem prompt assembly"]
STATE["state.go\nResultTier, toolResult"]
COST["cost.go\npricing table"]
FIFO["fifo.go\nprompt queue"]
COMPLETE["complete.go\ncode completion"]
CONFIG["agent_config.go\nConfig schema"]
CONFIG_PATHS["config_paths.go\nXDG config discovery"]
RUNTIME["runtime.go\nRuntime struct"]
NEW["new.go\nNew() factory"]
USAGE["usage_log.go\nusage tracking"]
end
subgraph Backend["backend/ package"]
BACKEND["backend.go\nBackend interface"]
OPENAI["openai.go"]
ANTHROPIC["anthropic.go"]
OLLAMA["ollama.go"]
GEMINI["gemini.go"]
COPILOT["copilot.go"]
CW["codewhisperer.go"]
POOL["pool.go\nbackend pool"]
end
subgraph Toolsrv["toolsrv/ package"]
SERVER["server.go\nServer struct"]
SHELL["shell.go\nshell execution"]
TOOLS["tools.go\nRunner interface"]
REGISTRY["registry.go\ntool discovery"]
DISCOVER["discover.go\ntool script scanning"]
SCHEMA["schema.go\nJSON Schema parsing"]
TIER["tier.go\nresult tiering"]
SKILLS["skills.go\nskill management"]
STREAM["stream.go\noutput streaming"]
REMOTE["remote.go\nRemoteServer"]
DIAL["dial.go\nSSH bootstrap"]
ACCESSORS["accessors.go\nIsParallelRead, ResultTier"]
BOOTSTRAP["bootstrap.sh\nSSH bootstrap script"]
end
subgraph Session["session/ package"]
SESS["session.go\nSession, Config"]
end
subgraph Fs["fs/ package"]
TREE["tree.go\n*Tree"]
NEWROOT["newroot.go\nNewRoot()"]
LIFECYCLE["lifecycle.go\nCreate/Kill/Rename"]
SESSFILES["sessionfiles.go\nsession file handlers"]
AGENTFILES["agentfiles.go\nagent file handlers"]
ROOTFILES["rootfiles.go\nroot file handlers"]
ELEVATEFILES["elevatefiles.go\nelevation handlers"]
PROCFILES["procfiles.go\nprocess handlers"]
PERSIST["persist.go\nsession persistence"]
FS_GO["fs.go\n9P File/FileConfig"]
FORMAT["format.go\nevent formatting"]
TYPES["types.go\nSession struct"]
SPEC["spec.go\nEDSL namespace declaration"]
BUILDER["builder.go\nBuildTree"]
FSNODE["fsnode.go\nFsNodeDecl"]
end
subgraph Support["Supporting Packages"]
DETACH["detach/\nprocess management"]
ELEVATE["elevate/\nelevation broker"]
SANDBOX["sandbox/\nLandlock config"]
ENV["env/\nenvironment"]
LOG["log/\nstructured logging"]
PATHS["paths/\nXDG resolution"]
end
Agent --> Backend
Agent --> Toolsrv
Agent --> Session
Fs --> Session
Toolsrv --> SANDBOX
Toolsrv --> DETACH
Toolsrv --> ELEVATE
The Agent Struct (agent/agent.go)
Agent is the boundary between the engine and any consumer. It is a concrete struct, not an interface — the interface was removed during the Great Flattening because no consumer ever implemented it independently.
type Agent struct {
history *History
runtime *Runtime
cfg agentConfig
agentName string
agentsDir string
baseLayers []string
promptEnvExtra []string
newToolServer func() toolsrv.Runner
newBackend func(string) (backend.Backend, error)
currentAction atomic.Pointer[actionHandle]
// ... execution state, sync primitives, session metadata
}
func New(cfg AgentConfig) *Agent
All output flows through the event bus as Event{Role, Name, Content} values. Roles include: assistant, tool, call, info, exec, error, usage, reasoning, limitretry.
Agent Construction (agent/new.go)
AgentConfig
The full construction config:
| Field | Purpose |
|---|---|
History |
Pre-existing session history to restore (nil = new) |
Runtime |
Pre-built *Runtime (tools, hooks, prompt, params) |
AgentName |
Name of the active agent config |
AgentsDir |
Path to agent JSON directory |
SessionsDir |
Path to persisted session files |
SessionID |
Unique session identifier |
AgentID |
Immutable agent principal |
CWD |
Working directory for tools and prompt |
NewToolServer |
Factory for fresh toolsrv.Runner on agent switch |
NewBackend |
Factory for backend by name |
Log |
Logger instance |
MaxSteps |
Override agent config's maxSteps |
ReadPlanStep |
Function to read next unchecked plan step |
ListHandlers |
Additional list handlers |
Runtime
The Runtime struct holds all per-agent configuration that changes on an /agent switch but remains stable across turns. The agent struct stores a pointer to the active Runtime; switching agents replaces it atomically.
type Runtime struct {
Backend backend.Backend
Runner toolsrv.Runner
Hooks Hooks
Preamble string // compiled system prompt
Tools []backend.Tool
ClassifyTool func(string) bool // parallelism check
ClassifyTier func(string, json.RawMessage) string
GenParams backend.GenerationParams
MaxSteps int
ToolResultMaxBytes int
CfgBackend string
CfgModel string
Messages []string // non-fatal warnings
}
BuildRuntime
BuildRuntime(cfg, runner, cwd, env) wires everything together:
- Lists all tools from the runner
- Resolves the system prompt by executing the config's
promptcommands - Extracts generation parameters from the config
- Builds the tool executor closure (routes calls through the runner)
- Extracts parallelism and tier classifiers
- Applies tool/executor restrictions if configured
- Returns a
*Runtimestruct ready forNew
New
Creates the Agent struct from AgentConfig:
- Sweeps stale tmpdirs from previous crashes
- Stores the provided
Runtime(or creates an empty one) - Sets the backend model if
ModelNameis specified - Creates
/tmp/ollie/{sessionID}for this session - Sets up the plan file reader
- Wires the
turnErrorhook with a detached context - Pushes
OLLIE_SESSION_IDandOLLIE_AGENT_IDinto the runner env - Sets the flock directory for parallel coordination
The Agent Loop (agent/loop.go)
The run() function is the core execution engine. It takes an agentConfig and a state interface, and loops until the model stops calling tools.
Per-Iteration Flow
flowchart TB
START["Start turn"]
GATE["Context gate:\nstrip cold material\nif over budget"]
BUILD["Build message history:\nsystem prompt + task state + hot tail"]
STREAM["Stream from backend\n(with retry logic)"]
CHECK{"Tool calls\nin response?"}
EXEC["Execute tool calls\n(serial or parallel)"]
UPDATE["Update state:\nappend messages + results"]
INFER["Infer TaskState\nupdates"]
SAFETY["Safety checks:\nerror limits, stall,\nreplan, max steps"]
COMPACT["Auto-compact\nif needed"]
DONE["Turn complete"]
START --> GATE --> BUILD --> STREAM --> CHECK
CHECK -- yes --> EXEC --> UPDATE --> INFER --> SAFETY --> COMPACT --> BUILD
CHECK -- no --> DONE
Streaming & Retry
The streaming loop handles three failure modes:
| Failure | Detection | Recovery |
|---|---|---|
| Rate limit (429) | *RateLimitError |
Exponential backoff (5x2^attempt seconds), up to 3 retries |
| Transient (5xx, network) | *TransientError |
Same retry logic |
| Stream drop | Done never received, channel closes |
Retry with 2x2^attempt second delay |
| Context overflow | *ContextOverflowError |
Trigger compaction, retry |
| Tool unsupported | *ToolUnsupportedError |
Propagate to turnError hook |
On the first error, the turnError hook fires. If it exits 0 (handled), the loop returns immediately — the hook is responsible for recovery (e.g., switching models via freeloader).
Tool Execution
Tool calls are dispatched with parallelism awareness:
- Classify each tool call via
ClassifyTool(name)(readsreadOnlyfrom.metafile) - Batch consecutive parallel-read-safe calls into a group
- Fan out the batch concurrently with deduplication (identical calls share one execution)
- Serial calls run one at a time
Per-call lifecycle:
preTool hook → execute → postTool hook → inject check → truncation → cache (if read-safe)
preToolexit 2 blocks executionpostToolexit 2 replaces the result- Results exceeding
ToolResultMaxBytesare hard-truncated - Read-safe results are cached in a
sync.Mapfor the duration of the turn
Safety Mechanisms
| Mechanism | Threshold | Action |
|---|---|---|
| Consecutive errors (soft) | 5 rounds | Inject nudge: "try a different approach" |
| Consecutive errors (hard) | 10 rounds | Abort the turn |
| Replan gate | 8 rounds without PLAN: block | Inject replan nudge |
| Stall detection | 5 rounds with same LastAction | Inject stall nudge |
| Max steps | Configurable per-agent | Soft exit with nudge |
| Plan re-injection | Every 10 tool rounds | Re-inject task state |
TaskState Inference
After each tool round, inferTaskStateUpdate updates the structured overlay:
LastAction← most recent tool call + truncated resultPlanStep← next unchecked item from the plan file (or heuristic from assistant text)
Session & Context (agent/history.go)
Message History
History holds a flat []backend.Message slice. No rolling window — grows until compaction.
Checkpoint: History.Checkpoint(ts TaskState) forks the session — returns a new History that inherits the given TaskState but starts with a clean message history. Used for narrow-context sub-agents that know what to do without inheriting parent message noise.
Token tracking per session:
TotalInputTokens,TotalOutputTokens,TotalRequestsTotalCachedInputTokens,TotalCacheCreationTokens- Per-turn accumulators reset at turn start
LastTurnCostUSD,SessionCostUSD
Compaction
Triggered when estimated tokens reach 50% of model context length.
Three-zone strategy:
| Zone | Content | Size |
|---|---|---|
| Cold | Structured TaskState JSON | 1 message |
| Warm | One-line decision index | ~10 messages |
| Hot | Verbatim recent messages | 8 messages |
Process:
- Flatten tool call/result pairs to plain text
- Send full history + compaction prompt to the LLM
- Parse structured JSON response into TaskState
- Rebuild history as cold + warm + hot zones
- Write pre-compaction snapshot to
{id}.compaction.jsonlfor recovery
Result Tiers
Tool results are classified at execution time:
| Tier | Behavior | Examples |
|---|---|---|
| Hot | Stays verbatim in messages | file_write, file_edit, reasoning_think |
| Warm | Summarized on next compaction | (custom tools) |
| Cold | Immediately collapsed to one-line summary | file_read, memory_recall, lsp_* |
Classification sources (in priority order):
- Built-in table in
toolsrv/tier.go(defaultColdTools) tierfield in the tool's.metasidecar file- Default: hot
Persistence
Sessions are saved as PersistedSession JSON:
{
"id": "1716300000000-abc123",
"agent": "coding",
"messages": [...],
"taskState": {...}
}
Saved after every tool round (saveSession callback) and on clean exit.
Hooks (agent/hooks.go)
Hooks are shell commands executed with:
- stdin: JSON payload with context (session_id, cwd, model, error, etc.)
- env:
OLLIE_*variables derived from the payload map - timeout: 60 seconds (configurable in tests)
- process group:
Setpgid=truefor clean kill on timeout
Exit Code Semantics
| Code | Meaning |
|---|---|
| 0 | Success. Stdout appended to context. |
| 2 | Block. Stops the action (e.g., blocks a tool call, prevents compaction). |
| Other | Non-blocking warning. Execution continues. |
Hook Execution Order
Multiple commands per hook run sequentially. If any exits 2, execution stops immediately and returns Blocked=true.
Commands (agent/commands.go)
Submit() first checks for commands before starting an agent turn:
| Prefix | Dispatch |
|---|---|
!<cmd> |
Shell command in session CWD |
/backend [name] |
Show or switch LLM backend |
/model [name] |
Show or switch model |
/models |
List available models |
/agent [name] |
Show or switch agent config (rebuilds Runtime) |
/agents |
List available agents |
/maxsteps [n] |
Get/set max steps (0 = unlimited) |
/compact |
Manual compaction |
/context |
Show context size and message breakdown |
/usage |
Show token usage and context percentage |
/cost |
Show last turn and session cost |
/history |
Dump bounded message history |
/sessions |
List saved sessions |
/clear |
Clear session (reset context) |
/kill |
Kill session |
/rn <name> |
Rename session |
/cwd [path] |
Show or change working directory |
/skills |
List available skills |
/tools |
List available tools |
/sp |
Show rendered system prompt |
/i <prompt> |
Inject into running turn |
/irw <prompt> |
Inject-rewrite |
| `/queued [pop | clear]` |
/help |
Show command help |
Agent Switching
When /agent <name> is issued:
- Load the new agent config JSON
- Create a fresh runner via
newToolServer() - Call
BuildRuntimewith the new config - If the config specifies a backend/model, switch those too
- Replace the entire
Runtimepointer on the agent struct - Clear the session (fresh context for the new persona)
Tool System (toolsrv/)
Runner Interface
The Runner is the minimal interface satisfied by any tool server (local or remote):
type Runner interface {
ListTools() ([]ToolInfo, error)
CallTool(ctx context.Context, tool string, args json.RawMessage) (json.RawMessage, error)
}
Server
The concrete Server struct implements Runner and provides:
- shell: sandboxed bash command execution
- tool registry:
tool_load,tool_list,tool_active— lazy tool discovery - skill registry:
skill_load,skill_list,skill_active— skill management
shell
Runs a single bash command in a sandbox.
Input schema:
{
"cmd": "...",
"timeout": 30,
"sandbox": "default",
"elevated": false
}
Code Validation
Before execution, inline code is checked against dangerous patterns:
- Universal:
mkfs,dd if=/dev/,sudo/su,/etc/shadow - Bash: recursive force-delete of system paths, fork bombs,
/dev/sdwrites, eval injection - Python:
shutil.rmtree("/"), subprocess rm -rf, os.remove of system paths - Perl: system/exec rm -rf, backtick rm, unlink system paths
- Lua: os.execute rm -rf, os.remove system paths
Validation failures are rate-limited: 5 failures within 1 minute triggers a 5-minute block.
Parallelism & Locking
Tool scripts declare their concurrency class via their .meta sidecar file:
.meta field |
Lock Class | Behavior |
|---|---|---|
"readOnly": true |
Read | Shared lock (LOCK_SH); concurrent with other reads |
"readOnly": false (default) |
Write | Exclusive lock (LOCK_EX); serialized |
Lock files live in /tmp/ollie/{sessionID}/ and are acquired via flock(2).
Result Tier Classification (toolsrv/tier.go)
Built-in cold tools (results consumed immediately, don't need verbatim retention):
file_read,memory_recall,web_fetch,web_searchlsp_definition,lsp_references,lsp_hover
Custom tools declare their tier via "tier": "cold|warm|hot" in their .meta file.
Elevation
Steps marked elevated: true bypass the sandbox entirely. They are dispatched to the elevation backend (x/elevate) which runs outside landrun. Only bash is supported for elevated steps.
Remote Execution
When a session is configured with remote=user@host, the runner is a RemoteServer that forwards tool calls over SSH to an ollie-remote binary:
type RemoteServer struct {
conn *rpc.Conn // JSON-RPC over SSH stdin/stdout
}
func (r *RemoteServer) CallTool(ctx context.Context, tool string, args json.RawMessage) (json.RawMessage, error) {
return r.conn.Call(ctx, "execute", tool, args)
}
The agent loop doesn't know or care whether execution is local or remote. The NewToolServer factory returns a RemoteServer instead of a local Server when the session is configured for remote execution.
SSH bootstrap:
- Open SSH connection to target host
- Send bootstrap script over stdin
- Bootstrap checks for cached
ollie-remotebinary (by SHA-256 hash) - If not cached: receive binary over stdin (gzipped + base64)
- Launch
ollie-remotewith the working directory - Switch to JSON-RPC over stdin/stdout
Backend System (backend/)
Shared Types
| Type | Purpose |
|---|---|
Message |
Conversation turn (role, content, tool_calls, tool_call_id) |
Tool |
Function declaration (name, description, JSON Schema params) |
ToolCall |
Model's request to invoke a function |
StreamEvent |
Incremental streaming delta (content, reasoning, tool_calls, done, usage) |
Usage |
Token counts + optional cost |
GenerationParams |
Unified sampling parameters |
Error Types
| Error | HTTP Code | Loop Behavior |
|---|---|---|
RateLimitError |
429 | Retry with backoff |
TransientError |
5xx, network | Retry with backoff |
ContextOverflowError |
400 (specific) | Compact and retry |
ToolUnsupportedError |
400/422 (specific) | Propagate to turnError hook |
streamRequest
Shared HTTP streaming infrastructure used by all backends:
- Execute the HTTP request
- Classify response status into typed errors
- Spawn a goroutine that parses SSE/NDJSON into
StreamEventchannel - Return the channel for the loop to consume
Backend Implementations
OllamaBackend (/api/chat):
- Local inference, no API key
- Context length from
/api/show(cached) - Model list from
/api/tags - NDJSON streaming
OpenAIBackend (/v1/chat/completions):
- Works with OpenAI, OpenRouter, any compatible API
- Configurable base URL, name, extra headers
- Context length from model list endpoint (cached)
- SSE streaming (
data: {...}lines) - Handles both
tool_callsand legacyfunction_call
AnthropicBackend (/v1/messages):
- Extended thinking (reasoning budget via
ThinkingBudget) - Prompt caching (cache_control blocks)
- 200k context window (hardcoded)
- SSE streaming with content_block_delta events
GeminiBackend (Gemini API):
- Google Gemini models
CopilotBackend (GitHub Copilot API):
- Copilot Chat integration
CodeWhispererBackend (AWS CodeWhisperer API):
- AWS CodeWhisperer integration
Config System (agent/)
Agent Config Schema
type Config struct {
Hooks map[string]HookCmds
Prompt Prompt
Backend string
Model string
Tools *bool
AllowExecutors []string
AllowTools []string
MaxTokens int
MaxCompletionTokens int
MaxSteps int
Temperature *float64
TopP, TopK, MinP, TopA *float64/*int
FrequencyPenalty *float64
PresencePenalty *float64
RepetitionPenalty *float64
Reasoning int
ReasoningEffort string
IncludeReasoning *bool
ResponseFormat string
Stop []string
Verbosity string
}
Prompt Type
Dual-mode unmarshaling:
- JSON string: treated as literal text (with
!command, file path, or env expansion) - JSON array: each element is a shell command; stdout concatenated
The array form is preferred — it enables composable prompt assembly from template fragments.
HookCmds Type
Dual-mode: a single string "cmd" or an array ["cmd1", "cmd2"]. Multiple commands run sequentially.
Prompt Resolution (agent/prompt_resolver.go)
Two paths based on Prompt.IsExec:
Exec mode (array):
- Each command runs via
sh -cwith 30s timeout - CWD set to session working directory
- Extra env vars injected (OLLIE_SESSION_ID, OLLIE mount path)
- Non-zero exit skips that command (no error propagation)
- All stdout concatenated with newlines
String mode (legacy):
- Environment variables expanded
- If contains newline: literal text
- If starts with
!: execute as shell command - If names an existing file: read file contents
- Otherwise: use as-is
Cost Tracking (agent/cost.go)
When the backend doesn't report cost directly (only OpenRouter does), cost is computed from a built-in pricing table:
- Matches model name prefix (case-insensitive)
- Covers Anthropic Claude, OpenAI GPT/o-series, Google Gemini
- Cached input tokens charged at discount (10% Anthropic, 50% others)
- Cache creation tokens charged at 125% of input rate
- Unknown/local models return zero cost
Sandbox (sandbox/)
Config Format
general:
best_effort: true
log_level: ""
filesystem:
ro: ["/usr", "/lib", "/etc"]
rox: ["/usr/bin", "/usr/local/bin"]
rw: ["{CWD}", "{TMPDIR}"]
rwx: []
network:
enabled: true
unrestricted: true
bind_tcp: []
connect_tcp: []
env: []
advanced:
ldd: false
add_exec: false
Template Variables
| Variable | Expansion |
|---|---|
{CWD} |
Session working directory |
{HOME} |
User home directory |
{TMPDIR} |
Temp directory (default: /tmp) |
{XDG_CONFIG_HOME} |
XDG config (default: ~/.config) |
{XDG_DATA_HOME} |
XDG data (default: ~/.local/share) |
{XDG_CACHE_HOME} |
XDG cache (default: ~/.cache) |
{XDG_STATE_HOME} |
XDG state (default: ~/.local/state) |
{XDG_RUNTIME_DIR} |
XDG runtime (default: /run/user/UID) |
{OLLIE_CFG_PATH} |
ollie config directory |
{OLLIE_DATA_PATH} |
ollie data directory |
Event System
The agent communicates exclusively through events published on a pubsub.Bus:
type Event struct {
Role string
Name string
Content string
}
| Role | When |
|---|---|
assistant |
Streaming text delta from the model |
reasoning |
Thinking/reasoning delta |
call |
Tool call initiated (Name=tool, Content=args JSON) |
tool |
Tool result returned (Name=tool, Content=result) |
info |
Informational message (command output, status) |
exec |
Shell command output (!cmd) |
error |
Error message |
usage |
Token usage report |
limitretry |
Rate limit retry initiated |
Consumers subscribe to the "event" topic on the bus. The 9P server's session manager subscribes to build the chat log; frontends subscribe to render output.
Concurrency Model
- One turn at a time:
submitMuserializesSubmit()calls - Interrupt:
currentActionholds aCancelCauseFunc;Interrupt()calls it - Queue drain: after each turn completes, queued prompts are drained sequentially
- State observation:
WaitChange()uses a condition variable (changeCond) that broadcasts on any state/usage/cwd change - Tool parallelism: within a turn, read-safe tools fan out via goroutines with
sync.Mapcaching andflockcoordination
Session ID Format
<unix-nanoseconds>-<6-hex-chars>
Lexicographically sortable by creation time. Unique even within the same nanosecond due to the random suffix.