ollie/doc/resources/core.md

24 KiB

ollie-core: Detailed Architecture

This document covers the internal design of the Go library — the packages that implement the agent engine. It is the dependency imported by all consumers (olliesrv, frontends, tests).

Package Map

flowchart TB
    subgraph Agent["agent/ package"]
        AGENT["agent.go\nAgent struct, New()"]
        LOOP["loop.go\nrun() — agent loop"]
        HISTORY["history.go\nHistory, TaskState, compaction"]
        TURN["turn.go\nper-turn state machine"]
        HOOKS["hooks.go\nlifecycle callbacks"]
        CMDS["commands.go\n/command dispatch"]
        BUILD["build_runtime.go\nBuildRuntime()"]
        PROMPT["prompt_resolver.go\nsystem prompt assembly"]
        STATE["state.go\nResultTier, toolResult"]
        COST["cost.go\npricing table"]
        FIFO["fifo.go\nprompt queue"]
        COMPLETE["complete.go\ncode completion"]
        CONFIG["agent_config.go\nConfig schema"]
        CONFIG_PATHS["config_paths.go\nXDG config discovery"]
        RUNTIME["runtime.go\nRuntime struct"]
        NEW["new.go\nNew() factory"]
        USAGE["usage_log.go\nusage tracking"]
    end

    subgraph Backend["backend/ package"]
        BACKEND["backend.go\nBackend interface"]
        OPENAI["openai.go"]
        ANTHROPIC["anthropic.go"]
        OLLAMA["ollama.go"]
        GEMINI["gemini.go"]
        COPILOT["copilot.go"]
        CW["codewhisperer.go"]
        POOL["pool.go\nbackend pool"]
    end

    subgraph Toolsrv["toolsrv/ package"]
        SERVER["server.go\nServer struct"]
        SHELL["shell.go\nshell execution"]
        TOOLS["tools.go\nRunner interface"]
        REGISTRY["registry.go\ntool discovery"]
        DISCOVER["discover.go\ntool script scanning"]
        SCHEMA["schema.go\nJSON Schema parsing"]
        TIER["tier.go\nresult tiering"]
        SKILLS["skills.go\nskill management"]
        STREAM["stream.go\noutput streaming"]
        REMOTE["remote.go\nRemoteServer"]
        DIAL["dial.go\nSSH bootstrap"]
        ACCESSORS["accessors.go\nIsParallelRead, ResultTier"]
        BOOTSTRAP["bootstrap.sh\nSSH bootstrap script"]
    end

    subgraph Session["session/ package"]
        SESS["session.go\nSession, Config"]
    end

    subgraph Fs["fs/ package"]
        TREE["tree.go\n*Tree"]
        NEWROOT["newroot.go\nNewRoot()"]
        LIFECYCLE["lifecycle.go\nCreate/Kill/Rename"]
        SESSFILES["sessionfiles.go\nsession file handlers"]
        AGENTFILES["agentfiles.go\nagent file handlers"]
        ROOTFILES["rootfiles.go\nroot file handlers"]
        ELEVATEFILES["elevatefiles.go\nelevation handlers"]
        PROCFILES["procfiles.go\nprocess handlers"]
        PERSIST["persist.go\nsession persistence"]
        FS_GO["fs.go\n9P File/FileConfig"]
        FORMAT["format.go\nevent formatting"]
        TYPES["types.go\nSession struct"]
        SPEC["spec.go\nEDSL namespace declaration"]
        BUILDER["builder.go\nBuildTree"]
        FSNODE["fsnode.go\nFsNodeDecl"]
    end

    subgraph Support["Supporting Packages"]
        DETACH["detach/\nprocess management"]
        ELEVATE["elevate/\nelevation broker"]
        SANDBOX["sandbox/\nLandlock config"]
        ENV["env/\nenvironment"]
        LOG["log/\nstructured logging"]
        PATHS["paths/\nXDG resolution"]
    end

    Agent --> Backend
    Agent --> Toolsrv
    Agent --> Session
    Fs --> Session
    Toolsrv --> SANDBOX
    Toolsrv --> DETACH
    Toolsrv --> ELEVATE

The Agent Struct (agent/agent.go)

Agent is the boundary between the engine and any consumer. It is a concrete struct, not an interface — the interface was removed during the Great Flattening because no consumer ever implemented it independently.

type Agent struct {
    history        *History
    runtime        *Runtime
    cfg            agentConfig
    profileName    string
    agentID        string
    displayName    string
    agentsDir      string
    baseLayers     []string
    promptEnvExtra []string
    newToolServer  func() toolsrv.Runner
    newBackend     func(string) (backend.Backend, error)
    currentAction  atomic.Pointer[actionHandle]
    // ... execution state, sync primitives, session metadata
}

func New(cfg AgentConfig) *Agent

All output flows through the event bus as Event{Role, Name, Content} values. Roles include: assistant, tool, call, info, exec, error, usage, reasoning, limitretry.

Agent Construction (agent/new.go)

AgentConfig

The full construction config:

Field Purpose
History Pre-existing session history to restore (nil = new)
Runtime Pre-built *Runtime (tools, hooks, prompt, params)
Profile Name of the active agent configuration profile
ID Immutable agent ID (uname)
Name Mutable display name
AgentsDir Path to agent JSON directory
SessionsDir Path to persisted session files
SessionID Unique session identifier
AgentID Immutable agent principal
CWD Working directory for tools and prompt
NewToolServer Factory for fresh toolsrv.Runner on agent switch
NewBackend Factory for backend by name
Log Logger instance
MaxSteps Override agent config's maxSteps
ReadPlanStep Function to read next unchecked plan step
ListHandlers Additional list handlers

Runtime

The Runtime struct holds all per-agent configuration that changes on an /agent switch but remains stable across turns. The agent struct stores a pointer to the active Runtime; switching agents replaces it atomically.

type Runtime struct {
    Backend      backend.Backend
    Runner       toolsrv.Runner
    Hooks        Hooks
    Preamble     string              // compiled system prompt
    Tools        []backend.Tool
    ClassifyTool func(string) bool   // parallelism check
    ClassifyTier func(string, json.RawMessage) string
    GenParams    backend.GenerationParams
    MaxSteps           int
    ToolResultMaxBytes int
    CfgBackend   string
    CfgModel     string
    Messages     []string            // non-fatal warnings
}

BuildRuntime

BuildRuntime(cfg, runner, cwd, env) wires everything together:

  1. Lists all tools from the runner
  2. Resolves the system prompt by executing the config's prompt commands
  3. Extracts generation parameters from the config
  4. Builds the tool executor closure (routes calls through the runner)
  5. Extracts parallelism and tier classifiers
  6. Applies tool/executor restrictions if configured
  7. Returns a *Runtime struct ready for New

New

Creates the Agent struct from AgentConfig:

  1. Sweeps stale tmpdirs from previous crashes
  2. Stores the provided Runtime (or creates an empty one)
  3. Sets the backend model if ModelName is specified
  4. Creates /tmp/ollie/{sessionID} for this session
  5. Sets up the plan file reader
  6. Wires the turnError hook with a detached context
  7. Pushes OLLIE_SESSION_ID and OLLIE_AGENT_ID into the runner env
  8. Sets the flock directory for parallel coordination

The Agent Loop (agent/loop.go)

The run() function is the core execution engine. It takes an agentConfig and a state interface, and loops until the model stops calling tools.

Per-Iteration Flow

flowchart TB
    START["Start turn"]
    GATE["Context gate:\nstrip cold material\nif over budget"]
    BUILD["Build message history:\nsystem prompt + task state + hot tail"]
    STREAM["Stream from backend\n(with retry logic)"]
    CHECK{"Tool calls\nin response?"}
    EXEC["Execute tool calls\n(serial or parallel)"]
    UPDATE["Update state:\nappend messages + results"]
    INFER["Infer TaskState\nupdates"]
    SAFETY["Safety checks:\nerror limits, stall,\nreplan, max steps"]
    COMPACT["Auto-compact\nif needed"]
    DONE["Turn complete"]

    START --> GATE --> BUILD --> STREAM --> CHECK
    CHECK -- yes --> EXEC --> UPDATE --> INFER --> SAFETY --> COMPACT --> BUILD
    CHECK -- no --> DONE

Streaming & Retry

The streaming loop handles three failure modes:

Failure Detection Recovery
Rate limit (429) *RateLimitError Exponential backoff (5x2^attempt seconds), up to 3 retries
Transient (5xx, network) *TransientError Same retry logic
Stream drop Done never received, channel closes Retry with 2x2^attempt second delay
Context overflow *ContextOverflowError Trigger compaction, retry
Tool unsupported *ToolUnsupportedError Propagate to turnError hook

On the first error, the turnError hook fires. If it exits 0 (handled), the loop returns immediately — the hook is responsible for recovery (e.g., switching models via freeloader).

Tool Execution

Tool calls are dispatched with parallelism awareness:

  1. Classify each tool call via ClassifyTool(name) (reads readOnly from .meta file)
  2. Batch consecutive parallel-read-safe calls into a group
  3. Fan out the batch concurrently with deduplication (identical calls share one execution)
  4. Serial calls run one at a time

Per-call lifecycle:

preTool hook → execute → postTool hook → inject check → truncation → cache (if read-safe)
  • preTool exit 2 blocks execution
  • postTool exit 2 replaces the result
  • Results exceeding ToolResultMaxBytes are hard-truncated
  • Read-safe results are cached in a sync.Map for the duration of the turn

Safety Mechanisms

Mechanism Threshold Action
Consecutive errors (soft) 5 rounds Inject nudge: "try a different approach"
Consecutive errors (hard) 10 rounds Abort the turn
Replan gate 8 rounds without PLAN: block Inject replan nudge
Stall detection 5 rounds with same LastAction Inject stall nudge
Max steps Configurable per-agent Soft exit with nudge
Plan re-injection Every 10 tool rounds Re-inject task state

TaskState Inference

After each tool round, inferTaskStateUpdate updates the structured overlay:

  • LastAction ← most recent tool call + truncated result
  • PlanStep ← next unchecked item from the plan file (or heuristic from assistant text)

Session & Context (agent/history.go)

Message History

History holds a flat []backend.Message slice. No rolling window — grows until compaction.

Checkpoint: History.Checkpoint(ts TaskState) forks the session — returns a new History that inherits the given TaskState but starts with a clean message history. Used for narrow-context sub-agents that know what to do without inheriting parent message noise.

Token tracking per session:

  • TotalInputTokens, TotalOutputTokens, TotalRequests
  • TotalCachedInputTokens, TotalCacheCreationTokens
  • Per-turn accumulators reset at turn start
  • LastTurnCostUSD, SessionCostUSD

Compaction

Triggered when estimated tokens reach 50% of model context length.

Three-zone strategy:

Zone Content Size
Cold Structured TaskState JSON 1 message
Warm One-line decision index ~10 messages
Hot Verbatim recent messages 8 messages

Process:

  1. Flatten tool call/result pairs to plain text
  2. Send full history + compaction prompt to the LLM
  3. Parse structured JSON response into TaskState
  4. Rebuild history as cold + warm + hot zones
  5. Write pre-compaction snapshot to {id}.compaction.jsonl for recovery

Result Tiers

Tool results are classified at execution time:

Tier Behavior Examples
Hot Stays verbatim in messages file_write, file_edit, reasoning_think
Warm Summarized on next compaction (custom tools)
Cold Immediately collapsed to one-line summary file_read, memory_recall, lsp_*

Classification sources (in priority order):

  1. Built-in table in toolsrv/tier.go (defaultColdTools)
  2. tier field in the tool's .meta sidecar file
  3. Default: hot

Persistence

Sessions are saved as PersistedSession JSON:

{
  "id": "1716300000000-abc123",
  "agent": "coding",
  "messages": [...],
  "taskState": {...}
}

Saved after every tool round (saveSession callback) and on clean exit.

Hooks (agent/hooks.go)

Hooks are shell commands executed with:

  • stdin: JSON payload with context (session_id, cwd, model, error, etc.)
  • env: OLLIE_* variables derived from the payload map
  • timeout: 60 seconds (configurable in tests)
  • process group: Setpgid=true for clean kill on timeout

Exit Code Semantics

Code Meaning
0 Success. Stdout appended to context.
2 Block. Stops the action (e.g., blocks a tool call, prevents compaction).
Other Non-blocking warning. Execution continues.

Hook Execution Order

Multiple commands per hook run sequentially. If any exits 2, execution stops immediately and returns Blocked=true.

Commands (agent/commands.go)

Submit() first checks for commands before starting an agent turn:

Prefix Dispatch
!<cmd> Shell command in session CWD
/backend [name] Show or switch LLM backend
/model [name] Show or switch model
/models List available models
/agent [name] Show or switch agent config (rebuilds Runtime)
/agents List available agents
/maxsteps [n] Get/set max steps (0 = unlimited)
/compact Manual compaction
/context Show context size and message breakdown
/usage Show token usage and context percentage
/cost Show last turn and session cost
/history Dump bounded message history
/sessions List saved sessions
/clear Clear session (reset context)
/kill Kill session
/rn <name> Rename session
/cwd [path] Show or change working directory
/skills List available skills
/tools List available tools
/sp Show rendered system prompt
/i <prompt> Inject into running turn
/irw <prompt> Inject-rewrite
`/queued [pop clear]`
/help Show command help

Agent Switching

When /agent <name> is issued:

  1. Load the new agent config JSON
  2. Create a fresh runner via newToolServer()
  3. Call BuildRuntime with the new config
  4. If the config specifies a backend/model, switch those too
  5. Replace the entire Runtime pointer on the agent struct
  6. Clear the session (fresh context for the new persona)

Tool System (toolsrv/)

Runner Interface

The Runner is the minimal interface satisfied by any tool server (local or remote):

type Runner interface {
    ListTools() ([]ToolInfo, error)
    CallTool(ctx context.Context, tool string, args json.RawMessage) (json.RawMessage, error)
}

Server

The concrete Server struct implements Runner and provides:

  • shell: sandboxed bash command execution
  • tool registry: tool_load, tool_list — lazy tool discovery
  • skill registry: skill_load, skill_list — skill management

shell

Runs a single bash command in a sandbox.

Input schema:

{
  "cmd": "...",
  "timeout": 30,
  "sandbox": "default",
  "elevated": false
}

Code Validation

Before execution, inline code is checked against dangerous patterns:

  • Universal: mkfs, dd if=/dev/, sudo/su, /etc/shadow
  • Bash: recursive force-delete of system paths, fork bombs, /dev/sd writes, eval injection
  • Python: shutil.rmtree("/"), subprocess rm -rf, os.remove of system paths
  • Perl: system/exec rm -rf, backtick rm, unlink system paths
  • Lua: os.execute rm -rf, os.remove system paths

Validation failures are rate-limited: 5 failures within 1 minute triggers a 5-minute block.

Parallelism & Locking

Tool scripts declare their concurrency class via their .meta sidecar file:

.meta field Lock Class Behavior
"readOnly": true Read Shared lock (LOCK_SH); concurrent with other reads
"readOnly": false (default) Write Exclusive lock (LOCK_EX); serialized

Lock files live in /tmp/ollie/{sessionID}/ and are acquired via flock(2).

Result Tier Classification (toolsrv/tier.go)

Built-in cold tools (results consumed immediately, don't need verbatim retention):

  • file_read, memory_recall, web_fetch, web_search
  • lsp_definition, lsp_references, lsp_hover

Custom tools declare their tier via "tier": "cold|warm|hot" in their .meta file.

Elevation

Steps marked elevated: true bypass the sandbox entirely. They are dispatched to the elevation backend (x/elevate) which runs outside landrun. Only bash is supported for elevated steps.

Remote Execution

When a session is configured with remote=user@host, the runner is a RemoteServer that forwards tool calls over SSH to an ollie-remote binary:

type RemoteServer struct {
    conn *rpc.Conn  // JSON-RPC over SSH stdin/stdout
}

func (r *RemoteServer) CallTool(ctx context.Context, tool string, args json.RawMessage) (json.RawMessage, error) {
    return r.conn.Call(ctx, "execute", tool, args)
}

The agent loop doesn't know or care whether execution is local or remote. The NewToolServer factory returns a RemoteServer instead of a local Server when the session is configured for remote execution.

SSH bootstrap:

  1. Open SSH connection to target host
  2. Send bootstrap script over stdin
  3. Bootstrap checks for cached ollie-remote binary (by SHA-256 hash)
  4. If not cached: receive binary over stdin (gzipped + base64)
  5. Launch ollie-remote with the working directory
  6. Switch to JSON-RPC over stdin/stdout

Backend System (backend/)

Shared Types

Type Purpose
Message Conversation turn (role, content, tool_calls, tool_call_id)
Tool Function declaration (name, description, JSON Schema params)
ToolCall Model's request to invoke a function
StreamEvent Incremental streaming delta (content, reasoning, tool_calls, done, usage)
Usage Token counts + optional cost
GenerationParams Unified sampling parameters

Error Types

Error HTTP Code Loop Behavior
RateLimitError 429 Retry with backoff
TransientError 5xx, network Retry with backoff
ContextOverflowError 400 (specific) Compact and retry
ToolUnsupportedError 400/422 (specific) Propagate to turnError hook

streamRequest

Shared HTTP streaming infrastructure used by all backends:

  1. Execute the HTTP request
  2. Classify response status into typed errors
  3. Spawn a goroutine that parses SSE/NDJSON into StreamEvent channel
  4. Return the channel for the loop to consume

Backend Implementations

OllamaBackend (/api/chat):

  • Local inference, no API key
  • Context length from /api/show (cached)
  • Model list from /api/tags
  • NDJSON streaming

OpenAIBackend (/v1/chat/completions):

  • Works with OpenAI, OpenRouter, any compatible API
  • Configurable base URL, name, extra headers
  • Context length from model list endpoint (cached)
  • SSE streaming (data: {...} lines)
  • Handles both tool_calls and legacy function_call

AnthropicBackend (/v1/messages):

  • Extended thinking (reasoning budget via ThinkingBudget)
  • Prompt caching (cache_control blocks)
  • 200k context window (hardcoded)
  • SSE streaming with content_block_delta events

GeminiBackend (Gemini API):

  • Google Gemini models

CopilotBackend (GitHub Copilot API):

  • Copilot Chat integration

CodeWhispererBackend (AWS CodeWhisperer API):

  • AWS CodeWhisperer integration

Config System (agent/)

Agent Config Schema

type Config struct {
    Hooks               map[string]HookCmds
    Prompt              Prompt
    Backend             string
    Model               string
    Tools               *bool
    AllowExecutors      []string
    AllowTools          []string
    MaxTokens           int
    MaxCompletionTokens int
    MaxSteps            int
    Temperature         *float64
    TopP, TopK, MinP, TopA  *float64/*int
    FrequencyPenalty    *float64
    PresencePenalty     *float64
    RepetitionPenalty   *float64
    Reasoning           int
    ReasoningEffort     string
    IncludeReasoning    *bool
    ResponseFormat      string
    Stop                []string
    Verbosity           string
}

Prompt Type

Dual-mode unmarshaling:

  • JSON string: treated as literal text (with !command, file path, or env expansion)
  • JSON array: each element is a shell command; stdout concatenated

The array form is preferred — it enables composable prompt assembly from template fragments.

HookCmds Type

Dual-mode: a single string "cmd" or an array ["cmd1", "cmd2"]. Multiple commands run sequentially.

Prompt Resolution (agent/prompt_resolver.go)

Two paths based on Prompt.IsExec:

Exec mode (array):

  • Each command runs via sh -c with 30s timeout
  • CWD set to session working directory
  • Extra env vars injected (OLLIE_SESSION_ID, OLLIE mount path)
  • Non-zero exit skips that command (no error propagation)
  • All stdout concatenated with newlines

String mode (legacy):

  • Environment variables expanded
  • If contains newline: literal text
  • If starts with !: execute as shell command
  • If names an existing file: read file contents
  • Otherwise: use as-is

Cost Tracking (agent/cost.go)

When the backend doesn't report cost directly (only OpenRouter does), cost is computed from a built-in pricing table:

  • Matches model name prefix (case-insensitive)
  • Covers Anthropic Claude, OpenAI GPT/o-series, Google Gemini
  • Cached input tokens charged at discount (10% Anthropic, 50% others)
  • Cache creation tokens charged at 125% of input rate
  • Unknown/local models return zero cost

Sandbox (sandbox/)

Config Format

general:
  best_effort: true
  log_level: ""

filesystem:
  ro:  ["/usr", "/lib", "/etc"]
  rox: ["/usr/bin", "/usr/local/bin"]
  rw:  ["{CWD}", "{TMPDIR}"]
  rwx: []

network:
  enabled: true
  unrestricted: true
  bind_tcp: []
  connect_tcp: []

env: []

advanced:
  ldd: false
  add_exec: false

Template Variables

Variable Expansion
{CWD} Session working directory
{HOME} User home directory
{TMPDIR} Temp directory (default: /tmp)
{XDG_CONFIG_HOME} XDG config (default: ~/.config)
{XDG_DATA_HOME} XDG data (default: ~/.local/share)
{XDG_CACHE_HOME} XDG cache (default: ~/.cache)
{XDG_STATE_HOME} XDG state (default: ~/.local/state)
{XDG_RUNTIME_DIR} XDG runtime (default: /run/user/UID)
{OLLIE_CFG_PATH} ollie config directory
{OLLIE_DATA_PATH} ollie data directory

Event System

The agent communicates exclusively through events published on a pubsub.Bus:

type Event struct {
    Role    string
    Name    string
    Content string
}
Role When
assistant Streaming text delta from the model
reasoning Thinking/reasoning delta
call Tool call initiated (Name=tool, Content=args JSON)
tool Tool result returned (Name=tool, Content=result)
info Informational message (command output, status)
exec Shell command output (!cmd)
error Error message
usage Token usage report
limitretry Rate limit retry initiated

Consumers subscribe to the "event" topic on the bus. The 9P server's session manager subscribes to build the chat log; frontends subscribe to render output.

Concurrency Model

  • One turn at a time: submitMu serializes Submit() calls
  • Interrupt: currentAction holds a CancelCauseFunc; Interrupt() calls it
  • Queue drain: after each turn completes, queued prompts are drained sequentially
  • State observation: WaitChange() uses a condition variable (changeCond) that broadcasts on any state/usage/cwd change
  • Tool parallelism: within a turn, read-safe tools fan out via goroutines with sync.Map caching and flock coordination

Session ID Format

<unix-nanoseconds>-<6-hex-chars>

Lexicographically sortable by creation time. Unique even within the same nanosecond due to the random suffix.