ollie/doc/architecture-embedding.md

6.1 KiB

Embedding Architecture: Tool and Skill Discovery

Ollie uses a local text-embedding model to select relevant skills and tool hints before assembling each model prompt. This is a discovery and progressive disclosure subsystem. It does not execute tools, replace the tool registry, or act as the conversational backend.

Purpose

Exact-name discovery is fragile for small models. A user may ask to “inspect source structure” without naming code_outline, or ask for domain guidance without knowing the skill name. Embedding descriptions makes the capability surface searchable by meaning while keeping the prompt bounded.

The mechanism is deliberately local. User requests, tool descriptions, and skill descriptions are embedded by an ONNX sentence-transformer on the Ollie host. No discovery request is sent to the configured LLM provider.

Components

flowchart TB
    Q[User request] --> A[Agent turn]
    A --> SI[Skill index]
    A --> TI[Tool index]
    SI --> E[embedding.Model]
    TI --> E
    E --> ONNX[all-MiniLM-L6-v2 ONNX model]
    SI --> SH[Relevant skill content]
    TI --> TH[Relevant tool hints]
    SH --> P[Prompt assembly]
    TH --> P
    P --> L[Conversational backend]

embedding package

embedding.Model loads three runtime assets from the model directory:

  • model.onnx — the all-MiniLM-L6-v2 sentence-transformer.
  • tokenizer.json — tokenizer vocabulary and configuration.
  • libonnxruntime.so — ONNX Runtime shared library.

The model produces 384-dimensional vectors. Embed tokenizes text, runs the ONNX session, performs attention-masked mean pooling, and L2-normalizes the result. CosineSimilarity compares query and description vectors.

ONNX Runtime initialization is process-global and occurs once. Model sessions are mutex-protected because inference uses shared session state. Each index owns its model session and releases it through Close.

embedding.Index[T]

The generic index type handles arbitrary entry types:

Index[T] {
    model: *Model
    entries: []T
    vectors: []Vector
    textFn:  func(T) string  // extracts text for embedding
}

Skill and tool indexes use the same generic structure with different entry types. For skills, entries are skills.Skill structs with name, description, and content. For tools, entries are toolsrv.Meta structs from loaded tool metadata.

Description vectors are precomputed when an index is built or reloaded. A user request is embedded at match time, then compared against all indexed vectors. Results below the threshold are discarded; remaining results are sorted by cosine score and capped by the caller.

Agent integration

The agent keeps independent cached indexes for skills and tools. Skill index initialization is lazy and guarded by sync.Once. Tool index is rebuilt per turn from the agent's currently loaded tools:

first turn
  └── build skill index from configured SKILL.md directories

each turn
  ├── build tool index from agent's loaded tools
  ├── embed request
  ├── match skills: threshold 0.35, limit 3
  └── match tools:  threshold 0.35, limit 5

Relevant skills are injected as a <skills> block containing complete skill content. Relevant tools are injected as a <tool-hints> block containing loaded tool names and descriptions. Tools must be explicitly loaded through the agent's autoLoad config or via the ctl file before they appear in matching results.

The bounded result limits prevent discovery from becoming a prompt-size or context-budget problem. This progressive disclosure is especially useful for models at or below 8B parameters: the model receives a small, request-specific capability set instead of a large undifferentiated catalog.

Configuration and installation

The default model directory is:

$XDG_DATA_HOME/ollie/models

or ~/.local/share/ollie/models when XDG_DATA_HOME is unset. make install-models downloads the model, tokenizer, and ONNX Runtime. The default skill directory is:

$XDG_CONFIG_HOME/ollie/skills

or ~/.config/ollie/skills. Skill search paths can be overridden with one path per line in embedding.conf in the Ollie configuration directory. Blank lines and comments are ignored; environment variables and ~ are expanded.

Failure behavior and lifecycle

Embedding is an enhancement, not a hard dependency for the agent loop. If the model cannot load, a skill directory is missing, a metadata file is malformed, or an individual embedding fails, semantic injection is skipped or that entry is omitted. The request continues through ordinary prompt assembly and tool loading.

Indexes are process-local. They are not persisted as vector databases. A new process rebuilds vectors from current skill files and tool metadata. The existing lazy index cache means repeated turns reuse loaded models and vectors. Changes to skills or tool metadata require index reload or a new process before they affect ranking.

Design constraints

  • Keep the conversational backend independent from discovery.
  • Keep toolsrv authoritative for metadata, loading, execution, sandboxing, and process state.
  • Keep descriptions short and meaningful because they are the indexed corpus.
  • Keep ranking bounded; semantic discovery must not inflate every prompt.
  • Treat embedding failure as a capability-discovery miss, not a turn failure.
  • Keep the model local to preserve privacy, deterministic deployment, and operation with small or offline conversational models.

Source map

  • embedding/embedding.go — model loading, tokenization, inference, pooling, normalization, and similarity.
  • embedding/index.go — generic Index[T] for vector-indexed entries.
  • skills/skills.go — skill scanning, frontmatter parsing, and Skill type.
  • cmd/olliesrv/internal/agent/skill_match.go — lazy skill index and prompt injection for skills.
  • cmd/olliesrv/internal/agent/tool_match.go — per-turn tool index from loaded tools and prompt injection for tool hints.
  • Makefile — model and runtime asset installation.
  • architecture-tools.md — tool metadata and execution.