From ba91920a88034b37ed194829210204632fad73c1 Mon Sep 17 00:00:00 2001 From: Ollie Agent Date: Thu, 20 Aug 2026 17:41:56 +0200 Subject: [PATCH] Document embedding discovery architecture --- doc/architecture-embedding.md | 156 ++++++++++++++++++++++++++++++++++ doc/architecture-tools.md | 23 +---- 2 files changed, 159 insertions(+), 20 deletions(-) create mode 100644 doc/architecture-embedding.md diff --git a/doc/architecture-embedding.md b/doc/architecture-embedding.md new file mode 100644 index 0000000..871c3f3 --- /dev/null +++ b/doc/architecture-embedding.md @@ -0,0 +1,156 @@ +# Embedding Architecture: Tool and Skill Discovery + +Ollie uses a local text-embedding model to select relevant skills and tool +hints before assembling each model prompt. This is a discovery and progressive +disclosure subsystem. It does not execute tools, replace the tool registry, or +act as the conversational backend. + +## Purpose + +Exact-name discovery is fragile for small models. A user may ask to “inspect +source structure” without naming `code_outline`, or ask for domain guidance +without knowing the skill name. Embedding descriptions makes the capability +surface searchable by meaning while keeping the prompt bounded. + +The mechanism is deliberately local. User requests, tool descriptions, and +skill descriptions are embedded by an ONNX sentence-transformer on the Ollie +host. No discovery request is sent to the configured LLM provider. + +## Components + +```mermaid +flowchart TB + Q[User request] --> A[Agent turn] + A --> SI[Skill index] + A --> TI[Tool index] + SI --> E[embedding.Model] + TI --> E + E --> ONNX[all-MiniLM-L6-v2 ONNX model] + SI --> SH[Relevant skill content] + TI --> TH[Relevant tool hints] + SH --> P[Prompt assembly] + TH --> P + P --> L[Conversational backend] +``` + +### `embedding` package + +`embedding.Model` loads three runtime assets from the model directory: + +- `model.onnx` — the `all-MiniLM-L6-v2` sentence-transformer. +- `tokenizer.json` — tokenizer vocabulary and configuration. +- `libonnxruntime.so` — ONNX Runtime shared library. + +The model produces 384-dimensional vectors. `Embed` tokenizes text, runs the +ONNX session, performs attention-masked mean pooling, and L2-normalizes the +result. `CosineSimilarity` compares query and description vectors. + +ONNX Runtime initialization is process-global and occurs once. Model sessions +are mutex-protected because inference uses shared session state. Each index +owns its model session and releases it through `Close`. + +### `skills.Index` + +The shared index type handles both skills and tools: + +```text +Index { + model: embedding.Model + entries: []Skill + vectors: []Vector +} +``` + +For skills, the index scans configured directories for `/SKILL.md`, +parses `name` and `description` from YAML frontmatter, and retains the full +content for injection. For tools, the agent converts `.meta` descriptions into +the same `Skill` representation; only the name and description are needed for +ranking. + +Description vectors are precomputed when an index is built or reloaded. A user +request is embedded at match time, then compared against all indexed vectors. +Results below the threshold are discarded; remaining results are sorted by +cosine score and capped by the caller. + +## Agent integration + +The agent keeps independent cached indexes for skills and tools. Initialization +is lazy and guarded by `sync.Once`: + +```text +first turn + ├── build skill index from configured SKILL.md directories + └── build tool index from installed .meta files + +each turn + ├── embed request + ├── match skills: threshold 0.35, limit 3 + └── match tools: threshold 0.35, limit 5 +``` + +Relevant skills are injected as a `` block containing complete skill +content. Relevant tools are injected as a `` block containing +callable names and descriptions. The normal tool registry still controls +whether a tool is loaded and executable; semantic matching only improves what +the model sees first. + +The bounded result limits prevent discovery from becoming a prompt-size or +context-budget problem. This progressive disclosure is especially useful for +models at or below 8B parameters: the model receives a small, request-specific +capability set instead of a large undifferentiated catalog. + +## Configuration and installation + +The default model directory is: + +```text +$XDG_DATA_HOME/ollie/models +``` + +or `~/.local/share/ollie/models` when `XDG_DATA_HOME` is unset. `make +install-models` downloads the model, tokenizer, and ONNX Runtime. The default +skill directory is: + +```text +$XDG_CONFIG_HOME/ollie/skills +``` + +or `~/.config/ollie/skills`. Skill search paths can be overridden with one path +per line in `embedding.conf` in the Ollie configuration directory. Blank lines +and comments are ignored; environment variables and `~` are expanded. + +## Failure behavior and lifecycle + +Embedding is an enhancement, not a hard dependency for the agent loop. If the +model cannot load, a skill directory is missing, a metadata file is malformed, +or an individual embedding fails, semantic injection is skipped or that entry +is omitted. The request continues through ordinary prompt assembly and tool +loading. + +Indexes are process-local. They are not persisted as vector databases. A new +process rebuilds vectors from current skill files and tool metadata. The +existing lazy index cache means repeated turns reuse loaded models and vectors. +Changes to skills or tool metadata require index reload or a new process before +they affect ranking. + +## Design constraints + +- Keep the conversational backend independent from discovery. +- Keep `toolsrv` authoritative for metadata, loading, execution, sandboxing, + and process state. +- Keep descriptions short and meaningful because they are the indexed corpus. +- Keep ranking bounded; semantic discovery must not inflate every prompt. +- Treat embedding failure as a capability-discovery miss, not a turn failure. +- Keep the model local to preserve privacy, deterministic deployment, and + operation with small or offline conversational models. + +## Source map + +- `embedding/embedding.go` — model loading, tokenization, inference, pooling, + normalization, and similarity. +- `skills/skills.go` — skill scanning, frontmatter parsing, indexing, matching, + and reload. +- `cmd/olliesrv/internal/agent/skill_match.go` — lazy indexes and prompt + injection for skills and tool hints. +- `Makefile` — model and runtime asset installation. +- [`architecture-tools.md`](architecture-tools.md) — tool metadata and execution. diff --git a/doc/architecture-tools.md b/doc/architecture-tools.md index 2691390..741d52a 100644 --- a/doc/architecture-tools.md +++ b/doc/architecture-tools.md @@ -72,26 +72,9 @@ The `.meta` file is the sole source of discovery and schema. The registry never Tool discovery remains metadata-driven: `toolsrv` scans `.meta` files and manages which tools are loaded for an agent. Ollie adds a separate semantic ranking -layer before prompt assembly: - -```text -user request - │ - ├── embedding.Model.Embed() - │ - ├── cosine similarity → skill descriptions → up to 3 injected skills - └── cosine similarity → tool descriptions → up to 5 injected hints -``` - -The shared `skills.Index` precomputes vectors for skill or tool descriptions and -returns results above a configured threshold in descending relevance order. The -agent keeps the full tool registry and callable tool loading semantics -unchanged; embeddings only improve which capabilities are disclosed first. - -The model is loaded lazily from `skills.DefaultModelDir()` and shared through -per-process cached indexes. It is not the conversational model and does not -send user requests or descriptions to a provider. See the `embedding/` and -`skills/` packages for the implementation. +layer before prompt assembly. See [`architecture-embedding.md`](architecture-embedding.md) +for the embedding model, index lifecycle, matching thresholds, and failure +behavior. ## Command resolution