Document embedding discovery architecture

This commit is contained in:
Ollie Agent 2026-08-20 17:41:56 +02:00
parent 7bf9b6da9b
commit ba91920a88
2 changed files with 159 additions and 20 deletions

View File

@ -0,0 +1,156 @@
# Embedding Architecture: Tool and Skill Discovery
Ollie uses a local text-embedding model to select relevant skills and tool
hints before assembling each model prompt. This is a discovery and progressive
disclosure subsystem. It does not execute tools, replace the tool registry, or
act as the conversational backend.
## Purpose
Exact-name discovery is fragile for small models. A user may ask to “inspect
source structure” without naming `code_outline`, or ask for domain guidance
without knowing the skill name. Embedding descriptions makes the capability
surface searchable by meaning while keeping the prompt bounded.
The mechanism is deliberately local. User requests, tool descriptions, and
skill descriptions are embedded by an ONNX sentence-transformer on the Ollie
host. No discovery request is sent to the configured LLM provider.
## Components
```mermaid
flowchart TB
Q[User request] --> A[Agent turn]
A --> SI[Skill index]
A --> TI[Tool index]
SI --> E[embedding.Model]
TI --> E
E --> ONNX[all-MiniLM-L6-v2 ONNX model]
SI --> SH[Relevant skill content]
TI --> TH[Relevant tool hints]
SH --> P[Prompt assembly]
TH --> P
P --> L[Conversational backend]
```
### `embedding` package
`embedding.Model` loads three runtime assets from the model directory:
- `model.onnx` — the `all-MiniLM-L6-v2` sentence-transformer.
- `tokenizer.json` — tokenizer vocabulary and configuration.
- `libonnxruntime.so` — ONNX Runtime shared library.
The model produces 384-dimensional vectors. `Embed` tokenizes text, runs the
ONNX session, performs attention-masked mean pooling, and L2-normalizes the
result. `CosineSimilarity` compares query and description vectors.
ONNX Runtime initialization is process-global and occurs once. Model sessions
are mutex-protected because inference uses shared session state. Each index
owns its model session and releases it through `Close`.
### `skills.Index`
The shared index type handles both skills and tools:
```text
Index {
model: embedding.Model
entries: []Skill
vectors: []Vector
}
```
For skills, the index scans configured directories for `<name>/SKILL.md`,
parses `name` and `description` from YAML frontmatter, and retains the full
content for injection. For tools, the agent converts `.meta` descriptions into
the same `Skill` representation; only the name and description are needed for
ranking.
Description vectors are precomputed when an index is built or reloaded. A user
request is embedded at match time, then compared against all indexed vectors.
Results below the threshold are discarded; remaining results are sorted by
cosine score and capped by the caller.
## Agent integration
The agent keeps independent cached indexes for skills and tools. Initialization
is lazy and guarded by `sync.Once`:
```text
first turn
├── build skill index from configured SKILL.md directories
└── build tool index from installed .meta files
each turn
├── embed request
├── match skills: threshold 0.35, limit 3
└── match tools: threshold 0.35, limit 5
```
Relevant skills are injected as a `<skills>` block containing complete skill
content. Relevant tools are injected as a `<tool-hints>` block containing
callable names and descriptions. The normal tool registry still controls
whether a tool is loaded and executable; semantic matching only improves what
the model sees first.
The bounded result limits prevent discovery from becoming a prompt-size or
context-budget problem. This progressive disclosure is especially useful for
models at or below 8B parameters: the model receives a small, request-specific
capability set instead of a large undifferentiated catalog.
## Configuration and installation
The default model directory is:
```text
$XDG_DATA_HOME/ollie/models
```
or `~/.local/share/ollie/models` when `XDG_DATA_HOME` is unset. `make
install-models` downloads the model, tokenizer, and ONNX Runtime. The default
skill directory is:
```text
$XDG_CONFIG_HOME/ollie/skills
```
or `~/.config/ollie/skills`. Skill search paths can be overridden with one path
per line in `embedding.conf` in the Ollie configuration directory. Blank lines
and comments are ignored; environment variables and `~` are expanded.
## Failure behavior and lifecycle
Embedding is an enhancement, not a hard dependency for the agent loop. If the
model cannot load, a skill directory is missing, a metadata file is malformed,
or an individual embedding fails, semantic injection is skipped or that entry
is omitted. The request continues through ordinary prompt assembly and tool
loading.
Indexes are process-local. They are not persisted as vector databases. A new
process rebuilds vectors from current skill files and tool metadata. The
existing lazy index cache means repeated turns reuse loaded models and vectors.
Changes to skills or tool metadata require index reload or a new process before
they affect ranking.
## Design constraints
- Keep the conversational backend independent from discovery.
- Keep `toolsrv` authoritative for metadata, loading, execution, sandboxing,
and process state.
- Keep descriptions short and meaningful because they are the indexed corpus.
- Keep ranking bounded; semantic discovery must not inflate every prompt.
- Treat embedding failure as a capability-discovery miss, not a turn failure.
- Keep the model local to preserve privacy, deterministic deployment, and
operation with small or offline conversational models.
## Source map
- `embedding/embedding.go` — model loading, tokenization, inference, pooling,
normalization, and similarity.
- `skills/skills.go` — skill scanning, frontmatter parsing, indexing, matching,
and reload.
- `cmd/olliesrv/internal/agent/skill_match.go` — lazy indexes and prompt
injection for skills and tool hints.
- `Makefile` — model and runtime asset installation.
- [`architecture-tools.md`](architecture-tools.md) — tool metadata and execution.

View File

@ -72,26 +72,9 @@ The `.meta` file is the sole source of discovery and schema. The registry never
Tool discovery remains metadata-driven: `toolsrv` scans `.meta` files and manages
which tools are loaded for an agent. Ollie adds a separate semantic ranking
layer before prompt assembly:
```text
user request
│
├── embedding.Model.Embed()
│
├── cosine similarity → skill descriptions → up to 3 injected skills
└── cosine similarity → tool descriptions → up to 5 injected hints
```
The shared `skills.Index` precomputes vectors for skill or tool descriptions and
returns results above a configured threshold in descending relevance order. The
agent keeps the full tool registry and callable tool loading semantics
unchanged; embeddings only improve which capabilities are disclosed first.
The model is loaded lazily from `skills.DefaultModelDir()` and shared through
per-process cached indexes. It is not the conversational model and does not
send user requests or descriptions to a provider. See the `embedding/` and
`skills/` packages for the implementation.
layer before prompt assembly. See [`architecture-embedding.md`](architecture-embedding.md)
for the embedding model, index lifecycle, matching thresholds, and failure
behavior.
## Command resolution