6.1 KiB
Embedding Architecture: Tool and Skill Discovery
Ollie uses a local text-embedding model to select relevant skills and tool hints before assembling each model prompt. This is a discovery and progressive disclosure subsystem. It does not execute tools, replace the tool registry, or act as the conversational backend.
Purpose
Exact-name discovery is fragile for small models. A user may ask to “inspect
source structure” without naming code_outline, or ask for domain guidance
without knowing the skill name. Embedding descriptions makes the capability
surface searchable by meaning while keeping the prompt bounded.
The mechanism is deliberately local. User requests, tool descriptions, and skill descriptions are embedded by an ONNX sentence-transformer on the Ollie host. No discovery request is sent to the configured LLM provider.
Components
flowchart TB
Q[User request] --> A[Agent turn]
A --> SI[Skill index]
A --> TI[Tool index]
SI --> E[embedding.Model]
TI --> E
E --> ONNX[all-MiniLM-L6-v2 ONNX model]
SI --> SH[Relevant skill content]
TI --> TH[Relevant tool hints]
SH --> P[Prompt assembly]
TH --> P
P --> L[Conversational backend]
embedding package
embedding.Model loads three runtime assets from the model directory:
model.onnx— theall-MiniLM-L6-v2sentence-transformer.tokenizer.json— tokenizer vocabulary and configuration.libonnxruntime.so— ONNX Runtime shared library.
The model produces 384-dimensional vectors. Embed tokenizes text, runs the
ONNX session, performs attention-masked mean pooling, and L2-normalizes the
result. CosineSimilarity compares query and description vectors.
ONNX Runtime initialization is process-global and occurs once. Model sessions
are mutex-protected because inference uses shared session state. Each index
owns its model session and releases it through Close.
embedding.Index[T]
The generic index type handles arbitrary entry types:
Index[T] {
model: *Model
entries: []T
vectors: []Vector
textFn: func(T) string // extracts text for embedding
}
Skill and tool indexes use the same generic structure with different entry
types. For skills, entries are skills.Skill structs with name, description,
and content. For tools, entries are toolsrv.Meta structs from loaded tool
metadata.
Description vectors are precomputed when an index is built or reloaded. A user request is embedded at match time, then compared against all indexed vectors. Results below the threshold are discarded; remaining results are sorted by cosine score and capped by the caller.
Agent integration
The agent keeps independent cached indexes for skills and tools. Skill index
initialization is lazy and guarded by sync.Once. Tool index is rebuilt per
turn from the agent's currently loaded tools:
first turn
└── build skill index from configured SKILL.md directories
each turn
├── build tool index from agent's loaded tools
├── embed request
├── match skills: threshold 0.35, limit 3
└── match tools: threshold 0.35, limit 5
Relevant skills are injected as a <skills> block containing complete skill
content. Relevant tools are injected as a <tool-hints> block containing
loaded tool names and descriptions. Tools must be explicitly loaded through
the agent's autoLoad config or via the ctl file before they appear in
matching results.
The bounded result limits prevent discovery from becoming a prompt-size or context-budget problem. This progressive disclosure is especially useful for models at or below 8B parameters: the model receives a small, request-specific capability set instead of a large undifferentiated catalog.
Configuration and installation
The default model directory is:
$XDG_DATA_HOME/ollie/models
or ~/.local/share/ollie/models when XDG_DATA_HOME is unset. make install-models downloads the model, tokenizer, and ONNX Runtime. The default
skill directory is:
$XDG_CONFIG_HOME/ollie/skills
or ~/.config/ollie/skills. Skill search paths can be overridden with one path
per line in embedding.conf in the Ollie configuration directory. Blank lines
and comments are ignored; environment variables and ~ are expanded.
Failure behavior and lifecycle
Embedding is an enhancement, not a hard dependency for the agent loop. If the model cannot load, a skill directory is missing, a metadata file is malformed, or an individual embedding fails, semantic injection is skipped or that entry is omitted. The request continues through ordinary prompt assembly and tool loading.
Indexes are process-local. They are not persisted as vector databases. A new process rebuilds vectors from current skill files and tool metadata. The existing lazy index cache means repeated turns reuse loaded models and vectors. Changes to skills or tool metadata require index reload or a new process before they affect ranking.
Design constraints
- Keep the conversational backend independent from discovery.
- Keep
toolsrvauthoritative for metadata, loading, execution, sandboxing, and process state. - Keep descriptions short and meaningful because they are the indexed corpus.
- Keep ranking bounded; semantic discovery must not inflate every prompt.
- Treat embedding failure as a capability-discovery miss, not a turn failure.
- Keep the model local to preserve privacy, deterministic deployment, and operation with small or offline conversational models.
Source map
embedding/embedding.go— model loading, tokenization, inference, pooling, normalization, and similarity.embedding/index.go— genericIndex[T]for vector-indexed entries.skills/skills.go— skill scanning, frontmatter parsing, andSkilltype.cmd/olliesrv/internal/agent/skill_match.go— lazy skill index and prompt injection for skills.cmd/olliesrv/internal/agent/tool_match.go— per-turn tool index from loaded tools and prompt injection for tool hints.Makefile— model and runtime asset installation.architecture-tools.md— tool metadata and execution.