Document embedding discovery architecture
This commit is contained in:
parent
7bf9b6da9b
commit
ba91920a88
|
|
@ -0,0 +1,156 @@
|
|||
# Embedding Architecture: Tool and Skill Discovery
|
||||
|
||||
Ollie uses a local text-embedding model to select relevant skills and tool
|
||||
hints before assembling each model prompt. This is a discovery and progressive
|
||||
disclosure subsystem. It does not execute tools, replace the tool registry, or
|
||||
act as the conversational backend.
|
||||
|
||||
## Purpose
|
||||
|
||||
Exact-name discovery is fragile for small models. A user may ask to “inspect
|
||||
source structure” without naming `code_outline`, or ask for domain guidance
|
||||
without knowing the skill name. Embedding descriptions makes the capability
|
||||
surface searchable by meaning while keeping the prompt bounded.
|
||||
|
||||
The mechanism is deliberately local. User requests, tool descriptions, and
|
||||
skill descriptions are embedded by an ONNX sentence-transformer on the Ollie
|
||||
host. No discovery request is sent to the configured LLM provider.
|
||||
|
||||
## Components
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
Q[User request] --> A[Agent turn]
|
||||
A --> SI[Skill index]
|
||||
A --> TI[Tool index]
|
||||
SI --> E[embedding.Model]
|
||||
TI --> E
|
||||
E --> ONNX[all-MiniLM-L6-v2 ONNX model]
|
||||
SI --> SH[Relevant skill content]
|
||||
TI --> TH[Relevant tool hints]
|
||||
SH --> P[Prompt assembly]
|
||||
TH --> P
|
||||
P --> L[Conversational backend]
|
||||
```
|
||||
|
||||
### `embedding` package
|
||||
|
||||
`embedding.Model` loads three runtime assets from the model directory:
|
||||
|
||||
- `model.onnx` — the `all-MiniLM-L6-v2` sentence-transformer.
|
||||
- `tokenizer.json` — tokenizer vocabulary and configuration.
|
||||
- `libonnxruntime.so` — ONNX Runtime shared library.
|
||||
|
||||
The model produces 384-dimensional vectors. `Embed` tokenizes text, runs the
|
||||
ONNX session, performs attention-masked mean pooling, and L2-normalizes the
|
||||
result. `CosineSimilarity` compares query and description vectors.
|
||||
|
||||
ONNX Runtime initialization is process-global and occurs once. Model sessions
|
||||
are mutex-protected because inference uses shared session state. Each index
|
||||
owns its model session and releases it through `Close`.
|
||||
|
||||
### `skills.Index`
|
||||
|
||||
The shared index type handles both skills and tools:
|
||||
|
||||
```text
|
||||
Index {
|
||||
model: embedding.Model
|
||||
entries: []Skill
|
||||
vectors: []Vector
|
||||
}
|
||||
```
|
||||
|
||||
For skills, the index scans configured directories for `<name>/SKILL.md`,
|
||||
parses `name` and `description` from YAML frontmatter, and retains the full
|
||||
content for injection. For tools, the agent converts `.meta` descriptions into
|
||||
the same `Skill` representation; only the name and description are needed for
|
||||
ranking.
|
||||
|
||||
Description vectors are precomputed when an index is built or reloaded. A user
|
||||
request is embedded at match time, then compared against all indexed vectors.
|
||||
Results below the threshold are discarded; remaining results are sorted by
|
||||
cosine score and capped by the caller.
|
||||
|
||||
## Agent integration
|
||||
|
||||
The agent keeps independent cached indexes for skills and tools. Initialization
|
||||
is lazy and guarded by `sync.Once`:
|
||||
|
||||
```text
|
||||
first turn
|
||||
├── build skill index from configured SKILL.md directories
|
||||
└── build tool index from installed .meta files
|
||||
|
||||
each turn
|
||||
├── embed request
|
||||
├── match skills: threshold 0.35, limit 3
|
||||
└── match tools: threshold 0.35, limit 5
|
||||
```
|
||||
|
||||
Relevant skills are injected as a `<skills>` block containing complete skill
|
||||
content. Relevant tools are injected as a `<tool-hints>` block containing
|
||||
callable names and descriptions. The normal tool registry still controls
|
||||
whether a tool is loaded and executable; semantic matching only improves what
|
||||
the model sees first.
|
||||
|
||||
The bounded result limits prevent discovery from becoming a prompt-size or
|
||||
context-budget problem. This progressive disclosure is especially useful for
|
||||
models at or below 8B parameters: the model receives a small, request-specific
|
||||
capability set instead of a large undifferentiated catalog.
|
||||
|
||||
## Configuration and installation
|
||||
|
||||
The default model directory is:
|
||||
|
||||
```text
|
||||
$XDG_DATA_HOME/ollie/models
|
||||
```
|
||||
|
||||
or `~/.local/share/ollie/models` when `XDG_DATA_HOME` is unset. `make
|
||||
install-models` downloads the model, tokenizer, and ONNX Runtime. The default
|
||||
skill directory is:
|
||||
|
||||
```text
|
||||
$XDG_CONFIG_HOME/ollie/skills
|
||||
```
|
||||
|
||||
or `~/.config/ollie/skills`. Skill search paths can be overridden with one path
|
||||
per line in `embedding.conf` in the Ollie configuration directory. Blank lines
|
||||
and comments are ignored; environment variables and `~` are expanded.
|
||||
|
||||
## Failure behavior and lifecycle
|
||||
|
||||
Embedding is an enhancement, not a hard dependency for the agent loop. If the
|
||||
model cannot load, a skill directory is missing, a metadata file is malformed,
|
||||
or an individual embedding fails, semantic injection is skipped or that entry
|
||||
is omitted. The request continues through ordinary prompt assembly and tool
|
||||
loading.
|
||||
|
||||
Indexes are process-local. They are not persisted as vector databases. A new
|
||||
process rebuilds vectors from current skill files and tool metadata. The
|
||||
existing lazy index cache means repeated turns reuse loaded models and vectors.
|
||||
Changes to skills or tool metadata require index reload or a new process before
|
||||
they affect ranking.
|
||||
|
||||
## Design constraints
|
||||
|
||||
- Keep the conversational backend independent from discovery.
|
||||
- Keep `toolsrv` authoritative for metadata, loading, execution, sandboxing,
|
||||
and process state.
|
||||
- Keep descriptions short and meaningful because they are the indexed corpus.
|
||||
- Keep ranking bounded; semantic discovery must not inflate every prompt.
|
||||
- Treat embedding failure as a capability-discovery miss, not a turn failure.
|
||||
- Keep the model local to preserve privacy, deterministic deployment, and
|
||||
operation with small or offline conversational models.
|
||||
|
||||
## Source map
|
||||
|
||||
- `embedding/embedding.go` — model loading, tokenization, inference, pooling,
|
||||
normalization, and similarity.
|
||||
- `skills/skills.go` — skill scanning, frontmatter parsing, indexing, matching,
|
||||
and reload.
|
||||
- `cmd/olliesrv/internal/agent/skill_match.go` — lazy indexes and prompt
|
||||
injection for skills and tool hints.
|
||||
- `Makefile` — model and runtime asset installation.
|
||||
- [`architecture-tools.md`](architecture-tools.md) — tool metadata and execution.
|
||||
|
|
@ -72,26 +72,9 @@ The `.meta` file is the sole source of discovery and schema. The registry never
|
|||
|
||||
Tool discovery remains metadata-driven: `toolsrv` scans `.meta` files and manages
|
||||
which tools are loaded for an agent. Ollie adds a separate semantic ranking
|
||||
layer before prompt assembly:
|
||||
|
||||
```text
|
||||
user request
|
||||
│
|
||||
├── embedding.Model.Embed()
|
||||
│
|
||||
├── cosine similarity → skill descriptions → up to 3 injected skills
|
||||
└── cosine similarity → tool descriptions → up to 5 injected hints
|
||||
```
|
||||
|
||||
The shared `skills.Index` precomputes vectors for skill or tool descriptions and
|
||||
returns results above a configured threshold in descending relevance order. The
|
||||
agent keeps the full tool registry and callable tool loading semantics
|
||||
unchanged; embeddings only improve which capabilities are disclosed first.
|
||||
|
||||
The model is loaded lazily from `skills.DefaultModelDir()` and shared through
|
||||
per-process cached indexes. It is not the conversational model and does not
|
||||
send user requests or descriptions to a provider. See the `embedding/` and
|
||||
`skills/` packages for the implementation.
|
||||
layer before prompt assembly. See [`architecture-embedding.md`](architecture-embedding.md)
|
||||
for the embedding model, index lifecycle, matching thresholds, and failure
|
||||
behavior.
|
||||
|
||||
## Command resolution
|
||||
|
||||
|
|
|
|||
Loading…
Reference in New Issue