Document lessons from embedding discovery
This commit is contained in:
parent
f407fa2604
commit
a0636b88aa
|
|
@ -97,7 +97,7 @@ Shell scripts, Acme, KDE components, and other clients create sessions, write pr
|
||||||
|
|
||||||
See [`doc/architecture-ide.md`](doc/architecture-ide.md) for integrating Ollie with an editor or IDE.
|
See [`doc/architecture-ide.md`](doc/architecture-ide.md) for integrating Ollie with an editor or IDE.
|
||||||
|
|
||||||
See [`doc/architecture.md`](doc/architecture.md) for component boundaries and [`doc/evolution.md`](doc/evolution.md) for the design history. See [`doc/usage.md`](doc/usage.md) for setup.
|
See [`doc/architecture.md`](doc/architecture.md) for component boundaries and [`doc/evolution.md`](doc/evolution.md) for the design history. See [`doc/lessons-learned.md`](doc/lessons-learned.md) for durable engineering lessons. See [`doc/usage.md`](doc/usage.md) for setup.
|
||||||
|
|
||||||
## Dependencies
|
## Dependencies
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -58,6 +58,8 @@ The architecture is split by responsibility:
|
||||||
- [`architecture-core.md`](architecture-core.md) — agent loop, sessions, history, backends, hooks, and concurrency.
|
- [`architecture-core.md`](architecture-core.md) — agent loop, sessions, history, backends, hooks, and concurrency.
|
||||||
- [`architecture-prompting.md`](architecture-prompting.md) — prompt resolution and runtime preamble assembly.
|
- [`architecture-prompting.md`](architecture-prompting.md) — prompt resolution and runtime preamble assembly.
|
||||||
- [`architecture-tools.md`](architecture-tools.md) — tool metadata and authoring, including metadata-only tools.
|
- [`architecture-tools.md`](architecture-tools.md) — tool metadata and authoring, including metadata-only tools.
|
||||||
|
- [`architecture-embedding.md`](architecture-embedding.md) — semantic tool and skill discovery.
|
||||||
|
- [`lessons-learned.md`](lessons-learned.md) — durable engineering lessons and design principles.
|
||||||
- [`architecture-toolsrv.md`](architecture-toolsrv.md) — tool registry, 9P service, process execution, sandbox, and cancellation.
|
- [`architecture-toolsrv.md`](architecture-toolsrv.md) — tool registry, 9P service, process execution, sandbox, and cancellation.
|
||||||
- [`architecture-remote.md`](architecture-remote.md) — SSH deployment of a remote toolsrv.
|
- [`architecture-remote.md`](architecture-remote.md) — SSH deployment of a remote toolsrv.
|
||||||
- [`architecture-virtfs.md`](architecture-virtfs.md) — the declaration DSL and in-memory filesystem tree.
|
- [`architecture-virtfs.md`](architecture-virtfs.md) — the declaration DSL and in-memory filesystem tree.
|
||||||
|
|
|
||||||
|
|
@ -0,0 +1,45 @@
|
||||||
|
# Lessons Learned
|
||||||
|
|
||||||
|
This document records durable engineering lessons from Ollie development. It
|
||||||
|
captures general principles and observed outcomes rather than a chronological
|
||||||
|
change log. See [`evolution.md`](evolution.md) for the implementation timeline.
|
||||||
|
|
||||||
|
## Small models benefit from pre-optimized discovery
|
||||||
|
|
||||||
|
Embedding-guided discovery acts as a **capability pre-selector** for the
|
||||||
|
conversational model. A local embedding index matches the user's request to
|
||||||
|
tool and skill descriptions before prompt assembly, so the model receives a
|
||||||
|
small, relevant capability set instead of reasoning over the entire catalog.
|
||||||
|
|
||||||
|
This is particularly effective for smaller models, including models around or
|
||||||
|
below 8B parameters. They often have less reliable discovery reasoning and can
|
||||||
|
waste multiple round trips trying the wrong tool or failing to identify a
|
||||||
|
relevant skill. Semantic pre-selection removes much of that discovery burden
|
||||||
|
before generation begins.
|
||||||
|
|
||||||
|
The embedding layer is therefore not a replacement for reasoning. It is a
|
||||||
|
pre-optimizer for the model's action space: it narrows the likely useful
|
||||||
|
capabilities while leaving final selection, tool loading, argument generation,
|
||||||
|
and execution to the normal agent loop.
|
||||||
|
|
||||||
|
## Progressive disclosure beats a complete catalog
|
||||||
|
|
||||||
|
Injecting every tool and skill into every prompt consumes context and makes
|
||||||
|
capability selection harder. Ranking descriptions locally, applying a
|
||||||
|
similarity threshold, and injecting only bounded results preserves context for
|
||||||
|
the task itself. The current discovery limits are three skills and five tool
|
||||||
|
hints per request.
|
||||||
|
|
||||||
|
## Keep discovery separate from execution
|
||||||
|
|
||||||
|
The embedding index suggests relevant capabilities. The skill loader and
|
||||||
|
`toolsrv` remain authoritative for content, metadata, loading, execution,
|
||||||
|
sandboxing, and process state. Separating ranking from execution keeps the
|
||||||
|
system replaceable and prevents a failed or stale index from changing the
|
||||||
|
security boundary.
|
||||||
|
|
||||||
|
## Local models improve privacy and small-model operation
|
||||||
|
|
||||||
|
A local embedding model avoids sending discovery data to the conversational
|
||||||
|
provider and works with offline or inexpensive backends. The embedding model
|
||||||
|
and conversational model can evolve independently.
|
||||||
Loading…
Reference in New Issue