From a0636b88aafd1ff4330e6d9120b0f9773abc0229 Mon Sep 17 00:00:00 2001 From: Ollie Agent Date: Thu, 20 Aug 2026 17:44:40 +0200 Subject: [PATCH] Document lessons from embedding discovery --- README.md | 2 +- doc/architecture.md | 2 ++ doc/lessons-learned.md | 45 ++++++++++++++++++++++++++++++++++++++++++ 3 files changed, 48 insertions(+), 1 deletion(-) create mode 100644 doc/lessons-learned.md diff --git a/README.md b/README.md index 1441d24..67e0614 100644 --- a/README.md +++ b/README.md @@ -97,7 +97,7 @@ Shell scripts, Acme, KDE components, and other clients create sessions, write pr See [`doc/architecture-ide.md`](doc/architecture-ide.md) for integrating Ollie with an editor or IDE. -See [`doc/architecture.md`](doc/architecture.md) for component boundaries and [`doc/evolution.md`](doc/evolution.md) for the design history. See [`doc/usage.md`](doc/usage.md) for setup. +See [`doc/architecture.md`](doc/architecture.md) for component boundaries and [`doc/evolution.md`](doc/evolution.md) for the design history. See [`doc/lessons-learned.md`](doc/lessons-learned.md) for durable engineering lessons. See [`doc/usage.md`](doc/usage.md) for setup. ## Dependencies diff --git a/doc/architecture.md b/doc/architecture.md index 9acf23d..52b4af9 100644 --- a/doc/architecture.md +++ b/doc/architecture.md @@ -58,6 +58,8 @@ The architecture is split by responsibility: - [`architecture-core.md`](architecture-core.md) — agent loop, sessions, history, backends, hooks, and concurrency. - [`architecture-prompting.md`](architecture-prompting.md) — prompt resolution and runtime preamble assembly. - [`architecture-tools.md`](architecture-tools.md) — tool metadata and authoring, including metadata-only tools. +- [`architecture-embedding.md`](architecture-embedding.md) — semantic tool and skill discovery. +- [`lessons-learned.md`](lessons-learned.md) — durable engineering lessons and design principles. - [`architecture-toolsrv.md`](architecture-toolsrv.md) — tool registry, 9P service, process execution, sandbox, and cancellation. - [`architecture-remote.md`](architecture-remote.md) — SSH deployment of a remote toolsrv. - [`architecture-virtfs.md`](architecture-virtfs.md) — the declaration DSL and in-memory filesystem tree. diff --git a/doc/lessons-learned.md b/doc/lessons-learned.md new file mode 100644 index 0000000..e84f6a2 --- /dev/null +++ b/doc/lessons-learned.md @@ -0,0 +1,45 @@ +# Lessons Learned + +This document records durable engineering lessons from Ollie development. It +captures general principles and observed outcomes rather than a chronological +change log. See [`evolution.md`](evolution.md) for the implementation timeline. + +## Small models benefit from pre-optimized discovery + +Embedding-guided discovery acts as a **capability pre-selector** for the +conversational model. A local embedding index matches the user's request to +tool and skill descriptions before prompt assembly, so the model receives a +small, relevant capability set instead of reasoning over the entire catalog. + +This is particularly effective for smaller models, including models around or +below 8B parameters. They often have less reliable discovery reasoning and can +waste multiple round trips trying the wrong tool or failing to identify a +relevant skill. Semantic pre-selection removes much of that discovery burden +before generation begins. + +The embedding layer is therefore not a replacement for reasoning. It is a +pre-optimizer for the model's action space: it narrows the likely useful +capabilities while leaving final selection, tool loading, argument generation, +and execution to the normal agent loop. + +## Progressive disclosure beats a complete catalog + +Injecting every tool and skill into every prompt consumes context and makes +capability selection harder. Ranking descriptions locally, applying a +similarity threshold, and injecting only bounded results preserves context for +the task itself. The current discovery limits are three skills and five tool +hints per request. + +## Keep discovery separate from execution + +The embedding index suggests relevant capabilities. The skill loader and +`toolsrv` remain authoritative for content, metadata, loading, execution, +sandboxing, and process state. Separating ranking from execution keeps the +system replaceable and prevents a failed or stale index from changing the +security boundary. + +## Local models improve privacy and small-model operation + +A local embedding model avoids sending discovery data to the conversational +provider and works with offline or inexpensive backends. The embedding model +and conversational model can evolve independently.