ollie/doc/lessons-learned.md

94 lines
4.6 KiB
Markdown

# Lessons Learned
This document records durable engineering lessons from Ollie development. It
captures general principles and observed outcomes rather than a chronological
change log. See [`evolution.md`](evolution.md) for the implementation timeline.
## Small models benefit from pre-optimized discovery
Embedding-guided discovery acts as a **capability pre-selector** for the
conversational model. A local embedding index matches the user's request to
tool and skill descriptions before prompt assembly, so the model receives a
small, relevant capability set instead of reasoning over the entire catalog.
This is particularly effective for smaller models, including models around or
below 8B parameters. They often have less reliable discovery reasoning and can
waste multiple round trips trying the wrong tool or failing to identify a
relevant skill. Semantic pre-selection removes much of that discovery burden
before generation begins.
The embedding layer is therefore not a replacement for reasoning. It is a
pre-optimizer for the model's action space: it narrows the likely useful
capabilities while leaving final selection, tool loading, argument generation,
and execution to the normal agent loop.
## Architecture beats prompting alone
The larger lesson is that reliable agent behavior cannot be achieved by prompt
wording alone. Prompting tells the model what behavior is desired, but it does
not reliably solve capability discovery for smaller models. The local embedding
model and injected user-facing guidance work together as runtime architecture:
- embeddings select the relevant capabilities before generation;
- injected prompts explain how and when to use those capabilities;
- bounded tool and skill disclosure preserves context for the task;
- the agent loop and toolsrv enforce the resulting execution path.
This combination moves behavior that would otherwise require repeated model
reasoning into deterministic runtime preparation. The model still decides and
acts, but it starts with the right capabilities and explicit operational
instructions. In practice, this is more effective than adding more prompt text
or expecting a small model to discover the complete tool surface unaided.
## Progressive disclosure beats a complete catalog
Injecting every tool and skill into every prompt consumes context and makes
capability selection harder. Ranking descriptions locally, applying a
similarity threshold, and injecting only bounded results preserves context for
the task itself. The current discovery limits are three skills and five tool
hints per request.
## Keep discovery separate from execution
The embedding index suggests relevant capabilities. The skill loader and
`toolsrv` remain authoritative for content, metadata, loading, execution,
sandboxing, and process state. Separating ranking from execution keeps the
system replaceable and prevents a failed or stale index from changing the
security boundary.
## Local models improve privacy and small-model operation
A local embedding model avoids sending discovery data to the conversational
provider and works with offline or inexpensive backends. The embedding model
and conversational model can evolve independently.
## Lazy tool loading is a dead end
Phase 36 attempted a one-tool bootstrap: agents started with only `client_9p`
and loaded other tools on demand through semantic hints. The theory was that
this would create minimal capability surfaces that expanded only as needed.
In practice, this failed for three reasons:
1. **Models don't follow multi-step loading protocols.** When a model wants to
read a file, it calls `file_read`. It doesn't first call `client_9p` to load
`file_read`, then call `file_read`. The indirection step is almost always
skipped regardless of prompt instructions.
2. **The indirection doubles round-trips.** Even when the model follows the
protocol, every tool use requires two generation cycles: one to decide to
load, one to use. This halves throughput for common operations.
3. **Capability boundaries become implicit.** If any agent can load any tool,
the permission model depends on what the model decides to do, not what it's
configured to do. A code reviewer could load `shell` and execute arbitrary
commands. The boundary is only as strong as the model's compliance.
The fix was explicit `autoLoad` declarations in agent configs. Each profile
lists exactly which tools it starts with. Role-specific profiles get role-
appropriate tools. The capability surface is explicit, auditable, and
enforced at config time rather than generation time.
**Lesson:** Don't rely on model compliance for capability boundaries. If
a tool shouldn't be available, don't make it loadable.