94 lines
4.6 KiB
Markdown
94 lines
4.6 KiB
Markdown
# Lessons Learned
|
|
|
|
This document records durable engineering lessons from Ollie development. It
|
|
captures general principles and observed outcomes rather than a chronological
|
|
change log. See [`evolution.md`](evolution.md) for the implementation timeline.
|
|
|
|
## Small models benefit from pre-optimized discovery
|
|
|
|
Embedding-guided discovery acts as a **capability pre-selector** for the
|
|
conversational model. A local embedding index matches the user's request to
|
|
tool and skill descriptions before prompt assembly, so the model receives a
|
|
small, relevant capability set instead of reasoning over the entire catalog.
|
|
|
|
This is particularly effective for smaller models, including models around or
|
|
below 8B parameters. They often have less reliable discovery reasoning and can
|
|
waste multiple round trips trying the wrong tool or failing to identify a
|
|
relevant skill. Semantic pre-selection removes much of that discovery burden
|
|
before generation begins.
|
|
|
|
The embedding layer is therefore not a replacement for reasoning. It is a
|
|
pre-optimizer for the model's action space: it narrows the likely useful
|
|
capabilities while leaving final selection, tool loading, argument generation,
|
|
and execution to the normal agent loop.
|
|
|
|
## Architecture beats prompting alone
|
|
|
|
The larger lesson is that reliable agent behavior cannot be achieved by prompt
|
|
wording alone. Prompting tells the model what behavior is desired, but it does
|
|
not reliably solve capability discovery for smaller models. The local embedding
|
|
model and injected user-facing guidance work together as runtime architecture:
|
|
|
|
- embeddings select the relevant capabilities before generation;
|
|
- injected prompts explain how and when to use those capabilities;
|
|
- bounded tool and skill disclosure preserves context for the task;
|
|
- the agent loop and toolsrv enforce the resulting execution path.
|
|
|
|
This combination moves behavior that would otherwise require repeated model
|
|
reasoning into deterministic runtime preparation. The model still decides and
|
|
acts, but it starts with the right capabilities and explicit operational
|
|
instructions. In practice, this is more effective than adding more prompt text
|
|
or expecting a small model to discover the complete tool surface unaided.
|
|
|
|
## Progressive disclosure beats a complete catalog
|
|
|
|
Injecting every tool and skill into every prompt consumes context and makes
|
|
capability selection harder. Ranking descriptions locally, applying a
|
|
similarity threshold, and injecting only bounded results preserves context for
|
|
the task itself. The current discovery limits are three skills and five tool
|
|
hints per request.
|
|
|
|
## Keep discovery separate from execution
|
|
|
|
The embedding index suggests relevant capabilities. The skill loader and
|
|
`toolsrv` remain authoritative for content, metadata, loading, execution,
|
|
sandboxing, and process state. Separating ranking from execution keeps the
|
|
system replaceable and prevents a failed or stale index from changing the
|
|
security boundary.
|
|
|
|
## Local models improve privacy and small-model operation
|
|
|
|
A local embedding model avoids sending discovery data to the conversational
|
|
provider and works with offline or inexpensive backends. The embedding model
|
|
and conversational model can evolve independently.
|
|
|
|
## Lazy tool loading is a dead end
|
|
|
|
Phase 36 attempted a one-tool bootstrap: agents started with only `client_9p`
|
|
and loaded other tools on demand through semantic hints. The theory was that
|
|
this would create minimal capability surfaces that expanded only as needed.
|
|
|
|
In practice, this failed for three reasons:
|
|
|
|
1. **Models don't follow multi-step loading protocols.** When a model wants to
|
|
read a file, it calls `file_read`. It doesn't first call `client_9p` to load
|
|
`file_read`, then call `file_read`. The indirection step is almost always
|
|
skipped regardless of prompt instructions.
|
|
|
|
2. **The indirection doubles round-trips.** Even when the model follows the
|
|
protocol, every tool use requires two generation cycles: one to decide to
|
|
load, one to use. This halves throughput for common operations.
|
|
|
|
3. **Capability boundaries become implicit.** If any agent can load any tool,
|
|
the permission model depends on what the model decides to do, not what it's
|
|
configured to do. A code reviewer could load `shell` and execute arbitrary
|
|
commands. The boundary is only as strong as the model's compliance.
|
|
|
|
The fix was explicit `autoLoad` declarations in agent configs. Each profile
|
|
lists exactly which tools it starts with. Role-specific profiles get role-
|
|
appropriate tools. The capability surface is explicit, auditable, and
|
|
enforced at config time rather than generation time.
|
|
|
|
**Lesson:** Don't rely on model compliance for capability boundaries. If
|
|
a tool shouldn't be available, don't make it loadable.
|