4.6 KiB
Lessons Learned
This document records durable engineering lessons from Ollie development. It
captures general principles and observed outcomes rather than a chronological
change log. See evolution.md for the implementation timeline.
Small models benefit from pre-optimized discovery
Embedding-guided discovery acts as a capability pre-selector for the conversational model. A local embedding index matches the user's request to tool and skill descriptions before prompt assembly, so the model receives a small, relevant capability set instead of reasoning over the entire catalog.
This is particularly effective for smaller models, including models around or below 8B parameters. They often have less reliable discovery reasoning and can waste multiple round trips trying the wrong tool or failing to identify a relevant skill. Semantic pre-selection removes much of that discovery burden before generation begins.
The embedding layer is therefore not a replacement for reasoning. It is a pre-optimizer for the model's action space: it narrows the likely useful capabilities while leaving final selection, tool loading, argument generation, and execution to the normal agent loop.
Architecture beats prompting alone
The larger lesson is that reliable agent behavior cannot be achieved by prompt wording alone. Prompting tells the model what behavior is desired, but it does not reliably solve capability discovery for smaller models. The local embedding model and injected user-facing guidance work together as runtime architecture:
- embeddings select the relevant capabilities before generation;
- injected prompts explain how and when to use those capabilities;
- bounded tool and skill disclosure preserves context for the task;
- the agent loop and toolsrv enforce the resulting execution path.
This combination moves behavior that would otherwise require repeated model reasoning into deterministic runtime preparation. The model still decides and acts, but it starts with the right capabilities and explicit operational instructions. In practice, this is more effective than adding more prompt text or expecting a small model to discover the complete tool surface unaided.
Progressive disclosure beats a complete catalog
Injecting every tool and skill into every prompt consumes context and makes capability selection harder. Ranking descriptions locally, applying a similarity threshold, and injecting only bounded results preserves context for the task itself. The current discovery limits are three skills and five tool hints per request.
Keep discovery separate from execution
The embedding index suggests relevant capabilities. The skill loader and
toolsrv remain authoritative for content, metadata, loading, execution,
sandboxing, and process state. Separating ranking from execution keeps the
system replaceable and prevents a failed or stale index from changing the
security boundary.
Local models improve privacy and small-model operation
A local embedding model avoids sending discovery data to the conversational provider and works with offline or inexpensive backends. The embedding model and conversational model can evolve independently.
Lazy tool loading is a dead end
Phase 36 attempted a one-tool bootstrap: agents started with only client_9p
and loaded other tools on demand through semantic hints. The theory was that
this would create minimal capability surfaces that expanded only as needed.
In practice, this failed for three reasons:
-
Models don't follow multi-step loading protocols. When a model wants to read a file, it calls
file_read. It doesn't first callclient_9pto loadfile_read, then callfile_read. The indirection step is almost always skipped regardless of prompt instructions. -
The indirection doubles round-trips. Even when the model follows the protocol, every tool use requires two generation cycles: one to decide to load, one to use. This halves throughput for common operations.
-
Capability boundaries become implicit. If any agent can load any tool, the permission model depends on what the model decides to do, not what it's configured to do. A code reviewer could load
shelland execute arbitrary commands. The boundary is only as strong as the model's compliance.
The fix was explicit autoLoad declarations in agent configs. Each profile
lists exactly which tools it starts with. Role-specific profiles get role-
appropriate tools. The capability surface is explicit, auditable, and
enforced at config time rather than generation time.
Lesson: Don't rely on model compliance for capability boundaries. If a tool shouldn't be available, don't make it loadable.