ollie/doc/lessons-learned.md

132 lines
6.5 KiB
Markdown

# Lessons Learned
This document records durable engineering lessons from Ollie development. It
captures general principles and observed outcomes rather than a chronological
change log. See [`evolution.md`](evolution.md) for the implementation timeline.
## Small models benefit from pre-optimized discovery
Embedding-guided discovery acts as a **capability pre-selector** for the
conversational model. A local embedding index matches the user's request to
tool and skill descriptions before prompt assembly, so the model receives a
small, relevant capability set instead of reasoning over the entire catalog.
This is particularly effective for smaller models, including models around or
below 8B parameters. They often have less reliable discovery reasoning and can
waste multiple round trips trying the wrong tool or failing to identify a
relevant skill. Semantic pre-selection removes much of that discovery burden
before generation begins.
The embedding layer is therefore not a replacement for reasoning. It is a
pre-optimizer for the model's action space: it narrows the likely useful
capabilities while leaving final selection, tool loading, argument generation,
and execution to the normal agent loop.
## Architecture beats prompting alone
The larger lesson is that reliable agent behavior cannot be achieved by prompt
wording alone. Prompting tells the model what behavior is desired, but it does
not reliably solve capability discovery for smaller models. The local embedding
model and injected user-facing guidance work together as runtime architecture:
- embeddings select the relevant capabilities before generation;
- injected prompts explain how and when to use those capabilities;
- bounded tool and skill disclosure preserves context for the task;
- the agent loop and toolsrv enforce the resulting execution path.
This combination moves behavior that would otherwise require repeated model
reasoning into deterministic runtime preparation. The model still decides and
acts, but it starts with the right capabilities and explicit operational
instructions. In practice, this is more effective than adding more prompt text
or expecting a small model to discover the complete tool surface unaided.
## Progressive disclosure beats a complete catalog
Injecting every tool and skill into every prompt consumes context and makes
capability selection harder. Ranking descriptions locally, applying a
similarity threshold, and injecting only bounded results preserves context for
the task itself. The current discovery limits are three skills and five tool
hints per request.
## Keep discovery separate from execution
The embedding index suggests relevant capabilities. The skill loader and
`toolsrv` remain authoritative for content, metadata, loading, execution,
sandboxing, and process state. Separating ranking from execution keeps the
system replaceable and prevents a failed or stale index from changing the
security boundary.
## Local models improve privacy and small-model operation
A local embedding model avoids sending discovery data to the conversational
provider and works with offline or inexpensive backends. The embedding model
and conversational model can evolve independently.
## Lazy tool loading is a dead end
Phase 36 attempted a one-tool bootstrap: agents started with only `client_9p`
and loaded other tools on demand through semantic hints. The theory was that
this would create minimal capability surfaces that expanded only as needed.
In practice, this failed for three reasons:
1. **Models don't follow multi-step loading protocols.** When a model wants to
read a file, it calls `file_read`. It doesn't first call `client_9p` to load
`file_read`, then call `file_read`. The indirection step is almost always
skipped regardless of prompt instructions.
2. **The indirection doubles round-trips.** Even when the model follows the
protocol, every tool use requires two generation cycles: one to decide to
load, one to use. This halves throughput for common operations.
3. **Capability boundaries become implicit.** If any agent can load any tool,
the permission model depends on what the model decides to do, not what it's
configured to do. A code reviewer could load `shell` and execute arbitrary
commands. The boundary is only as strong as the model's compliance.
The fix was explicit `autoLoad` declarations in agent configs. Each profile
lists exactly which tools it starts with. Role-specific profiles get role-
appropriate tools. The capability surface is explicit, auditable, and
enforced at config time rather than generation time.
**Lesson:** Don't rely on model compliance for capability boundaries. If
a tool shouldn't be available, don't make it loadable.
## Model compliance is not a security boundary
The lazy loading failure (above) reveals a deeper principle: **behavioral
defenses are not structural guarantees**. Security properties that depend on
the model following instructions are only as strong as the model's compliance.
This distinction matters for agent security:
| Defense Type | Mechanism | Failure Mode |
|-------------|-----------|--------------|
| **Behavioral** | Prompt instructions, model refusal | Model ignores/misinterprets instructions |
| **Structural** | Config-time enforcement, OS-level sandbox | Requires implementation bug to bypass |
Ollie's security model uses structural defenses:
1. **Tool access**: Per-agent registry populated at config time via `autoLoad`.
An agent cannot call tools not in its registry — `Lookup()` fails before
any tool code runs. This is structural: the model cannot comply its way
past a missing registry entry.
2. **Execution sandbox**: Landlock kernel enforcement. Even if an agent calls
`shell` with a malicious command, the kernel blocks filesystem access
outside allowed paths. This is structural: the model cannot prompt-inject
its way past a syscall filter.
3. **Bypass broker**: Escape from sandbox requires explicit user approval.
The model can request bypass, but cannot grant it to itself.
The empirical finding that models skip loading protocols is one instance of
a general pattern: models vary in whether they follow behavioral expectations.
Some models (Claude) refuse malicious instructions reliably. Others (GPT-4o)
attempt them. A security model that depends on refusal breaks when the model
doesn't refuse.
**Lesson:** Structural enforcement (registry checks, kernel sandboxing) holds
regardless of model behavior. Behavioral enforcement (prompt instructions,
model refusal) is defense-in-depth, not a primary security boundary.