ollie/doc/lessons-learned.md

6.5 KiB

Lessons Learned

This document records durable engineering lessons from Ollie development. It captures general principles and observed outcomes rather than a chronological change log. See evolution.md for the implementation timeline.

Small models benefit from pre-optimized discovery

Embedding-guided discovery acts as a capability pre-selector for the conversational model. A local embedding index matches the user's request to tool and skill descriptions before prompt assembly, so the model receives a small, relevant capability set instead of reasoning over the entire catalog.

This is particularly effective for smaller models, including models around or below 8B parameters. They often have less reliable discovery reasoning and can waste multiple round trips trying the wrong tool or failing to identify a relevant skill. Semantic pre-selection removes much of that discovery burden before generation begins.

The embedding layer is therefore not a replacement for reasoning. It is a pre-optimizer for the model's action space: it narrows the likely useful capabilities while leaving final selection, tool loading, argument generation, and execution to the normal agent loop.

Architecture beats prompting alone

The larger lesson is that reliable agent behavior cannot be achieved by prompt wording alone. Prompting tells the model what behavior is desired, but it does not reliably solve capability discovery for smaller models. The local embedding model and injected user-facing guidance work together as runtime architecture:

  • embeddings select the relevant capabilities before generation;
  • injected prompts explain how and when to use those capabilities;
  • bounded tool and skill disclosure preserves context for the task;
  • the agent loop and toolsrv enforce the resulting execution path.

This combination moves behavior that would otherwise require repeated model reasoning into deterministic runtime preparation. The model still decides and acts, but it starts with the right capabilities and explicit operational instructions. In practice, this is more effective than adding more prompt text or expecting a small model to discover the complete tool surface unaided.

Progressive disclosure beats a complete catalog

Injecting every tool and skill into every prompt consumes context and makes capability selection harder. Ranking descriptions locally, applying a similarity threshold, and injecting only bounded results preserves context for the task itself. The current discovery limits are three skills and five tool hints per request.

Keep discovery separate from execution

The embedding index suggests relevant capabilities. The skill loader and toolsrv remain authoritative for content, metadata, loading, execution, sandboxing, and process state. Separating ranking from execution keeps the system replaceable and prevents a failed or stale index from changing the security boundary.

Local models improve privacy and small-model operation

A local embedding model avoids sending discovery data to the conversational provider and works with offline or inexpensive backends. The embedding model and conversational model can evolve independently.

Lazy tool loading is a dead end

Phase 36 attempted a one-tool bootstrap: agents started with only client_9p and loaded other tools on demand through semantic hints. The theory was that this would create minimal capability surfaces that expanded only as needed.

In practice, this failed for three reasons:

  1. Models don't follow multi-step loading protocols. When a model wants to read a file, it calls file_read. It doesn't first call client_9p to load file_read, then call file_read. The indirection step is almost always skipped regardless of prompt instructions.

  2. The indirection doubles round-trips. Even when the model follows the protocol, every tool use requires two generation cycles: one to decide to load, one to use. This halves throughput for common operations.

  3. Capability boundaries become implicit. If any agent can load any tool, the permission model depends on what the model decides to do, not what it's configured to do. A code reviewer could load shell and execute arbitrary commands. The boundary is only as strong as the model's compliance.

The fix was explicit autoLoad declarations in agent configs. Each profile lists exactly which tools it starts with. Role-specific profiles get role- appropriate tools. The capability surface is explicit, auditable, and enforced at config time rather than generation time.

Lesson: Don't rely on model compliance for capability boundaries. If a tool shouldn't be available, don't make it loadable.

Model compliance is not a security boundary

The lazy loading failure (above) reveals a deeper principle: behavioral defenses are not structural guarantees. Security properties that depend on the model following instructions are only as strong as the model's compliance.

This distinction matters for agent security:

Defense Type Mechanism Failure Mode
Behavioral Prompt instructions, model refusal Model ignores/misinterprets instructions
Structural Config-time enforcement, OS-level sandbox Requires implementation bug to bypass

Ollie's security model uses structural defenses:

  1. Tool access: Per-agent registry populated at config time via autoLoad. An agent cannot call tools not in its registry — Lookup() fails before any tool code runs. This is structural: the model cannot comply its way past a missing registry entry.

  2. Execution sandbox: Landlock kernel enforcement. Even if an agent calls shell with a malicious command, the kernel blocks filesystem access outside allowed paths. This is structural: the model cannot prompt-inject its way past a syscall filter.

  3. Bypass broker: Escape from sandbox requires explicit user approval. The model can request bypass, but cannot grant it to itself.

The empirical finding that models skip loading protocols is one instance of a general pattern: models vary in whether they follow behavioral expectations. Some models (Claude) refuse malicious instructions reliably. Others (GPT-4o) attempt them. A security model that depends on refusal breaks when the model doesn't refuse.

Lesson: Structural enforcement (registry checks, kernel sandboxing) holds regardless of model behavior. Behavioral enforcement (prompt instructions, model refusal) is defense-in-depth, not a primary security boundary.