doc: update architecture docs for explicit tool autoLoad

- architecture-embedding.md: skills.Index → generic embedding.Index[T],
  tool index now per-turn from loaded tools, split source map entries
- architecture-tools.md: explain explicit autoLoad requirement, remove
  lazy loading mention
- evolution.md: add Phase 37 (lazy loading reversal, generic index,
  agent config alignment), add 4 dead-end entries
- lessons-learned.md: add 'Lazy tool loading is a dead end' section
- README.md: update line 88 to reflect explicit autoLoad

Reflects the removal of load-on-call tool loading (Phase 36 reversal)
and the separation of embedding/index.go as a generic type.
This commit is contained in:
Levi Neely 2026-08-21 10:43:58 +02:00
parent 6b5d3fd0b6
commit e0ac54c0fb
5 changed files with 178 additions and 27 deletions

View File

@ -85,7 +85,7 @@ session/
`session/new` accepts `name`, `remote`, `workflow`, `cwd`, and `yolo=true` key/value fields. Session-level `yolo` is persisted and applies to local or remote toolsrv startup; it disables native Landlock enforcement for explicit development use. Normal tools use the configured native Landlock policy, while approved escape requests go through the bypass broker.
Provider backends are the model-facing boundary; the agent loop is independent of vendor APIs. Tools are external executable programs served by the separate `toolsrv` 9P process. Their metadata is discovered and loaded lazily, so the model sees only the capabilities needed for a task.
Provider backends are the model-facing boundary; the agent loop is independent of vendor APIs. Tools are external executable programs served by the separate `toolsrv` 9P process. Each agent profile declares its tools via `autoLoad`; the model sees only the capabilities configured for its role.
Tool execution uses the configured native Landlock sandbox. Operations requiring an approved escape use the bypass namespace and broker, which applies policy and approval controls rather than providing an unrestricted escape.

View File

@ -49,23 +49,23 @@ ONNX Runtime initialization is process-global and occurs once. Model sessions
are mutex-protected because inference uses shared session state. Each index
owns its model session and releases it through `Close`.
### `skills.Index`
### `embedding.Index[T]`
The shared index type handles both skills and tools:
The generic index type handles arbitrary entry types:
```text
Index {
model: embedding.Model
entries: []Skill
Index[T] {
model: *Model
entries: []T
vectors: []Vector
textFn: func(T) string // extracts text for embedding
}
```
For skills, the index scans configured directories for `<name>/SKILL.md`,
parses `name` and `description` from YAML frontmatter, and retains the full
content for injection. For tools, the agent converts `.meta` descriptions into
the same `Skill` representation; only the name and description are needed for
ranking.
Skill and tool indexes use the same generic structure with different entry
types. For skills, entries are `skills.Skill` structs with name, description,
and content. For tools, entries are `toolsrv.Meta` structs from loaded tool
metadata.
Description vectors are precomputed when an index is built or reloaded. A user
request is embedded at match time, then compared against all indexed vectors.
@ -74,26 +74,26 @@ cosine score and capped by the caller.
## Agent integration
The agent keeps independent cached indexes for skills and tools. Initialization
is lazy and guarded by `sync.Once`:
The agent keeps independent cached indexes for skills and tools. Skill index
initialization is lazy and guarded by `sync.Once`. Tool index is rebuilt per
turn from the agent's currently loaded tools:
```text
first turn
├── build skill index from configured SKILL.md directories
└── build tool index from installed .meta files
└── build skill index from configured SKILL.md directories
each turn
├── build tool index from agent's loaded tools
├── embed request
├── match skills: threshold 0.35, limit 3
└── match tools: threshold 0.35, limit 5
```
Relevant skills are injected as a `<skills>` block containing complete skill
content. Relevant tools are injected as a `<tool-hints>` block containing names and
descriptions plus the `client_9p` operation needed to load them. Semantic
matching does not grant capability: the normal per-agent registry still
controls whether a tool is loaded and executable. The model must load a hinted
tool through the agent `ctl` file before calling it.
content. Relevant tools are injected as a `<tool-hints>` block containing
loaded tool names and descriptions. Tools must be explicitly loaded through
the agent's `autoLoad` config or via the `ctl` file before they appear in
matching results.
The bounded result limits prevent discovery from becoming a prompt-size or
context-budget problem. This progressive disclosure is especially useful for
@ -149,9 +149,11 @@ they affect ranking.
- `embedding/embedding.go` — model loading, tokenization, inference, pooling,
normalization, and similarity.
- `skills/skills.go` — skill scanning, frontmatter parsing, indexing, matching,
and reload.
- `cmd/olliesrv/internal/agent/skill_match.go` — lazy indexes and prompt
injection for skills and tool hints.
- `embedding/index.go` — generic `Index[T]` for vector-indexed entries.
- `skills/skills.go` — skill scanning, frontmatter parsing, and `Skill` type.
- `cmd/olliesrv/internal/agent/skill_match.go` — lazy skill index and prompt
injection for skills.
- `cmd/olliesrv/internal/agent/tool_match.go` — per-turn tool index from loaded
tools and prompt injection for tool hints.
- `Makefile` — model and runtime asset installation.
- [`architecture-tools.md`](architecture-tools.md) — tool metadata and execution.

View File

@ -184,9 +184,17 @@ A metadata-only tool needs only the first command. Test its command contract wit
echo '{"pattern":"TODO"}' | bash -c 'YOUR_CMD'
```
Agent profiles may load tools at startup with `autoLoad`, but the default profile
uses an empty list and starts with only `client_9p`. Semantic tool hints tell the
agent to load additional tools through its own `ctl` file:
Agent profiles load tools at startup with `autoLoad`. Each profile specifies
the tools it needs — there is no lazy loading or automatic tool discovery at
call time:
```json
{
"autoLoad": ["shell", "file_read", "file_edit", "file_grep", "file_glob"]
}
```
To add a tool to a running agent, use the `ctl` file:
```text
client_9p(op="rdwr", path="session/$OLLIE_SESSION_ID/agent/$OLLIE_UNAME/ctl", data="tool_load <name>")

View File

@ -1521,6 +1521,10 @@ Everything that was built and then killed, in roughly chronological order.
| `cascade` orchestrator script | Replaced by parallel `subagent_spawn` calls (scope: read, natural parallelism) |
| Agent-loop batching as correctness mechanism | Moved to toolsrv path-lock table; agent batching is now just an optimization |
| `virtfs.Request` naming | Renamed to `Rdwr` — it's an atomic operation, not a variant of read or write |
| One-tool bootstrap (Phase 36) | Models skip the load step and call tools directly; explicit `autoLoad` per profile replaced it |
| Lazy tool loading | Capability boundaries must be explicit; model compliance is not a security boundary |
| `client_9p` tool hints with load instructions | Removed — tool hints now show loaded tools only |
| `skills.Index` shared for skills and tools | Replaced by generic `embedding.Index[T]` with separate skill_match.go and tool_match.go |
## Feed file + Observer agents + BlockOnce/Stream refactor (Aug 13)
@ -1995,3 +1999,110 @@ The `client_9p` script also expands `$OLLIE_SESSION_ID` and `$OLLIE_UNAME` in
virtual namespace paths. This is required because JSON argument parsing does not
perform shell expansion, and keeps the documented namespace examples directly
callable by the model.
## Phase 37: Explicit Tool Loading and Generic Embedding Index (Aug 21)
Lazy tool loading — where the agent could call any tool by name and have it
loaded automatically — was removed. Tools now require explicit `autoLoad`
declarations in agent configs. This is a reversal of the Phase 36 progressive
loading experiment.
### Why lazy loading failed
The one-tool bootstrap approach (Phase 36) relied on semantic hints telling the
model to load tools through `client_9p`. In practice:
1. **Models called tools directly instead of loading them first.** Even with
explicit instructions, models would attempt to call `file_read` or `shell`
without the intermediate `client_9p tool_load` step.
2. **The indirection added latency and context cost.** Every tool use required
an extra round-trip: hint → model decides to load → load call → refresh →
actual tool call. This doubled the turns for common operations.
3. **Capability boundaries became unclear.** An agent's effective capability
was "whatever it decides to load," which is not the same as "what it's
configured to do." A code-review agent shouldn't have `shell` access just
because it asked for it.
### The fix: explicit autoLoad per profile
Each agent profile now declares exactly which tools it starts with:
```json
{
"name": "default",
"autoLoad": ["shell", "file_read", "file_edit", "file_grep", "file_glob",
"file_write", "lsp_hover", "lsp_definition", "lsp_references",
"lsp_symbols", "lsp_diagnostics", "lsp_completion", "lsp_rename",
"code_outline", "code_symbols", "code_query", "codebase_overview",
"code_dependencies", "code_rewrite", "client_9p", "subagent_spawn",
"skill_list", "skill_load", "memory_wake", "memory_remember",
"memory_recall", "memory_zoom", "web_fetch", "reasoning_think"]
}
```
Role-specific profiles (conductor, reviewer, panelist, etc.) load only what
they need. An observer loads read-only tools. A foreman loads coordination
tools. This makes capabilities explicit and auditable.
### Generic embedding index
The embedding package gained a generic `Index[T]` type that replaces the
skills-specific `skills.Index`. The same index structure now handles both
skill matching (entries are `skills.Skill`) and tool matching (entries are
`toolsrv.Meta`).
```go
type Index[T any] struct {
model *Model
entries []T
vectors []Vector
textFn func(T) string
}
```
Tool matching was split from skill matching:
- `skill_match.go` — builds skill index once per process, matches per turn
- `tool_match.go` — builds tool index per turn from loaded tools only
The tool index now reflects the agent's actual loaded tools, not all installed
metadata. This aligns semantic discovery with the explicit loading model.
### Agent config alignment
All 14 agent configs were updated to declare tools appropriate to their roles:
| Profile | Purpose | Key tools |
|---------|---------|-----------|
| default | General coding | Full tool set |
| conductor | Task decomposition | Code intel + subagent_spawn |
| reviewer | Code review | Read-only + LSP |
| foreman | Consensus synthesis | Coordination only |
| panelist | Independent analysis | Read + reasoning |
| observer | Watch and comment | Read-only |
| driver | Remote execution | Shell + file tools |
| copilot | IDE assistance | Code intel + LSP |
| writer | Documentation | File tools + reasoning |
| researcher | Investigation | Read + web + reasoning |
| planner | Architecture | Read + reasoning |
| debugger | Troubleshooting | Full tool set |
| tester | Test writing | Code + shell |
| refactorer | Code transformation | Code + LSP + rewrite |
### What was removed
- `load-on-call` tool loading (the Phase 36 mechanism)
- `client_9p` load hints in tool-hints injection
- `skills.Index` (replaced by generic `embedding.Index[T]`)
- Combined skill/tool matching in `skill_match.go`
### Source changes
```text
embedding/index.go +89 new generic Index[T]
skills/skills.go -70 removed Index, kept Skill type
skill_match.go -40 removed tool matching
tool_match.go +55 new per-turn tool index
data/agents/*.json +14 autoLoad declarations
```

View File

@ -61,3 +61,33 @@ security boundary.
A local embedding model avoids sending discovery data to the conversational
provider and works with offline or inexpensive backends. The embedding model
and conversational model can evolve independently.
## Lazy tool loading is a dead end
Phase 36 attempted a one-tool bootstrap: agents started with only `client_9p`
and loaded other tools on demand through semantic hints. The theory was that
this would create minimal capability surfaces that expanded only as needed.
In practice, this failed for three reasons:
1. **Models don't follow multi-step loading protocols.** When a model wants to
read a file, it calls `file_read`. It doesn't first call `client_9p` to load
`file_read`, then call `file_read`. The indirection step is almost always
skipped regardless of prompt instructions.
2. **The indirection doubles round-trips.** Even when the model follows the
protocol, every tool use requires two generation cycles: one to decide to
load, one to use. This halves throughput for common operations.
3. **Capability boundaries become implicit.** If any agent can load any tool,
the permission model depends on what the model decides to do, not what it's
configured to do. A code reviewer could load `shell` and execute arbitrary
commands. The boundary is only as strong as the model's compliance.
The fix was explicit `autoLoad` declarations in agent configs. Each profile
lists exactly which tools it starts with. Role-specific profiles get role-
appropriate tools. The capability surface is explicit, auditable, and
enforced at config time rather than generation time.
**Lesson:** Don't rely on model compliance for capability boundaries. If
a tool shouldn't be available, don't make it loadable.