docs: update ARCHITECTURE and PLANNING for new tool structure
Reflect the removal of PlanBackend/MemoryBackend abstractions and the addition of pkg/tools/memory and pkg/tools/task. Update tool table (12 tools, 5 servers), data flow diagram, planning lifecycle, and consumer setup example. Rewrite PLANNING.md around task_plan/task_complete. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
parent
f20c62dd3d
commit
ed4fe7e6f2
|
|
@ -12,9 +12,12 @@ pkg/
|
|||
backend/ — Backend interface + implementations
|
||||
config/ — Agent config struct and loader
|
||||
mcp/ — MCP client (concrete)
|
||||
tools/ — Server and Dispatcher interfaces, tool definitions (builtin.go)
|
||||
tools/ — Server and Dispatcher interfaces
|
||||
tools/execute/ — execute.Server: execute_code, execute_tool, execute_pipe
|
||||
tools/reasoning/ — reasoning.Server: reasoning_think, reasoning_plan
|
||||
tools/file/ — file.Server: file_read, file_write, file_edit, file_glob, file_grep
|
||||
tools/memory/ — memory.Server: memory_remember, memory_recall
|
||||
tools/reasoning/ — reasoning.Server: reasoning_think
|
||||
tools/task/ — task.Server: task_plan, task_complete
|
||||
|
||||
internal/
|
||||
sandbox/ — landrun sandbox config and command wrapper
|
||||
|
|
@ -35,21 +38,20 @@ Consumers extend ollie by implementing or composing its interfaces:
|
|||
- **`backend.Backend`** — swap or add LLM backends
|
||||
- **`tools.Server`** — add a new tool server (built-in or MCP-backed); all servers are equal
|
||||
- **`tools.Dispatcher`** — replace the tool router entirely (e.g. remote dispatcher, mock)
|
||||
- **task backend (MCP)** — any MCP server that exposes `task_create` is automatically wired as the persistence backend for `reasoning_plan` by `BuildAgentEnv`; no consumer code required. See [9beads-mcp](https://github.com/lneely/9beads-mcp) for the reference implementation and the interface contract.
|
||||
- **fallback plan backend** — consumers can supply a `tools.PlanBackend` via `agent.WithFallbackPlanBackend` that is used when no `task_create` MCP tool is found. ollie-9p uses this to enqueue plan steps into the session queue.
|
||||
- **`agent.Core`** — the agent's public API; frontends drive it without knowing internals
|
||||
|
||||
All tool servers implement the same `tools.Server` interface regardless of whether they are built-in or backed by MCP. There is no special "builtin" concept — `execute.Server` and `file.Server` are registered by name the same way MCP servers are, and are torn down and recreated on `/agent` switches just like MCP connections.
|
||||
|
||||
Each built-in server package exports a `Decl` function (`func Decl(...) func() tools.Server`) — a parameterized factory that produces a fresh server instance. `tools.NewDispatcherFunc` takes a map of name→Decl result and returns a `func() tools.Dispatcher` suitable for `agent.AgentCoreConfig.NewDispatcher`. `tools.NewServer(client)` wraps an `mcp.Client` as a `tools.Server`.
|
||||
|
||||
`execute.Decl(cwd string)` accepts a working directory that is set as `cmd.Dir` for sandboxed commands and used to expand `{CWD}` in the sandbox config. Pass `""` to fall back to `os.Getwd()`.
|
||||
`execute.Decl(cwd string)` and `file.Decl(cwd string)` accept a working directory. Pass `""` to fall back to `os.Getwd()`.
|
||||
|
||||
`task.Decl(planDir string)` accepts the directory where plan files are written.
|
||||
|
||||
**Adding a new tool server:**
|
||||
1. Implement `tools.Server` in a new package under `pkg/tools/`
|
||||
2. Export `func Decl(...) func() tools.Server`
|
||||
3. Add tool definitions to `pkg/tools/builtin.go` (avoids import cycles)
|
||||
4. Register the Decl result by name in `tools.NewDispatcherFunc` — no frontend changes needed
|
||||
3. Register the Decl result by name in `tools.NewDispatcherFunc` — no frontend changes needed
|
||||
|
||||
## Data Flow
|
||||
|
||||
|
|
@ -61,70 +63,65 @@ Frontend
|
|||
└── tools.Dispatcher.Dispatch(...) — tool dispatch
|
||||
├── "execute" → tools.Server
|
||||
│ execute_code / execute_tool / execute_pipe
|
||||
├── "file" → tools.Server
|
||||
│ file_read / file_write / file_edit / file_glob / file_grep
|
||||
├── "memory" → tools.Server
|
||||
│ memory_remember / memory_recall
|
||||
├── "reasoning" → tools.Server
|
||||
│ reasoning_think
|
||||
│ reasoning_plan
|
||||
│ └── tools.PlanBackend (optional, priority order)
|
||||
│ 1. dispatchPlanBackend → task_create (MCP)
|
||||
│ 2. WithFallbackPlanBackend (e.g. queuePlanBackend)
|
||||
│ 3. nil → in-context plan only
|
||||
└── "<servername>" → tools.Server
|
||||
├── "task" → tools.Server
|
||||
│ task_plan / task_complete
|
||||
└── "<servername>" → tools.Server (MCP)
|
||||
```
|
||||
|
||||
All tool calls go through one `tools.Dispatcher.Dispatch` path. Every registered server is a `tools.Server`; the dispatcher routes by name and is agnostic to how any server is implemented.
|
||||
|
||||
## Built-in Tools
|
||||
|
||||
Five tools are registered across two servers:
|
||||
Twelve tools across five servers:
|
||||
|
||||
| Server | Tool | What it does |
|
||||
|---|---|---|
|
||||
| `execute` | `execute_code` | Runs inline bash in a landrun sandbox |
|
||||
| `execute` | `execute_tool` | Reads a named script from `OLLIE_TOOLS_PATH` and runs it sandboxed |
|
||||
| `execute` | `execute_pipe` | Chains steps, piping stdout of each into stdin of the next |
|
||||
| `file` | `file_read` | Reads a file from the filesystem |
|
||||
| `file` | `file_write` | Writes a file to the filesystem |
|
||||
| `file` | `file_edit` | Edits a file with precise string replacement |
|
||||
| `file` | `file_glob` | Finds files by glob pattern |
|
||||
| `file` | `file_grep` | Searches file contents by regex |
|
||||
| `memory` | `memory_remember` | Saves a persistent memory to a flat markdown file |
|
||||
| `memory` | `memory_recall` | Searches saved memories by keyword |
|
||||
| `reasoning` | `reasoning_think` | Externalizes intermediate reasoning (no-op, recorded in history) |
|
||||
| `reasoning` | `reasoning_plan` | Decomposes a goal into ordered steps; persists to task backend if available |
|
||||
|
||||
File operations go through `execute_code` using standard shell tools (`cat`, `grep`, `sed`, `ed`, `ssam` if plan9port is available, etc.).
|
||||
|
||||
Tool definitions live in `pkg/tools/builtin.go` to avoid import cycles — the subpackages import `pkg/tools`, so `pkg/tools` cannot import them back.
|
||||
| `task` | `task_plan` | Decomposes a goal into ordered steps; saves a markdown checklist to pl/ |
|
||||
| `task` | `task_complete` | Marks a plan step complete; returns remaining steps |
|
||||
|
||||
`OLLIE_TOOLS_PATH` defaults to `~/.config/ollie/tools`. The directory can be a symlink or a mountpoint — `execute_tool` treats it as an ordinary filesystem path.
|
||||
|
||||
## Planning and Task Persistence
|
||||
`OLLIE_MEMORY_PATH` defaults to `~/.local/share/ollie/memory/`. Memory files follow a denote-like naming convention: `YYYYMMDDTHHmmss--title-slug__tag1_tag2.md`.
|
||||
|
||||
`reasoning_plan` is a meta-cognitive tool for executive planning. It decomposes a goal into a dependency graph of steps before execution begins. If a task backend is available, the plan is committed to persistent storage.
|
||||
## Planning
|
||||
|
||||
The task backend is any MCP server that exposes a `task_create` tool. `BuildAgentEnv` scans available tools after connecting MCP servers: if `task_create` is found, it wires a `dispatchPlanBackend` to the reasoning server's `Plan` field via the `tools.PlanBackendSetter` interface. If not found, it falls back to any `PlanBackend` supplied via `WithFallbackPlanBackend`. If neither is present, `reasoning_plan` produces an in-context plan only.
|
||||
`task_plan` saves a markdown checklist to the plan directory (`pl/` in ollie-9p) as a `__todo.md` file. The agent then:
|
||||
|
||||
ollie-9p supplies a `queuePlanBackend` fallback for every session: when `task_create` is absent, plan steps are enqueued to the session's `enqueue` file in topological order and returned as placeholder IDs (`q1`, `q2`, …). This implementation lives in `9p` because it has a hard dependency on the 9P filesystem layout; the extension point itself (`tools.PlanBackend` + `WithFallbackPlanBackend`) remains in core.
|
||||
1. Renames `__todo` → `__wip` when work begins (`file_edit` or shell)
|
||||
2. Calls `task_complete` after each step finishes — marks `[ ]` → `[x]`, returns remaining steps
|
||||
3. Uses `file_*` tools to add, remove, or revise steps as the plan evolves
|
||||
4. Renames `__wip` → `__done` when the goal is realized
|
||||
|
||||
The reference task backend is [9beads-mcp](https://github.com/lneely/9beads-mcp), which wraps the [9beads](https://github.com/lneely/9beads) 9P task server.
|
||||
`task_complete` takes the plan filename and a step title (substring match, case-insensitive). It errors if no matching unchecked step is found, so silent skips are caught.
|
||||
|
||||
### Task interface contract
|
||||
|
||||
Auto-wiring requires only `task_create`. The minimum viable contract is:
|
||||
|
||||
| Tool | Required | Purpose |
|
||||
|---|---|---|
|
||||
| `task_create` | yes | Create a task; must return a plain-text ID |
|
||||
| `task_delete` | recommended | Remove tasks when a plan is aborted or superseded |
|
||||
| `task_list` | optional | Orient the agent across sessions |
|
||||
| `task_read` | optional | Inspect a specific task |
|
||||
| `task_update` | optional | Claim, complete, fail, defer, label tasks |
|
||||
| `task_edit` | optional | Revise title, body, or parent of an existing task |
|
||||
| `task_dep` | optional | Add or remove blocking dependencies between tasks |
|
||||
|
||||
Ollie degrades gracefully: if a tool is absent, the agent falls back to `execute_code` (shell-out) or skips that lifecycle step. The only hard requirement for persistent planning is `task_create` returning an ID in the `{"content": [{"type": "text", "text": "<id>"}]}` format.
|
||||
|
||||
See [PLANNING.md](PLANNING.md) for the full design rationale.
|
||||
Everything else — listing plans, inspecting a specific step, bulk operations — goes through `file_*` or `execute_code`. The task server handles only creation and step completion.
|
||||
|
||||
## Typical Consumer Setup
|
||||
|
||||
```go
|
||||
newDispatcher := tools.NewDispatcherFunc(map[string]func() tools.Server{
|
||||
"execute": execute.Decl(cwd), // "" falls back to os.Getwd()
|
||||
"file": file.Decl(cwd),
|
||||
"memory": memory.Decl(),
|
||||
"reasoning": reasoning.Decl(),
|
||||
"task": task.Decl(planDir),
|
||||
})
|
||||
|
||||
env := agent.BuildAgentEnv(cfg, newDispatcher(), cwd) // also connects MCP servers from cfg
|
||||
|
|
@ -140,8 +137,6 @@ core := agent.NewAgentCore(agent.AgentCoreConfig{
|
|||
|
||||
`BuildAgentEnv` adds MCP servers from the config on top of the pre-registered servers. On `/agent` switches, `NewDispatcher` is called to produce a fresh dispatcher — all servers (built-in and MCP) are torn down and recreated for the new agent config. `CWD` is preserved across switches.
|
||||
|
||||
After connecting MCP servers, `BuildAgentEnv` scans the tool list for `task_create`. If found, it wires a `dispatchPlanBackend` to the reasoning server's `Plan` field via `tools.PlanBackendSetter`. If not found, it wires any fallback passed via `WithFallbackPlanBackend`. This auto-wiring runs on every agent start and `/agent` switch.
|
||||
|
||||
`Core.ListServers()` returns all registered tool servers and their tools, grouped by server name. Accessible via the `/mcp` command or `ollie/s/{sid}/mcp` in ollie-9p.
|
||||
|
||||
## Session and Context
|
||||
|
|
|
|||
179
doc/PLANNING.md
179
doc/PLANNING.md
|
|
@ -1,170 +1,55 @@
|
|||
# Planning and Task Persistence in Ollie
|
||||
# Planning in Ollie
|
||||
|
||||
## Problem
|
||||
## Design
|
||||
|
||||
An agent without planning is less effective but still useful. Planning capability
|
||||
should be optional — degraded but functional when no task backend is available,
|
||||
and persistent when one is.
|
||||
Planning is executive functioning — breaking a goal into ordered steps before acting. It lives in the `task` server alongside `task_complete`, which closes the loop by marking steps done as work progresses.
|
||||
|
||||
The challenge: how to integrate a planning tool with an optional external task
|
||||
backend without hard-coupling the core to any specific implementation.
|
||||
Two tools handle the full planning lifecycle:
|
||||
|
||||
## Design Decisions
|
||||
| Tool | When to call |
|
||||
|---|---|
|
||||
| `task_plan` | Before beginning a multi-step task |
|
||||
| `task_complete` | Immediately after each step finishes |
|
||||
|
||||
### reasoning_plan as a built-in tool
|
||||
Everything else — adding steps, editing descriptions, renaming the plan file — goes through the standard `file_*` tools. The task server handles creation and completion only.
|
||||
|
||||
Planning is executive functioning — breaking a goal into ordered steps before
|
||||
acting. It belongs in the `reasoning_*` namespace alongside `reasoning_think`,
|
||||
which handles moment-to-moment reflection. Both are meta-cognitive tools:
|
||||
`reasoning_think` externalizes intermediate thought; `reasoning_plan`
|
||||
externalizes the work breakdown.
|
||||
## task_plan
|
||||
|
||||
Tool schemas are more reliable than free-text instructions over long sessions.
|
||||
Agents rarely forget how to call `file_read`; they can forget shell conventions.
|
||||
The planning operation is complex enough (goal + structured steps + dependency
|
||||
edges) that a typed schema earns its place.
|
||||
`task_plan` takes a goal and an ordered list of steps (each with an optional description and `depends_on` index list). It:
|
||||
|
||||
`reasoning_plan` does not rely on shell-out. It takes structured JSON, formats
|
||||
the plan as readable text, and optionally persists it. This is the right
|
||||
minimum for a built-in tool.
|
||||
1. Topologically sorts the steps (blockers before dependents)
|
||||
2. Writes a markdown checklist to the plan directory as `YYYYMMDDTHHmmss_{uid}--{goal-slug}__todo.md`
|
||||
3. Returns the filename and step count
|
||||
|
||||
### Loose coupling via PlanBackend interface
|
||||
The returned filename is what the agent passes to `task_complete` and any `file_*` calls on the plan.
|
||||
|
||||
The reasoning server holds an optional `Plan tools.PlanBackend` field (nil by
|
||||
default). When nil, `reasoning_plan` produces an in-context plan only. When set,
|
||||
the plan is persisted (or queued) by the backend.
|
||||
### Plan lifecycle
|
||||
|
||||
`tools.PlanBackend` and `tools.PlanBackendSetter` are defined in `pkg/tools`.
|
||||
The reasoning server implements `PlanBackendSetter`. The agent package implements
|
||||
`dispatchPlanBackend`, which routes task creation through the dispatcher. No
|
||||
import cycles: reasoning → tools, agent → tools, agent ↛ reasoning.
|
||||
|
||||
Consumers can supply a fallback backend via `agent.WithFallbackPlanBackend(b)`:
|
||||
|
||||
```go
|
||||
env := agent.BuildAgentEnv(cfg, d, cwd, agent.WithFallbackPlanBackend(myFallback))
|
||||
```
|
||||
task_plan → __todo.md created
|
||||
file_edit (rename __todo → __wip) → work begins
|
||||
task_complete (per step) → [ ] → [x], remaining steps returned
|
||||
file_edit (rename __wip → __done) → goal realized
|
||||
```
|
||||
|
||||
The fallback is only used when no `task_create` MCP tool is found. If a task
|
||||
backend is available it takes priority unconditionally.
|
||||
Renaming is intentional: `__todo` / `__wip` / `__done` is a visible filesystem-level state machine. Plans survive crashes and context compaction because they live on disk, not in the conversation.
|
||||
|
||||
### Task backend as an MCP server
|
||||
### Re-planning
|
||||
|
||||
The task backend is any MCP server that exposes `task_create`. There is no
|
||||
built-in task server in ollie core — this is an intentional extension point.
|
||||
The convention (the interface contract) is:
|
||||
If a step fails or new information changes the approach, call `task_plan` again with the revised goal and steps. Abandon the old plan file (rename it `__abandoned` or delete it) and proceed with the new one.
|
||||
|
||||
| Tool | Required | Description |
|
||||
|---|---|---|
|
||||
| `task_create` | yes | Create a task; return its ID as plain text |
|
||||
| `task_delete` | recommended | Remove tasks when a plan is aborted or superseded |
|
||||
| `task_list` | no | List tasks by status |
|
||||
| `task_read` | no | Read a task by ID |
|
||||
| `task_update` | no | Update task status/metadata |
|
||||
| `task_edit` | no | Revise title, body, or parent of an existing task |
|
||||
| `task_dep` | no | Add or remove blocking dependencies between tasks |
|
||||
## task_complete
|
||||
|
||||
`task_create` must return a plain-text ID in the standard MCP content format:
|
||||
`{"content": [{"type": "text", "text": "<id>"}]}`.
|
||||
`task_complete` takes the plan filename and a step title (substring, case-insensitive). It finds the first unchecked step whose title contains the substring, flips `[ ]` to `[x]`, writes the file back, and returns the remaining unchecked step titles.
|
||||
|
||||
Any MCP server implementing this convention will be auto-detected and wired.
|
||||
Returning the remaining steps serves two purposes: the agent always knows what's left without a separate file read, and the explicit list nudges continuation rather than premature termination.
|
||||
|
||||
### Auto-wiring in BuildAgentEnv
|
||||
`task_complete` errors if no matching unchecked step is found. This prevents silent skips — the agent must match a real step or handle the error explicitly.
|
||||
|
||||
After connecting MCP servers, `BuildAgentEnv` scans the tool list for
|
||||
`task_create`. If found, it constructs a `dispatchPlanBackend` and sets it on
|
||||
the reasoning server via `PlanBackendSetter`. This wiring runs on every agent
|
||||
start and `/agent` switch, so the task backend is always current.
|
||||
### Discipline
|
||||
|
||||
If `task_create` is not found, `BuildAgentEnv` checks whether the caller
|
||||
supplied a fallback via `WithFallbackPlanBackend`. If so, that backend is wired
|
||||
instead. Priority: task MCP > caller fallback > nil (in-context only).
|
||||
Call `task_complete` immediately after each step finishes, before starting the next. This keeps the plan file current throughout execution and ensures that if the session is interrupted, the state on disk reflects actual progress.
|
||||
|
||||
No frontend changes are needed to gain planning capability — add a task MCP
|
||||
server to the agent config and it works.
|
||||
## Memory
|
||||
|
||||
### Graceful degradation
|
||||
|
||||
- No task MCP server configured, no fallback → `reasoning_plan` produces an
|
||||
in-context plan only.
|
||||
- No task MCP server configured, fallback supplied → steps are handed to the
|
||||
fallback backend (e.g. queued for sequential execution). See
|
||||
[Queue-based fallback](#queue-based-fallback) below.
|
||||
- Task MCP server configured but unreachable → `task_create` fails,
|
||||
`CreatePlan` returns an error, reasoning server degrades to in-context plan
|
||||
with a warning.
|
||||
- Task MCP server running → full persistence, steps get IDs, plan is durable.
|
||||
|
||||
The agent's behavior is identical in all cases: call `reasoning_plan`, get a
|
||||
formatted plan, proceed. The persistence difference is transparent.
|
||||
|
||||
### Queue-based fallback
|
||||
|
||||
ollie-9p registers a `queuePlanBackend` as the fallback for every session. When
|
||||
`task_create` is absent, `reasoning_plan` writes each step as a queued prompt to
|
||||
the session's `enqueue` file in topological order (blockers before dependents).
|
||||
The implementation lives in `9p` (not core) because it has a hard dependency on
|
||||
the 9P filesystem layout — the extension point is the interface, not this
|
||||
specific implementation.
|
||||
|
||||
Steps are returned as placeholder IDs (`q1`, `q2`, …) so the agent can refer to
|
||||
them in subsequent `reasoning_think` calls. The queue persists independently of
|
||||
context — unlike an in-context-only plan, steps survive compaction.
|
||||
|
||||
The goal description is prepended to the first enqueued step for context.
|
||||
Dependency annotations (`after: q1, q2`) are included in each step's prompt so
|
||||
the agent retains the relationship even if it processes steps across multiple
|
||||
turns.
|
||||
|
||||
Fan-out parallelism is achievable by spawning sub-sessions: the enqueued step
|
||||
can instruct the agent to create a child session per parallel branch, nudge each
|
||||
with its task via `prompt`, and have sub-agents write completion notifications
|
||||
back to the parent via `enqueue`.
|
||||
|
||||
### Shell-out for everything else
|
||||
|
||||
Queries, comments, bulk operations, event watching — all of these belong in
|
||||
`execute_code`, not in built-in tools. The `task_*` MCP tools cover the full
|
||||
planning and execution lifecycle (create, list, read, update, edit, delete,
|
||||
dependency management). For example, advanced 9beads operations are accessible
|
||||
via shell: `cat $TASK_DIR/list`, `grep`, etc.
|
||||
|
||||
This keeps the built-in surface minimal and relies on `execute_code` for
|
||||
flexibility.
|
||||
|
||||
## Reference Implementation: 9beads-mcp
|
||||
|
||||
[9beads-mcp](https://github.com/lneely/9beads-mcp) wraps the 9beads 9P task
|
||||
server as an MCP server. It:
|
||||
|
||||
- Resolves the project mount from `$PWD` (passed explicitly in the MCP server
|
||||
env config, since MCP subprocesses run in a minimal environment)
|
||||
- Auto-mounts the project directory if not already mounted
|
||||
- Exposes `task_create`, `task_list`, `task_read`, `task_update`, `task_edit`, `task_delete`, `task_dep`
|
||||
|
||||
Agent config example:
|
||||
|
||||
```yaml
|
||||
mcpServers:
|
||||
task:
|
||||
command: 9beads-mcp
|
||||
env:
|
||||
PWD: "$PWD"
|
||||
```
|
||||
|
||||
`$PWD` is expanded by `os.ExpandEnv` at connect time, giving the MCP server
|
||||
the agent's working directory.
|
||||
|
||||
## Future Direction: Event-Driven Planning
|
||||
|
||||
9beads exposes `~/mnt/beads/events` as a blocking JSON event stream. A
|
||||
goroutine in the frontend could watch this and inject events into the agent's
|
||||
interrupt queue (via `PromptFIFO`) when:
|
||||
|
||||
- A blocked step becomes unblocked (its dependency completed)
|
||||
- A step is assigned to this agent by an external actor
|
||||
- An external process marks a step complete
|
||||
|
||||
This would make agents reactive rather than polling — the agent yields after
|
||||
completing work and wakes up when the event stream fires. Not implemented yet;
|
||||
documented here as the natural next step.
|
||||
Memory is handled separately by the `memory` server (`memory_remember` / `memory_recall`). It is not part of the planning system. Memory persists facts across sessions; plans persist work-in-progress within a task.
|
||||
|
|
|
|||
Loading…
Reference in New Issue