- Delete agent/complete.go (110 lines)
- The kate plugin builds the completion prompt inline and calls
ollie-9p rdwr generate instead of ollie-9p rdwr complete
- Remove defaultCompactionModels map from compaction.go
- Add agent/models.go: load ~/.config/ollie/models.yaml for all backend/model settings
- Completion (Complete()) falls back to models.yaml completion section when env vars unset
- Compaction falls back to models.yaml compaction section per-backend
- Priority chain: agent config > OLLIE_*_MODEL env > models.yaml > session model
- Remove all OLLIE_*_PATH vars (TOOLS_PATH, CFG_PATH, DATA_PATH, etc.)
Use XDG_CONFIG_HOME/ollie/* and XDG_DATA_HOME/ollie/* instead
- Replace OLLIE_<TAG>_LOG per-component logging with single OLLIE_LOG={level}
- Remove Route()/RouteRequest/RouteResult (replaced by direct backend selection)
- Update sandbox config to use XDG paths instead of OLLIE_*_PATH tokens
- Update docs accordingly
If AGENTS.md exists in the working directory, its contents are injected
as a user message at session start and after compaction (via spawnContext).
Not part of the system prompt — lives in conversation history.
Registry.Summaries() and Load() now scan the tools directory on every
call instead of reading from a startup cache. New tools dropped into
the directory are immediately visible without restarting the server.
When tool_list is called, the OnToolsChanged hook fires and updates
the preamble's '# Available Tools' section in the live agent runtime.
No agent reload or session restart required.
tools.Server is now a concrete struct (the local execution engine).
tools.Runner is the minimal 2-method interface for polymorphism
(satisfied by both Server and RemoteServer).
Deleted: Dispatcher, CWDSetter, EnvSetter, ToolRestrictionSetter
interfaces. Agent uses inline type assertions where needed.
The execute/ package no longer exists.
Runtime.Dispatcher → Runtime.ExecServer (tools.Server). The Dispatcher
abstraction routed to exactly one server ('execute') — unnecessary
indirection. BuildRuntime now takes tools.Server directly.
Agent.newDispatcher → Agent.newToolServer. AgentCfg/Config updated.
All GetServer('execute') calls replaced with direct ExecServer access.
exec closure simplified: direct CallTool, no tool-name lookup.
Runtime is guaranteed non-nil by the caller (session.New ensures it).
Remove defensive checks that can never fire. execServer() retains
the Dispatcher nil check (agents without tools are valid).
No external callers — only used internally by the detach methods
(Detach, ListDetached, SignalDetached, GetDetachedOutput, DismissDetached).
The accessor is now private, no longer leaking the dispatcher interface.
Session.Submit now forwards directly to Agent.Submit. Slash commands
are agent-level concerns handled by Agent.HandleCommand.
Session-level operations (save, resume, kill, rename, cwd) are ctl
verbs issued via the ctl file — they are NOT slash commands and do
not flow through the prompt path.
This eliminates the artificial session/agent command routing and
makes the prompt path clean: everything in the prompt goes to the agent.
Agent is now in ollie/agent with proper encapsulation:
- Unexported fields, exported methods as the API
- Own constructor (agent.NewAgent)
- Owns: turn execution, history, runtime, hooks, commands, compaction
- Session never reaches into agent internals
Session (ollie/session) is a thin host:
- Owns: persistence, session ID, env, detach delegation
- Delegates all agent operations through exported Agent methods
- handleCommand dispatches to agent.HandleCommand for agent-level commands
Agent-level commands (/model, /backend, /compact, /agent, etc.) live
in agent/commands.go and access internals directly (same package).
Session-level commands (/sessions, /save, /resume, /cwd, /help)
remain in session/commands.go.
Test files temporarily removed pending rewrite against new API.
The fifo_test.go passes as a sanity check.
The agent/ package is now session/ — because a session is the
top-level concept. The package contains:
- harness: implements Core, the session orchestrator
- Agent: the reasoning entity (history, runtime, tools)
- History: conversation accumulator
- Fifo, loop, compaction, hooks, commands
The old session/ package (which just held extracted fields) is
dissolved back into the harness. One package, clean ownership.
The harness is the Core implementation that wires a session.Session
to an Agent and drives execution. A session has a harness; the
harness runs the agent.
The struct that implements Core is the session host — it owns a
session.Session and an Agent, orchestrating their interaction.
Renaming makes the architecture self-documenting:
sessionHost (implements Core)
├── sess *session.Session — runtime environment
└── agent *Agent — reasoning entity (swappable)
The inner struct that holds agent-specific state (history, runtime,
backend, tools, hooks) is now called Agent — because that's what it
is. A session owns an Agent; the Agent performs reasoning within
the session's environment.
Move agent-specific fields into a nested 'reasoning' struct:
- history, runtime, cfg (per-turn params)
- agentName, agentsDir
- baseLayers, promptEnvExtra
- newDispatcher, newBackend
- currentAction, warnedContext, resultCache
The agent struct is now the session host (owns sess + reasoning).
On /agent swap, a new reasoning can be built from the new config
while the session (identity, state, bus, env) remains stable.
The agent struct receiver was historically 's' (from when it was
called 'session'). Now that session is a separate concept, rename
to 'a' to match the type name.
The Session type in agent/ is purely a conversation accumulator
(message history, token tracking, compaction state). Rename it to
History to free up 'Session' for the runtime environment concept
in the upcoming agent/session split.
- backend/codewhisperer{,_internal}.go: full Amazon CodeWhisperer
backend implementation — binary AWS event stream decoding, SQLite
auth for both enterprise OIDC and personal social/GitHub flows, OIDC
token refresh, and message encoding to the Kiro wire format
- backend/anthropic.go, copilot.go: new backends wired into New()
- backend/new.go: register anthropic, copilot, kiro/codewhisperer cases
- backend/openai.go: add extraHeaders hook for future use
- agent/loop.go: surface non-standard stop reasons as errors instead of
silently dropping them
- main.go: defaultModelForBackend() sets a sensible default per backend
(ollama→qwen3.5:9b, openrouter→deepseek/deepseek-v3.2,
anthropic→claude-sonnet-4-5, kiro→auto); /backend switch now also
resets the model to avoid stale foreign model IDs causing
ValidationException; add -prompt flag for non-interactive batch mode
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The "no tools" stall condition (totalToolCalls==0 && hadContent) fired
whenever the bot replied in plain text without calling any tools, which
is the normal completion path. The UI would then see agentStalled before
the done message and never transition back to agentIdle.
Removed the "no tools" stall; only the max-steps limit case is a real
stall worth surfacing.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Compact was gated on EvictedMessages(), which only returns messages
that have spilled past the 120k char soft limit. Normal sessions never
hit that threshold, making /compact a permanent no-op.
New approach: compact ALL messages older than the tail window (i.e.
everything computeTailStart puts before the protected tail), regardless
of budget. This makes /compact always useful.
Added OlderMessages(), TailWindow(), and SystemMessages() to
ContextBuilder to support the rewrite cleanly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Session persistence:
- Sessions saved to ~/.config/ollie/sessions/{id}.json after each
completed agent turn and after /compact
- --session <id> flag resumes a saved session, auto-loading its agent
- /sessions command lists saved sessions with agent and goal preview
- /clear and /agent switch generate a new session ID
- Session ID shown at startup
Dedup:
- file_read: overlap check now returns a hard error (was a warning
prepended to the result); check uses requested range, not actual
read range, to avoid redundant dispatchFileRead on overlap
- tool calls: duplicate (name, args) now returns a hard error; key
is only recorded on successful execution
/compact: Compact() now returns the summary text; displayed in the
UI so the user can verify quality; session saved after compact
Write-then-write:
- After a successful file_write, repopulate fileRanges with the
written range so a follow-up write to the same region does not
require a re-read
- Whole-file writes record exact new line count from content;
range writes record the written [start, end] with totalLines=0
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Revised system prompt with concrete prohibitions against verbosity and
premature stopping
- Added GenerationParams (max_tokens, temperature, frequency/presence penalty)
threading from agent config through backend ChatStream calls
- Added stall detection: emits "stalled" role on max-steps hit or zero tool
calls with content, surfaces as "stalled" in status bar
- Added per-session file read range tracking: warns on overlapping re-reads,
blocks file_write unless the target range was previously read
- Added general tool-call dedup: warns on exact (name, args) repeats for
non-file tools
- Both caches invalidated on /compact and /clear; file read cache invalidated
per-path on file_write
- Updated all agent configs with documenting defaults for new generation params
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
- /compact: summarizes evicted context messages via LLM call, replaces
them with a single summary system message
- /clear: resets session and display
- ContextBuilder.EvictedMessages(): returns messages outside bounded window
- Session.Compact(): drives the summarization and history replacement
- file_read and file_write require explicit y/n confirmation before executing
- New agentConfirming UI state with confirm [y/n] status bar indicator
- y/yes approves, n/no denies, anything else denies and falls through to
normal prompt handling
- file_read output includes line numbers for precise file_write targeting
- Fix tool output display: remove per-line squashWhitespace (preserves indentation)
- file_write description hints to preserve formatting
- Pass caller context into Executor.Execute so cancellation reaches
the subprocess immediately
- Use Setpgid + SIGKILL on process group to kill grandchild processes
- Add context to ToolExecutor signature and thread it through dispatch
- Reset agent state to idle after drainAgent
- Rollback incomplete session turn on interrupt
When the model emits a text-only turn containing narration phrases
("let me", "i'll", "i will", etc.) with no tool calls, the loop now
injects an ephemeral "Continue. Act now." user message and runs another
step rather than treating the turn as completion. Capped at 2 nudges
per run to prevent infinite loops. UI shows "[nudge: continuing…]" so
the user can see what happened.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Return RateLimitError from openai backend on HTTP 429, parsing the
Retry-After header (integer seconds or HTTP-date). The agent loop retries
up to 3 times with exponential backoff (5s/10s/20s) when no header is
given, emitting per-second countdown ticks. The status bar renders
"retry {N}s" during the wait.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Five improvements for cost/focus efficiency:
1. Budget now accounts for fixed per-request overhead (system prompt +
tool schemas), which were previously invisible to ContextBuilder and
could silently exceed intended limits by 30-50%.
2. Tail counting now only counts user and plain-assistant (no tool_calls)
messages toward TailMessages quota. Processed tool exchanges between
conversational turns are included but do not crowd out genuine context.
Trailing in-progress exchanges are always preserved and free.
3. BoundedHistory / BoundedHistoryWithNotice deduplicated into a single
buildBounded(injectNotice bool) private helper.
4. O(n²) prepend in the greedy inclusion loop replaced with
append+slices.Reverse.
5. exec.Executor.MaxOutputChars replaces the hardcoded 8000-char const.
Initialised from OLLIE_TOOL_OUTPUT_CHARS env var, defaults to 8000.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Loop struct and New constructor removed; Run is now a package-level
function taking (ctx, Config, State)
- Stop condition de-nested: natural stop (no tool calls → MarkComplete
+ break) and step-limit stop (step >= maxSteps-1 → break) are now
separate, sequential checks instead of a redundant outer/inner pair
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Both backends already implement streaming; the non-streaming fallback
was dead code that added complexity.
- Backend interface now requires only ChatStream; the separate
StreamingBackend interface and Chat method are removed
- Both OpenAIBackend and OllamaBackend lose their Chat methods and the
stream=bool parameter on their internal doChat helpers
- Loop.Run is simplified: one streaming path, no streamed flag, no
skippedCalls map, no shouldStop helper, no runStreamStep indirection
- "call" events are now emitted solely in the act phase, not in the
stream phase, so dedup tracking is no longer needed
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two bugs that combined to produce unrecoverable 400 errors from backends:
1. loop.go: runStreamStep returned nil on stream-closed-without-done,
causing Run to commit a partial assistant message (with no ToolCalls)
to state even when the backend had already sent tool call frames.
Now returns a real error so state.Update is never called and the
session history stays clean for a retry.
2. context.go: BoundedHistory and BoundedHistoryWithNotice could evict
an assistant[tool_calls] message while its paired tool[result] messages
remained in the tail window, producing a tool message with no preceding
assistant — rejected by strict backends (OpenAI, DeepInfra) with a 400.
Fixed by dropping assistant+tool pairs atomically in the hard-limit loop
and adding a sanitizeHistory pass that strips any remaining orphaned
tool messages before the history is returned.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
In the streaming path all tool calls were being added to skippedCalls,
causing the post-stream Act loop to suppress every 'call' emit. The
'-> tool(args)' lines therefore never appeared in the UI.
Fix: emit the 'call' OutputMsg inside runStreamStep once the stream
reports Done and all tool-call arguments are fully accumulated. The
Act loop still skips re-emitting them (correct), while the UI now
sees one 'call' event per tool invocation.
OutputMsg.Usage carries the raw backend.Usage value when Role=="usage",
rather than formatting it into Content. main.go can now read em.Usage
directly instead of parsing the formatted string.
Add agent/context.go with a ContextBuilder that enforces a character budget
(soft + hard) over the assembled history, evicting oldest non-system messages
first and replacing large tool outputs with truncated summaries.
Update agent/session.go to use ContextBuilder in History(), so the loop
always sees a bounded context window regardless of how many steps have run.
The budget defaults are:
- SoftLimit: 24000 chars (~6k tokens at 4 chars/token)
- HardLimit: 96000 chars (~24k tokens)
- MaxToolOutputChars: 2000
- TailMessages: 6 (always keep last N messages verbatim)
All limits are configurable via ContextConfig.