- Revised system prompt with concrete prohibitions against verbosity and
premature stopping
- Added GenerationParams (max_tokens, temperature, frequency/presence penalty)
threading from agent config through backend ChatStream calls
- Added stall detection: emits "stalled" role on max-steps hit or zero tool
calls with content, surfaces as "stalled" in status bar
- Added per-session file read range tracking: warns on overlapping re-reads,
blocks file_write unless the target range was previously read
- Added general tool-call dedup: warns on exact (name, args) repeats for
non-file tools
- Both caches invalidated on /compact and /clear; file read cache invalidated
per-path on file_write
- Updated all agent configs with documenting defaults for new generation params
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
- Split execute_code into execute_code, execute_tool, execute_pipe for
clearer model comprehension, especially on smaller models
- Ollama: accumulate tool calls across all stream chunks (fixes llama3.1
which delivers tool calls only on the done event)
- Ollama: add HTTP 429 / RateLimitError handling
- Ollama: map ToolCallID on outbound messages
- System prompt: inject cwd, current time, available tools
- System prompt: clarify tool usage rules
- Fix spurious spaces in streamed output (remove addStreamingContent heuristic)
- Fix tool output display: preserve newlines, add blank line separation
Both backends already implement streaming; the non-streaming fallback
was dead code that added complexity.
- Backend interface now requires only ChatStream; the separate
StreamingBackend interface and Chat method are removed
- Both OpenAIBackend and OllamaBackend lose their Chat methods and the
stream=bool parameter on their internal doChat helpers
- Loop.Run is simplified: one streaming path, no streamed flag, no
skippedCalls map, no shouldStop helper, no runStreamStep indirection
- "call" events are now emitted solely in the act phase, not in the
stream phase, so dedup tracking is no longer needed
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add Usage{InputTokens, OutputTokens} to backend.Response. Both
OllamaBackend (prompt_eval_count/eval_count) and OpenAIBackend
(prompt_tokens/completion_tokens) populate it. The loop emits a
'usage' OutputMsg after each Chat call; the UI displays it as
[↑N ↓N tokens].
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>