Programmable test backend that satisfies Backend interface without
network calls. ChatStreamFunc can be set to control behavior
(blocking, custom responses, errors).
- backend/codewhisperer{,_internal}.go: full Amazon CodeWhisperer
backend implementation — binary AWS event stream decoding, SQLite
auth for both enterprise OIDC and personal social/GitHub flows, OIDC
token refresh, and message encoding to the Kiro wire format
- backend/anthropic.go, copilot.go: new backends wired into New()
- backend/new.go: register anthropic, copilot, kiro/codewhisperer cases
- backend/openai.go: add extraHeaders hook for future use
- agent/loop.go: surface non-standard stop reasons as errors instead of
silently dropping them
- main.go: defaultModelForBackend() sets a sensible default per backend
(ollama→qwen3.5:9b, openrouter→deepseek/deepseek-v3.2,
anthropic→claude-sonnet-4-5, kiro→auto); /backend switch now also
resets the model to avoid stale foreign model IDs causing
ValidationException; add -prompt flag for non-interactive batch mode
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Revised system prompt with concrete prohibitions against verbosity and
premature stopping
- Added GenerationParams (max_tokens, temperature, frequency/presence penalty)
threading from agent config through backend ChatStream calls
- Added stall detection: emits "stalled" role on max-steps hit or zero tool
calls with content, surfaces as "stalled" in status bar
- Added per-session file read range tracking: warns on overlapping re-reads,
blocks file_write unless the target range was previously read
- Added general tool-call dedup: warns on exact (name, args) repeats for
non-file tools
- Both caches invalidated on /compact and /clear; file read cache invalidated
per-path on file_write
- Updated all agent configs with documenting defaults for new generation params
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
- Split execute_code into execute_code, execute_tool, execute_pipe for
clearer model comprehension, especially on smaller models
- Ollama: accumulate tool calls across all stream chunks (fixes llama3.1
which delivers tool calls only on the done event)
- Ollama: add HTTP 429 / RateLimitError handling
- Ollama: map ToolCallID on outbound messages
- System prompt: inject cwd, current time, available tools
- System prompt: clarify tool usage rules
- Fix spurious spaces in streamed output (remove addStreamingContent heuristic)
- Fix tool output display: preserve newlines, add blank line separation
Return RateLimitError from openai backend on HTTP 429, parsing the
Retry-After header (integer seconds or HTTP-date). The agent loop retries
up to 3 times with exponential backoff (5s/10s/20s) when no header is
given, emitting per-second countdown ticks. The status bar renders
"retry {N}s" during the wait.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Both backends already implement streaming; the non-streaming fallback
was dead code that added complexity.
- Backend interface now requires only ChatStream; the separate
StreamingBackend interface and Chat method are removed
- Both OpenAIBackend and OllamaBackend lose their Chat methods and the
stream=bool parameter on their internal doChat helpers
- Loop.Run is simplified: one streaming path, no streamed flag, no
skippedCalls map, no shouldStop helper, no runStreamStep indirection
- "call" events are now emitted solely in the act phase, not in the
stream phase, so dedup tracking is no longer needed
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
OpenAI sends usage in a separate trailing SSE chunk, not in the same
chunk that carries finish_reason. The previous code read wire.Usage
from the finish_reason chunk (where it is always zero) and discarded
all subsequent chunks.
Fixes:
- Add stream_options:{include_usage:true} to the request so OpenAI
actually includes usage in the stream at all.
- Accumulate usage across every chunk instead of reading it once at
finish_reason time.
- On the Done event, use the accumulated usage rather than the
(always-zero) usage from the finish_reason chunk.
- Keep processing chunks after finish_reason until data:[DONE] so
the trailing usage chunk is not silently dropped.
openrouter is OpenAI-compatible; just set OLLIE_OPENAI_URL to the
OpenRouter endpoint. No separate backend case needed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Add Usage{InputTokens, OutputTokens} to backend.Response. Both
OllamaBackend (prompt_eval_count/eval_count) and OpenAIBackend
(prompt_tokens/completion_tokens) populate it. The loop emits a
'usage' OutputMsg after each Chat call; the UI displays it as
[↑N ↓N tokens].
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Consistent with OLLIE_OPENAI_URL; makes it clear the key only applies
to openai/openrouter backends.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
OLLIE_API_URL was shared across backends, causing the ollama backend
to use the OpenRouter URL when OLLIE_BACKEND was overridden on the
command line while the env file still set OLLIE_API_URL.
Replace with OLLIE_OLLAMA_URL and OLLIE_OPENAI_URL so each backend
only reads its own URL setting.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Env file is optional; existing environment variables take precedence.
Supports KEY=VALUE format with # comments and blank lines ignored.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>