- backend/codewhisperer{,_internal}.go: full Amazon CodeWhisperer
backend implementation — binary AWS event stream decoding, SQLite
auth for both enterprise OIDC and personal social/GitHub flows, OIDC
token refresh, and message encoding to the Kiro wire format
- backend/anthropic.go, copilot.go: new backends wired into New()
- backend/new.go: register anthropic, copilot, kiro/codewhisperer cases
- backend/openai.go: add extraHeaders hook for future use
- agent/loop.go: surface non-standard stop reasons as errors instead of
silently dropping them
- main.go: defaultModelForBackend() sets a sensible default per backend
(ollama→qwen3.5:9b, openrouter→deepseek/deepseek-v3.2,
anthropic→claude-sonnet-4-5, kiro→auto); /backend switch now also
resets the model to avoid stale foreign model IDs causing
ValidationException; add -prompt flag for non-interactive batch mode
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Revised system prompt with concrete prohibitions against verbosity and
premature stopping
- Added GenerationParams (max_tokens, temperature, frequency/presence penalty)
threading from agent config through backend ChatStream calls
- Added stall detection: emits "stalled" role on max-steps hit or zero tool
calls with content, surfaces as "stalled" in status bar
- Added per-session file read range tracking: warns on overlapping re-reads,
blocks file_write unless the target range was previously read
- Added general tool-call dedup: warns on exact (name, args) repeats for
non-file tools
- Both caches invalidated on /compact and /clear; file read cache invalidated
per-path on file_write
- Updated all agent configs with documenting defaults for new generation params
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Return RateLimitError from openai backend on HTTP 429, parsing the
Retry-After header (integer seconds or HTTP-date). The agent loop retries
up to 3 times with exponential backoff (5s/10s/20s) when no header is
given, emitting per-second countdown ticks. The status bar renders
"retry {N}s" during the wait.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Both backends already implement streaming; the non-streaming fallback
was dead code that added complexity.
- Backend interface now requires only ChatStream; the separate
StreamingBackend interface and Chat method are removed
- Both OpenAIBackend and OllamaBackend lose their Chat methods and the
stream=bool parameter on their internal doChat helpers
- Loop.Run is simplified: one streaming path, no streamed flag, no
skippedCalls map, no shouldStop helper, no runStreamStep indirection
- "call" events are now emitted solely in the act phase, not in the
stream phase, so dedup tracking is no longer needed
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
OpenAI sends usage in a separate trailing SSE chunk, not in the same
chunk that carries finish_reason. The previous code read wire.Usage
from the finish_reason chunk (where it is always zero) and discarded
all subsequent chunks.
Fixes:
- Add stream_options:{include_usage:true} to the request so OpenAI
actually includes usage in the stream at all.
- Accumulate usage across every chunk instead of reading it once at
finish_reason time.
- On the Done event, use the accumulated usage rather than the
(always-zero) usage from the finish_reason chunk.
- Keep processing chunks after finish_reason until data:[DONE] so
the trailing usage chunk is not silently dropped.
Add Usage{InputTokens, OutputTokens} to backend.Response. Both
OllamaBackend (prompt_eval_count/eval_count) and OpenAIBackend
(prompt_tokens/completion_tokens) populate it. The loop emits a
'usage' OutputMsg after each Chat call; the UI displays it as
[↑N ↓N tokens].
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>