a2223cef60
The context fullness gauge was driven by output token count, causing erratic jumps between turns (e.g. 84% -> 64%) with no compaction. Status bar and estimateContextTokens now use calculatePromptTokens() which returns input + cacheRead + cacheWrite — the actual input context size. Previously both used a formula that included the final output token count, which fluctuates with response length and is not part of the context window for the current request. isContextOverflow's usage-based fallback (z.ai silent overflow) was also missing cacheWrite (cache_creation_input_tokens). Per Anthropic docs the threshold is input + cache_read + cache_creation — all three. Ref: https://platform.claude.com/docs/en/about-claude/pricing#long-context-pricing google.ts and google-vertex.ts were double-counting cached tokens. Gemini's promptTokenCount already includes cachedContentTokenCount, so assigning input = promptTokenCount and cacheRead = cachedContentTokenCount overcounted by cachedContentTokenCount on every cached request. Fixed by subtracting first, matching the OpenAI convention: input = promptTokenCount - cachedContentTokenCount cacheRead = cachedContentTokenCount => input + cacheRead = promptTokenCount (total prompt, no double-count) Ref: https://ai.google.dev/api/generate-content#v1beta.GenerateContentResponse.UsageMetadata All other providers validated: amazon-bedrock (inputTokens is uncached by API contract), openai-completions/responses/azure (already subtract cached), kimi/gitlab-duo (delegate to correct implementations), cursor (API exposes output tokens only — input stays 0 by design). Co-authored-by: Miroslav Drbal <miroslav.drbal@gendigital.com>
@oh-my-pi/pi-coding-agent
Core implementation package for the omp coding agent in the oh-my-pi monorepo.
For installation, setup, provider configuration, model roles, slash commands, and full CLI reference, see:
Package-specific references: