Repository navigation
[issue-7773] [BE] fix: pick the cache cost calculator from the usage shape - #8723
anishmehta24 wants to merge 3 commits into
Conversation
|
@andrescrz could you take a look when you have time? |
…, cover priced cache writes
| int inputTokens = usage.getOrDefault("original_usage.prompt_tokens", | ||
| usage.getOrDefault("prompt_tokens", usage.getOrDefault("original_usage.input_tokens", 0))); |
There was a problem hiding this comment.
Misreads input/output totals across providers
This can reverse Responses billing and undercount Anthropic-shaped payloads because lines 74–75 always prefer normalized aliases over raw provider fields: swapped prompt/completion aliases override Responses input/output totals, while Anthropic input_tokens is treated as an OpenAI total and cache_read_input_tokens is ignored. Only use raw fields after confirming a Responses payload; otherwise preserve each provider’s fallback semantics.
Want Baz to fix this for you? Activate Fixer
Other fix methods
Prompt for AI Agents
Before applying, verify this suggestion against the current code. In
apps/opik-backend/src/main/java/com/comet/opik/domain/cost/SpanCostCalculator.java
around lines 74-75, update the input-token resolution logic in
`textGenerationWithCacheCostOpenAI` so it does not unconditionally mix normalized
aliases with raw provider fields. Detect a Responses-shaped payload before using
`original_usage.input_tokens`/`output_tokens` fallbacks, ensuring swapped
`prompt_tokens`/`completion_tokens` aliases cannot reverse billing; otherwise preserve
the existing provider-specific fallback behavior, including Anthropic `input_tokens` and
cache-read handling.
There was a problem hiding this comment.
Fair point, the fallback sat in the shared OpenAI calculator, so it ran for every provider routed there. Moved it in 00b241e: the OpenAI calculator is back to main's lookups, and only the Responses usage-shape route (after input_tokens_details.cached_tokens matched) fills missing prompt_tokens / completion_tokens from input_tokens / output_tokens. Existing aliases are never overridden.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Details
CostServicepicks the cache-aware calculator from the provider name, but the keys it has to read depend on the usage shape the span carries. When they don't line up, the cached tokens aren't seen and get billed at the full input rate, for example Anthropic-shaped usage ondeepinfra(OpenAI calculator) or OpenAI-shaped usage onanthropic.This picks the calculator from the raw
original_usage.*cache keys when exactly one shape (OpenAI chat, OpenAI responses, Anthropic, Bedrock, Gemini) has non-zero cache tokens, that shape's input count is present, and the model has a price for each cache bucket it carries. Anything else falls back to the provider map as before. Bare keys likecache_read_input_tokensmean different things depending on who wrote them, so they still go by provider. There's nocache <= prompt_tokenscomparison like in #7800. The Responses shape goes through a smallSpanCostCalculatorwrapper that fills missingprompt_tokens/completion_tokensfromoriginal_usage.input_tokens/output_tokens, so a raw Responses payload isn't billed 0 for input and output. The OpenAI calculator itself is unchanged.Change checklist
Issues
AI-WATERMARK
AI-WATERMARK: yes
Testing
Added a parameterized test in
CostServiceTestwith 14 cases. Nine fail on main: OpenAI chat / responses / Bedrock shape onanthropic, Bedrock shape with priced cache writes onanthropic, Anthropic shape with priced cache creation onbedrock, Anthropic shape ondeepinfra, Gemini shape onopenai, a Responses payload with no aliases, and a zero-valued key from another shape not blocking detection. Five keep the old cost: no cache tokens, bare cache key only, Anthropic cache key without its input count (Vercel-style), an unpriced cache bucket, and two shapes at once.mvn test -Dtest='CostServiceTest,SpanCostCalculatorTest'is 162/162 and 31/31 passing, andmvn spotless:checkis clean.Documentation
None needed.