Skip to content

[issue-7773] [BE] fix: pick the cache cost calculator from the usage shape - #8723

Open
anishmehta24 wants to merge 3 commits into
comet-ml:mainfrom
anishmehta24:anishmehta24/issue-7773-cache-cost-by-usage-shape
Open

anishmehta24 wants to merge 3 commits into
comet-ml:mainfrom
anishmehta24:anishmehta24/issue-7773-cache-cost-by-usage-shape

Conversation

@anishmehta24

@anishmehta24 anishmehta24 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Details

CostService picks the cache-aware calculator from the provider name, but the keys it has to read depend on the usage shape the span carries. When they don't line up, the cached tokens aren't seen and get billed at the full input rate, for example Anthropic-shaped usage on deepinfra (OpenAI calculator) or OpenAI-shaped usage on anthropic.

This picks the calculator from the raw original_usage.* cache keys when exactly one shape (OpenAI chat, OpenAI responses, Anthropic, Bedrock, Gemini) has non-zero cache tokens, that shape's input count is present, and the model has a price for each cache bucket it carries. Anything else falls back to the provider map as before. Bare keys like cache_read_input_tokens mean different things depending on who wrote them, so they still go by provider. There's no cache <= prompt_tokens comparison like in #7800. The Responses shape goes through a small SpanCostCalculator wrapper that fills missing prompt_tokens / completion_tokens from original_usage.input_tokens / output_tokens, so a raw Responses payload isn't billed 0 for input and output. The OpenAI calculator itself is unchanged.

Change checklist

  • User facing
  • Documentation update

Issues

AI-WATERMARK

AI-WATERMARK: yes

  • Tools: Claude Code
  • Scope: code and tests
  • Human verification: reviewed the change and ran the tests locally

Testing

Added a parameterized test in CostServiceTest with 14 cases. Nine fail on main: OpenAI chat / responses / Bedrock shape on anthropic, Bedrock shape with priced cache writes on anthropic, Anthropic shape with priced cache creation on bedrock, Anthropic shape on deepinfra, Gemini shape on openai, a Responses payload with no aliases, and a zero-valued key from another shape not blocking detection. Five keep the old cost: no cache tokens, bare cache key only, Anthropic cache key without its input count (Vercel-style), an unpriced cache bucket, and two shapes at once. mvn test -Dtest='CostServiceTest,SpanCostCalculatorTest' is 162/162 and 31/31 passing, and mvn spotless:check is clean.

Documentation

None needed.

@github-actions github-actions Bot added java Pull requests that update Java code Backend tests Including test files, or tests related like configuration. 🟡 size/M labels Oct 4, 2026
@anishmehta24

Copy link
Copy Markdown
Contributor Author

@andrescrz could you take a look when you have time?

Comment on lines +74 to +75
int inputTokens = usage.getOrDefault("original_usage.prompt_tokens",
usage.getOrDefault("prompt_tokens", usage.getOrDefault("original_usage.input_tokens", 0)));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Misreads input/output totals across providers

This can reverse Responses billing and undercount Anthropic-shaped payloads because lines 74–75 always prefer normalized aliases over raw provider fields: swapped prompt/completion aliases override Responses input/output totals, while Anthropic input_tokens is treated as an OpenAI total and cache_read_input_tokens is ignored. Only use raw fields after confirming a Responses payload; otherwise preserve each provider’s fallback semantics.

Severity

Want Baz to fix this for you? Activate Fixer

Other fix methods

Fix in Cursor

Prompt for AI Agents
Before applying, verify this suggestion against the current code. In
apps/opik-backend/src/main/java/com/comet/opik/domain/cost/SpanCostCalculator.java
around lines 74-75, update the input-token resolution logic in
`textGenerationWithCacheCostOpenAI` so it does not unconditionally mix normalized
aliases with raw provider fields. Detect a Responses-shaped payload before using
`original_usage.input_tokens`/`output_tokens` fallbacks, ensuring swapped
`prompt_tokens`/`completion_tokens` aliases cannot reverse billing; otherwise preserve
the existing provider-specific fallback behavior, including Anthropic `input_tokens` and
cache-read handling.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair point, the fallback sat in the shared OpenAI calculator, so it ran for every provider routed there. Moved it in 00b241e: the OpenAI calculator is back to main's lookups, and only the Responses usage-shape route (after input_tokens_details.cached_tokens matched) fills missing prompt_tokens / completion_tokens from input_tokens / output_tokens. Existing aliases are never overridden.

@anishmehta24
anishmehta24 marked this pull request as ready for review October 9, 2026 21:09
@anishmehta24
anishmehta24 requested a review from a team as a code owner October 9, 2026 21:09
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ⚠️ Failed 2026-10-09T21:11:59.884117Z 056d7c8 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Backend java Pull requests that update Java code 🟡 size/M tests Including test files, or tests related like configuration.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cache cost calculator is chosen by provider but should be chosen by usage shape

1 participant