# Why did OpenTelemetry split its token-usage metric into separate counters?

Because the metric it replaced, `gen_ai.client.token.usage`, silently mixed input and output tokens into one histogram keyed by a `gen_ai.token.type` attribute, so summing it without filtering by that type double-counted every call, and it had no way to record cache or reasoning tokens at all. A September 2026 change split it into five separate, monotonic per-modality counters, input, output, cache-read input, cache-write input, and reasoning output, plus two per-operation histograms for the distribution across calls rather than a running total.

The gap mattered because those two blind spots are exactly what makes a bill hard to reason about. Reasoning tokens on an extended-thinking call and cache reads on a repeated prompt each cost differently from a plain input token, and [tracking what actually drives an agent's spend](/answers/eval-costs/what-drives-the-cost-of-evaluating-an-ai-agent) needs to see each of those on its own rather than folded into one number a type filter has to unpack first.

The change removes the old histogram and its `gen_ai.token.type` attribute outright rather than keeping both, so a collector still reading the old field against an updated exporter gets nothing. Embeddings usage, which the old histogram covered, has no replacement yet; it's left uncovered until its own follow-up metric ships.

---

Sources:
- OpenTelemetry semantic-conventions-genai PR #374: Fix usage metrics to provide meaningful aggregation and break down by modality, reasoning, cache usage: https://github.com/open-telemetry/semantic-conventions-genai/pull/374 (fetched 2026-09-30)

Source: https://tessary.ai/answers/otel-genai-conventions/why-did-opentelemetry-split-its-token-usage-metric-into-separate-counters
More on Otel genai conventions: https://tessary.ai/answers/otel-genai-conventions
From Tessary, agent reliability for AI agents in production: https://tessary.ai
