Why did OpenTelemetry split its token-usage metric into separate counters?
Because the metric it replaced, gen_ai.client.token.usage, silently mixed input and output tokens into one histogram keyed by a gen_ai.token.type attribute, so summing it without filtering by that type double-counted every call, and it had no way to record cache or reasoning tokens at all. A September 2026 change split it into five separate, monotonic per-modality counters, input, output, cache-read input, cache-write input, and reasoning output, plus two per-operation histograms for the distribution across calls rather than a running total.
The gap mattered because those two blind spots are exactly what makes a bill hard to reason about. Reasoning tokens on an extended-thinking call and cache reads on a repeated prompt each cost differently from a plain input token, and tracking what actually drives an agent’s spend needs to see each of those on its own rather than folded into one number a type filter has to unpack first.
The change removes the old histogram and its gen_ai.token.type attribute outright rather than keeping both, so a collector still reading the old field against an updated exporter gets nothing. Embeddings usage, which the old histogram covered, has no replacement yet; it’s left uncovered until its own follow-up metric ships.