# Reasoning tokens count toward the bill and cap

> Thinking uses billed tokens before an answer appears. Check current API usage fields, a hypothetical cost example, and the limits of five archived records.

- Published: 2026-10-05
- Data as of: 2026-10-05
- Article slug: reasoning-models-explained
- Author: [Glevd](https://x.com/glevd)
- Reading time: 10 minutes
- Topics: reasoning-tokens, thinking-tokens, cost, tokens, evidence
- Data policy: Dated analysis; static benchmark claims are retained.
- Canonical URL: https://benchlm.ai/blog/posts/reasoning-models-explained

Reasoning tokens count toward the bill even when the final answer is short or absent.

In the September 14, 2026 support-extraction archive, three attempts consumed a 300-token output cap without a usable answer. Their output charges remained in the saved records.

Those are caller-supplied outputs whose model origin isn't independently verified. Current provider guides checked October 5 describe how a token limit can stop a reply before it appears.

## The record shows a charge without a usable answer

Our support-extraction kit asks for a small JSON object with four fields. Case 29 describes a payment error and asks to retry the payment. Its answer should fit in a short reply, but the saved September 14 long-prompt record ends three times with `finish_reason: "length"`.

Each attempt reports 300 completion tokens. The reasoning breakdown is 299, 300, and 299 tokens, and each output charge is $0.003. Each saved output contains `""`, not the requested JSON object. Our extractor uses that spelling when message content is absent, so the record cannot distinguish missing content from an empty JSON string. It does establish that no usable four-field answer was saved.

Those are reported charges in caller-supplied records, not invoices or independently authenticated model responses. The configuration's manifest labels it `anthropic/claude-sonnet-5`, routed through OpenRouter, with temperature 0 and a 300-token cap. The [receipt](/downloads/evidence-kits/support-extraction-v1/runs/2026-09-14-long-cached/receipt.json) preserves that provenance limit, and the [attempt log](/downloads/evidence-kits/support-extraction-v1/runs/2026-09-14-long-cached/attempts.jsonl) preserves every repetition.

We checked the saved accounting, including the attempts that failed the task. Across the configuration's 87 completed calls, 8,919 output tokens include 4,980 reasoning tokens. All three capped attempts belong in that total even though they produced nothing the task could accept. Dropping them would make both the apparent cost and the failure rate look better.

![Reported output tokens in three saved capped attempts: 299, 300, and 299 reasoning tokens within a 300-token output cap. Caller-supplied records. Model origin is not independently verified.](/images/blog/reasoning-cap-records.svg)

The [structured-output guide](/blog/posts/test-structured-output) covers whether a reply meets the schema and the task. Here the question is what the output limit and usage count contain.

## A short reply can use far more billed tokens

Reasoning tokens are generated while a model works through a request, including work that the answer does not show. Providers also call them thinking tokens. The [reasoning model shortlist](/best/reasoning-models) covers model selection. This post deals with the resulting token count.

For an inclusive output count, the arithmetic is straightforward:

```text
output cost = total billed output tokens × output rate / 1,000,000
reasoning cost = reasoning tokens × output rate / 1,000,000
```

Reasoning cost is part of the generated output bill. Adding it to an already inclusive output bill would charge for the reasoning twice. Check the endpoint's definition of the total before using either formula.

Consider a hypothetical request with 2,000 input tokens, a 200-token answer, and 1,800 reasoning tokens. Assume $2 per million input tokens and $10 per million output tokens. These are chosen example rates, not a quotation for any named model, and the token counts are invented to show the arithmetic.

| Part of the hypothetical request | Tokens | Assumed rate per million | Calculated cost |
| --- | --- | --- | --- |
| Input | 2,000 | $2 | $0.004 |
| Visible answer | 200 | $10 | $0.002 |
| Reasoning | 1,800 | $10 | $0.018 |
| Whole request | 4,000 | Input and output priced separately | $0.024 |

Counting only the answer would predict $0.006 for input plus output. Including the reasoning makes it $0.024, four times that estimate. Neither number is an observed result. This example isolates token volume. Cache charges, tools, discounts, and retries would need their own entries.

Compare current rates in the [API pricing table](/pricing), then enter the actual workload in the [cost calculator](/tools/cost-calculator). A rate without a generated-token count is an incomplete request estimate.

## Two reasoning percentages can describe the same record

Reasoning share means reasoning tokens divided by the inclusive output total. Extra reasoning above the remaining output uses a different denominator. A calculator that starts with the visible reply and adds a percentage needs that second denominator.

We recomputed the following figures from five archived attempt logs dated September 14 and September 25, 2026. Model names in the table are labels from the manifests. All outputs are caller-supplied, their model origin is not independently verified, and the fixture contains 30 authored support tickets rather than customer traffic.

| Manifest label and configuration | Record date | Calls with usage | Output tokens, including reasoning | Reasoning tokens | Reasoning share |
| --- | --- | --- | --- | --- | --- |
| Claude Sonnet 5, short prompt | September 14 | 87 | 6,517 | 2,468 | 37.9% |
| Claude Sonnet 5, long cached prompt | September 14 | 87 | 8,919 | 4,980 | 55.8% |
| GPT-5.6 Terra, short prompt | September 14 | 87 | 3,251 | 97 | 3.0% |
| GPT-5.6 Terra, long cached prompt | September 14 | 87 | 3,135 | 0 | 0.0% |
| GPT-5.6 Terra, Chat Completions | September 25 | 85 | 3,066 | 0 | 0.0% |
| GPT-5.6 Terra, Responses | September 25 | 87 | 3,135 | 0 | 0.0% |
| GPT-6 Sol, Chat Completions | September 25 | 87 | 4,624 | 1,416 | 30.6% |
| GPT-6 Sol, Responses | September 25 | 87 | 4,806 | 1,632 | 34.0% |
| GPT-6 Astra, Chat Completions | September 25 | 87 | 3,779 | 617 | 16.3% |
| GPT-6 Astra, Responses | September 25 | 87 | 3,766 | 610 | 16.2% |

Every configuration attempted 90 calls. Transport refusals reduce each configuration to 87 completed responses. Two completed Terra Chat Completions responses on September 25 lack usage, leaving 85 for that row. Its zeros describe the calls with usage. They do not fill the missing two.

In the long-prompt Sonnet-labeled record, reasoning is 4,980 / 8,919, or 55.8% of the inclusive output total. Subtracting the reasoning leaves 3,939 tokens. Dividing 4,980 by that remainder gives 126.4% extra, not 55.8%. That remainder can include formatting overhead, so it is not an exact count of text the reader sees.

Applying a 55.8% surcharge to that remainder would therefore underestimate the recorded total. Keep the denominator beside the percentage. The [short-prompt log](/downloads/evidence-kits/support-extraction-v1/runs/2026-09-14-short/attempts.jsonl) and the [Sol endpoint log](/downloads/evidence-kits/support-extraction-v1/runs/2026-09-25-endpoint-sol/attempts.jsonl) let you inspect the other configurations without treating this spread as a model ranking.

## The endpoint decides whether output already includes thinking

Providers do not count output fields the same way. Current OpenAI and Claude documentation describes inclusive totals. Google's Interactions documentation gives separate output and thought counts for pricing. We checked these field definitions on October 5, 2026.

| Endpoint | Total to read for generated-token pricing | Reasoning breakdown | Add the breakdown to that total? |
| --- | --- | --- | --- |
| OpenAI Responses | `usage.output_tokens` | `usage.output_tokens_details.reasoning_tokens` | No, reasoning is included |
| Claude Messages | `usage.output_tokens` | `usage.output_tokens_details.thinking_tokens` | No, thinking is included |
| Gemini Interactions | `usage.total_output_tokens`<br><br>plus<br><br>`usage.total_thought_tokens` | `usage.total_thought_tokens` | Yes, the documented pricing sum uses both |
| Archived OpenRouter records above | `usage.completion_tokens` | `usage.completion_tokens_details.reasoning_tokens` | No, the recorded completion charge includes reasoning |

[OpenAI's reasoning guide](https://developers.openai.com/api/docs/guides/reasoning) explains that reasoning is billed as output and counted within `max_output_tokens`. [Claude's steering guide](https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost) identifies `output_tokens` as the authoritative inclusive billing total. A returned summary can be shorter than the internal thinking that was billed.

[Gemini's thinking guide](https://ai.google.dev/gemini-api/docs/thinking) explicitly prices output plus thinking for Interactions. Applying an inclusive-output formula to those separate fields would undercount the bill. Applying the Google sum to an inclusive OpenAI or Claude total would double-count thinking instead.

Endpoint names belong in a cost log alongside model names. The [Chat Completions and Responses comparison](/blog/posts/chat-completions-to-responses-api) covers other differences between the OpenAI routes. If an SDK or gateway renames a usage field, preserve the original usage object until you have checked the adapter's definition.

## Thinking has to fit inside the output cap

Your output cap controls the generated-token allowance, which includes work beyond the answer the reader sees. OpenAI documents that a Responses request can become `incomplete` with `incomplete_details.reason` set to `max_output_tokens` before visible output begins. Google documents the same possibility for Interactions: the hard cutoff includes thought tokens, and a request can still accrue thinking charges.

Our archived example used a 300-token cap. We cannot learn the required uncapped budget from attempts that stopped at 300. Reaching the boundary tells us the cap was exhausted. It does not reveal how much further the work would have gone or whether a larger cap would have produced a correct answer.

Separate the output cap from reasoning effort. Effort guides how much work a supported model attempts. Your cap stops generation when the allowance is exhausted. Raising the cap gives more room, while lowering effort changes the work. Neither change proves that the reply will satisfy your task.

To pick a budget, record the generated total and stopping reason on a representative set of requests. Include ambiguous inputs and failures, then inspect the high end rather than sizing everything from the average reply length. Keep requests at the cap in a separate group because their full token demand is unknown.

If you test a lower effort setting, run the same task checks as before. Compare usable answers as well as spend. Reducing token use loses its advantage when enough failures or manual corrections outweigh the saving. Our archived records do not supply that crossover point for your workload.

## The useful denominator is an accepted result

Cost per request answers what one attempt consumed. Cost per accepted result answers what the work cost after attempts that failed the checks. For extraction, those checks can distinguish valid JSON from a correctly classified ticket, as the [custom benchmark guide](/blog/posts/building-custom-llm-benchmark) demonstrates.

We keep those counts separate in the evidence kit. A transport refusal, a capped reply, and a plausible but wrong field can all prevent acceptance, but each suggests a different fix. Every attempt belongs in the acceptance-rate denominator, and every reported charge belongs in the cost total. Cost per accepted result divides that total by accepted attempts only.

Reasoning-token share is useful for explaining a bill, but it cannot establish that the thinking was worthwhile. A different task can make the archived configuration with no recorded reasoning more expensive overall. Equally, a high share is not evidence of better judgment. Compare checked outputs under the settings you intend to use.

All five archived logs cover one synthetic fixture, caller-supplied model labels, default effort, and a small output cap. They establish arithmetic on the saved records. They do not establish current provider prices, general model quality, production savings, or an appropriate reasoning allowance for other tasks.

On your next evaluation, retain the inclusive generated count, its reasoning breakdown, the stopping reason, the charge, and whether the output passed. Those five fields let you spot the request that spent its allowance before delivering the answer you needed.

---

## Questions

### Are reasoning tokens billed as output tokens?

OpenAI and Claude charge for internal thinking at output-token rates. Their generated totals already contain that thinking, so adding the detail count again overstates the charge. Gemini Interactions documents separate output and thought fields whose costs are added together. Check your endpoint before turning usage into a bill.

### Do reasoning tokens count toward the maximum output tokens?

Yes on the Responses and Interactions routes documented here. Their limits include reasoning as well as the answer. With a small allowance, generation can end while the model is still thinking. Inspect the termination status and token count together. A reply that reaches the boundary has unknown uncapped demand.

### How many reasoning tokens will a request use?

Token demand depends on the request, model, settings, and prompt. Our archived synthetic fixture contains records with no reported reasoning and others that use nearly the entire allowance on it. Those records do not predict customer traffic. Measure representative tasks, including difficult inputs, and keep capped attempts separate.

### Why can an API return no answer and still charge?

An API can consume billed input and reasoning before the visible reply starts. If generation reaches the token limit at that point, the answer can be absent or incomplete while usage still accrues. That behavior is documented by OpenAI and Google. Check the stopping reason before automatically retrying the same configuration.

---

*Sources: OpenAI reasoning documentation, Claude thinking documentation, and Gemini thinking documentation, checked October 5, 2026. Archived support-extraction attempt records dated September 14 and September 25, 2026, recomputed October 5. Records are caller-supplied. Model origin is not independently verified.*

## Frequently asked questions

### Are reasoning tokens billed as output tokens?

OpenAI and Claude charge for internal thinking at output-token rates. Their generated totals already contain that thinking, so adding the detail count again overstates the charge. Gemini Interactions documents separate output and thought fields whose costs are added together. Check your endpoint before turning usage into a bill.

### Do reasoning tokens count toward the maximum output tokens?

Yes on the Responses and Interactions routes documented here. Their limits include reasoning as well as the answer. With a small allowance, generation can end while the model is still thinking. Inspect the termination status and token count together. A reply that reaches the boundary has unknown uncapped demand.

### How many reasoning tokens will a request use?

Token demand depends on the request, model, settings, and prompt. Our archived synthetic fixture contains records with no reported reasoning and others that use nearly the entire allowance on it. Those records do not predict customer traffic. Measure representative tasks, including difficult inputs, and keep capped attempts separate.

### Why can an API return no answer and still charge?

An API can consume billed input and reasoning before the visible reply starts. If generation reaches the token limit at that point, the answer can be absent or incomplete while usage still accrues. That behavior is documented by OpenAI and Google. Check the stopping reason before automatically retrying the same configuration.
