AI Cost Calculator
Estimate AI cost per blog post, website page, documentation article, PRD, or shipped feature without thinking in tokens first.
Choose a workload
Workload assumptions
Included In This Estimate
8 kickoff/alignment turns per month
6 sources reviewed per post at ~1200 words each
3 research turns and ~550 note words per turn
250 planning words and 200 formatting/meta output words per post
20% cache savings assumed on repeated source/context reuse
35% extra billed output assumed for reasoning-model thinking tokens
These workflow assumptions are baked into the preset for this use case so the front end stays simple while the estimate still includes kickoff back-and-forth, research, formatting, context reuse, and model thinking overhead.
Built-in web search and grounding tool-call fees are not yet added to the estimate. If your workflow uses provider search tools heavily, actual cost can still be higher.
Models to compare
32 selectedEstimated AI costs for this workload
Monthly input tokens
382,980
Visible output tokens
48,360
Estimated billed output
48,360 - 65,286
Derived from
8 posts x 1500 final words, plus 8 project setup turns, 3 research turns, 6 sources, formatting overhead, and 1 revision rounds
This workload uses about 294,600 prompt words and 37,200 visible response words per month before converting to tokens. Billed output can be higher for reasoning models because provider docs now count hidden thinking tokens too.
$0 per post · $0/M input · $0/M output
Overall 68.5 · #27
Billed output tokens: 65,286 (includes thinking overhead)
$0.02 per post · $0.3/M input · $1.2/M output
Overall 63.9 · #49Best score per dollar
Billed output tokens: 48,360
$0.02 per post · $0.25/M input · $1.5/M output
Overall 56.1 · #95
Billed output tokens: 48,360
$0.03 per post · $0.435/M input · $0.87/M output
Overall 66.4 · #33
Billed output tokens: 65,286 (includes thinking overhead)
$0.04 per post · $0.3/M input · $2.5/M output
Overall 60.5 · #67
Billed output tokens: 65,286 (includes thinking overhead)
$0.07 per post · $0.75/M input · $3.75/M output
Overall 78.4 · #6
Billed output tokens: 65,286 (includes thinking overhead)
$0.07 per post · $0.75/M input · $3.75/M output
Overall 70.2 · #17
Billed output tokens: 65,286 (includes thinking overhead)
$0.07 per post · $1/M input · $3.2/M output
Overall 59.2 · #72
Billed output tokens: 65,286 (includes thinking overhead)
$0.08 per post · $0.95/M input · $4/M output
Overall 65.4 · #40
Billed output tokens: 65,286 (includes thinking overhead)
$0.1 per post · $1.25/M input · $4.25/M output
Not publicly ranked
Billed output tokens: 65,286 (includes thinking overhead)
$0.1 per post · $1/M input · $6/M output
Overall 65.5 · #37
Billed output tokens: 65,286 (includes thinking overhead)
$0.1 per post · $1.4/M input · $4.4/M output
Overall 68.2 · #30
Billed output tokens: 65,286 (includes thinking overhead)
$0.1 per post · $1.4/M input · $4.4/M output
Overall 64.4 · #45
Billed output tokens: 65,286 (includes thinking overhead)
$0.13 per post · $1.5/M input · $7.5/M output
Overall 70.1 · #19
Billed output tokens: 65,286 (includes thinking overhead)
$0.14 per post · $2/M input · $6/M output
Overall 70.2 · #18
Billed output tokens: 65,286 (includes thinking overhead)
$0.14 per post · $1.5/M input · $9/M output
Overall 70 · #20
Billed output tokens: 65,286 (includes thinking overhead)
$0.18 per post · $2/M input · $10/M output
Overall 70.8 · #15
Billed output tokens: 65,286 (includes thinking overhead)
$0.19 per post · $2/M input · $12/M output
Overall 69.8 · #23
Billed output tokens: 65,286 (includes thinking overhead)
$0.24 per post · $2.5/M input · $15/M output
Overall 71.4 · #13
Billed output tokens: 65,286 (includes thinking overhead)
$0.27 per post · $3/M input · $15/M output
Overall 74.9 · #7
Billed output tokens: 65,286 (includes thinking overhead)
$0.39 per post · $5/M input · $25/M output
Overall 70.2 · #16
Billed output tokens: 48,360
$0.44 per post · $5/M input · $25/M output
Overall 80.7 · #4
Billed output tokens: 65,286 (includes thinking overhead)
$0.44 per post · $5/M input · $25/M output
Overall 72.3 · #10
Billed output tokens: 65,286 (includes thinking overhead)
$0.44 per post · $5/M input · $25/M output
Overall 69.9 · #21
Billed output tokens: 65,286 (includes thinking overhead)
$0.48 per post · $5/M input · $30/M output
Overall 79.6 · #5
Billed output tokens: 65,286 (includes thinking overhead)
$0.48 per post · $5/M input · $30/M output
Overall 73.3 · #8
Billed output tokens: 65,286 (includes thinking overhead)
$0.89 per post · $10/M input · $50/M output
Overall 83 · #1
Billed output tokens: 65,286 (includes thinking overhead)
$0.89 per post · $10/M input · $50/M output
Overall 81.1 · #2
Billed output tokens: 65,286 (includes thinking overhead)
$0.89 per post · $10/M input · $50/M output
Not publicly ranked
Billed output tokens: 65,286 (includes thinking overhead)
$0.89 per post · $10/M input · $50/M output
Overall 80.9 · #3
Billed output tokens: 65,286 (includes thinking overhead)
$2.91 per post · $30/M input · $180/M output
Overall 63.1 · #53
Billed output tokens: 65,286 (includes thinking overhead)
GLM-5.3 is the cheapest option for this workload at $0/mo. Compared with GPT-5.5 Pro, that saves $23.24/mo.
MiniMax M3 buys the most benchmark score per dollar here: overall 63.9 (#49) at $0.17/mo. Cheapest is not the same question as cheapest that clears your quality bar.
Pricing note: GLM-5.3: Z.AI publishes the GLM-5.3 FP8 checkpoint under the custom GLM-5.3 License for self-hosting and does not publish a distinct first-party hosted per-token rate for this exact checkpoint. We represent the open-weight row as self-host/free-per-token before infrastructure costs. The hosted Z.AI route and Coding Plan must not be read as free from this self-host placeholder. MiniMax M3: MiniMax's official API pricing page lists the permanent 50%-off standard rate for MiniMax-M3 at $0.30 input / $0.06 cached input / $1.20 output per million tokens for prompts up to 512K. The 512K-to-1M tier is $0.60 / $0.12 / $2.40; BenchLM stores the standard short-context rate. Gemini 3.1 Flash-Lite: Google's Gemini Developer API pricing page lists gemini-3.1-flash-lite at $0.25 input / $0.025 cached input / $1.50 output per million tokens for text, image, and video on the standard paid tier. DeepSeek V4 Pro 0813: DeepSeek's pricing page checked August 12, 2026 identifies the current `deepseek-v4-pro` API version as DeepSeek-V4-Pro-0813 and lists $0.435 cache-miss input / $0.003625 cache-hit input / $0.87 output per million tokens. DeepSeek warns that an overall API price increase is planned but has not published the new rates or effective date. DeepSeek has not published weights for the 0813 hosted checkpoint. Gemini 3.5 Flash-Lite: Google's Gemini Developer API pricing page lists gemini-3.5-flash-lite at $0.30 input / $0.03 cached input / $2.50 output per million tokens on the standard paid tier. Gemini 3.8 Flash: Google lists introductory Gemini 3.8 Flash rates through December 31, 2026: $0.75 input, $0.075 cached input, and $3.75 output per million tokens. Flex and Batch input and output rates are half those amounts. On January 1, 2027, standard rates rise to $1.50 input, $0.15 cached input, and $7.50 output; Flex and Batch rates rise to $0.75 input, $0.075 cached input, and $3.75 output. Cache storage is $0.50 per million tokens per hour through 2026 and $1.00 afterward. Gemini 3.7 Flash: Google lists introductory Gemini 3.7 Flash rates through December 31, 2026: $0.75 input, $0.075 cached input, and $3.75 output per million tokens. Flex and Batch input and output rates are half those amounts. On January 1, 2027, standard rates rise to $1.50 input, $0.15 cached input, and $7.50 output; Flex and Batch rates rise to $0.75 input, $0.075 cached input, and $3.75 output. Priority rates are $1.35 input, $0.135 cached input, and $6.75 output during the promotion, doubling in 2027. Kimi 2.6: Moonshot's Kimi API platform homepage lists K2.6 at $0.95 input / $4.00 output per million tokens. GLM-5 (Reasoning): Z.AI's model docs say GLM-5 supports thinking modes, and the official pricing page lists GLM-5 at $1.00 input / $3.20 output per million tokens. Kimi K2.7 Code: Moonshot's Kimi API platform lists kimi-k2.7-code at $0.95 cache-miss input / $4.00 output per million tokens, with cache-hit input priced at $0.19 per million tokens. Muse Spark 1.3: Meta Model API lists `muse-spark-1.3` at $1.25 input, $0.15 cached input, and $4.25 output per million tokens. Meta separately lists `muse-spark-1.3-contributor`, whose data may be used to improve Meta products, at $0.10 input, $0.002 cached input, and $0.20 output per million tokens; the standard non-contributor SKU is the headline rate. GPT-5.6 Luna: OpenAI's API billing price sheet lists GPT-5.6 Luna at $1.00 input / $0.10 cached input / $6.00 output per million tokens for short-context requests. Prompts above 272K input tokens use the published $2.00 input / $0.20 cached input / $9.00 output long-context tier; cache writes cost $1.25 per million tokens. GLM-5.2: Z.AI's official pricing page lists GLM-5.2 at $1.40 input / $4.40 output per million tokens. GLM-5.1: Z.AI's official pricing page lists GLM-5.1 at $1.40 input / $4.40 output per million tokens. Gemini 3.6 Flash: Google's Gemini Developer API pricing page lists gemini-3.6-flash at $1.50 input / $0.15 cached input / $7.50 output per million tokens on the standard paid tier. Grok 4.6: xAI's August 12, 2026 release notes list grok-4.6 at $2.00 input / $0.50 cached input / $6.00 output per million tokens below 200K prompt tokens. Prompts above 200K use the published $4.00 / $1.00 / $12.00 long-context tier; BenchLM stores the short-context rate. Gemini 3.5 Flash: Google's official Gemini Developer API pricing page lists Gemini 3.5 Flash at $1.50 input / $0.15 cached input / $9.00 output per million tokens on the standard paid tier. Claude Sonnet 5: Anthropic's June 30, 2026 Claude Sonnet 5 launch page lists introductory pricing of $2 input / $10 output per million tokens through August 31, 2026, then $3 input / $15 output per million tokens. The system card says standard configurations use 1M tokens, with BrowseComp using a 10M-token limit through compaction. Gemini 3.1 Pro: Google's Gemini Developer API pricing page lists gemini-3.1-pro-preview at $2.00 input / $0.20 cached input / $12.00 output per million tokens for prompts up to 200K tokens, rising to $4.00 / $0.40 / $18.00 above 200K. GPT-5.6 Terra: OpenAI's API billing price sheet lists GPT-5.6 Terra at $2.50 input / $0.25 cached input / $15.00 output per million tokens for short-context requests. Prompts above 272K input tokens use the published $5.00 input / $0.50 cached input / $22.50 output long-context tier; cache writes cost $3.125 per million tokens. Kimi K3: Moonshot AI's official Kimi K3 pricing page lists $0.30 cache-hit input, $3.00 cache-miss input, and $15.00 output per million tokens for model ID kimi-k3, with an exact 1,048,576-token context window. OpenRouter independently lists route moonshotai/kimi-k3 with one Moonshot AI INT4 endpoint at the same rates and context length. Prices exclude applicable taxes. Moonshot schedules the full weight release for July 27, 2026, so source type remains pending until the files are public. Claude Opus 4.7: Anthropic official pricing from the Claude Opus 4.7 announcement dated April 16, 2026. Anthropic states pricing remains the same as Opus 4.6: $5 input / $25 output per million tokens. Claude Opus 5: Anthropic's July 24, 2026 Claude Opus 5 announcement lists the claude-opus-5 API model at $5 input / $25 output per million tokens, the same base price as Opus 4.8. Fast mode runs around 2.5× the default speed and is available at twice the base price. Claude Opus 4.8: Anthropic's May 28, 2026 Claude Opus 4.8 launch announcement says Opus 4.8 is available at the same price as Opus 4.7. BenchLM maps that to the current Opus API rate of $5 input / $25 output per million tokens. Claude Opus 4.7 (Adaptive): Anthropic official pricing for Claude Opus 4.7 applies to the adaptive reasoning variant as the same API model with effort controls; launch announcement states pricing remains $5 input / $25 output per million tokens. GPT-5.6 Sol: OpenAI's current GPT-5.6 Sol model page lists $5.00 input / $0.50 cached input / $30.00 output per million tokens and a 1.05M-token context window. Prompts above 272K input tokens are charged at 2x input and 1.5x output for the full request; cache writes cost 1.25x uncached input. OpenAI's July 30, 2026 price update left Sol unchanged. GPT-5.5: OpenAI's current API pricing page lists gpt-5.5 at $5.00 input / $0.50 cached input / $30.00 output per million tokens for short-context requests. Batch and Flex are available at half the standard rate, and Priority is 2.5x the standard rate. Claude Fable 5.1: Anthropic's September 2026 launch page lists `claude-fable-5-1` at $10 input / $50 output per million tokens and cuts cache reads to $0.25 per million tokens. Anthropic estimates that lower cache-read rate reduces typical usage-based Fable workloads by about 25% and highly agentic workloads by as much as 45%; cache-write and Batch rates are not inferred from those estimates. GPT-6 Astra: OpenAI's GPT-6 Astra model page lists $10.00 input / $1.00 cached input / $12.50 cache writes / $50.00 output per million tokens, a 1,050,000-token context window, 922,000 maximum input tokens, and 128,000 maximum output tokens. Prompts above 272K input tokens are charged at 2x input and cache rates and 1.5x output for the full request; Batch and Flex are 50% of Standard and Fast mode is 2x. The September 3, 2026 price sheet lists Astra at 2.5x GPT-5.6 Sol's current $4.00 / $20.00 rates. Claude Mythos 5: Anthropic's June 9, 2026 Claude Fable 5 and Claude Mythos 5 launch says both models are offered at $10 input / $50 output per million tokens. Mythos 5 is restricted to Glasswing partners and trusted access programs. Claude Fable 5: Anthropic's June 9, 2026 Claude Fable 5 model page says Claude Fable 5 is available through the Claude API as `claude-fable-5` at $10 input / $50 output per million tokens, with a 90% input-token discount for prompt caching. US-only inference is available at 1.1x pricing. GPT-5.5 Pro: OpenAI's current API pricing page lists gpt-5.5-pro at $30.00 input / $180.00 output per million tokens for short-context requests and does not list a cached-input rate.
More tools
Start a prompt, improve an existing one, count tokens, or choose the right model for the workload.
How the AI cost calculator estimates a real workload
The calculator converts output such as blog posts, website pages, help articles, PRDs, and shipped features into estimated input and output tokens. It accounts for source material, formatting, review turns, and partial rewrites instead of requiring a token budget up front.
Reasoning-capable models can bill hidden thinking tokens in addition to visible output. If you already know the prompt size, count the tokens directly; for provider-level rates, compare current LLM API pricing.
How to estimate AI costs without understanding tokens
Most people do not budget AI work in tokens. They budget in outputs that matter to the business: how many blog posts need to go live, how many new landing pages the marketing team wants to publish, how many documentation articles support needs this quarter, or how many features the product team expects to ship. That is the gap this AI cost calculator is meant to close.
Instead of starting with input tokens per request and output tokens per request, this page starts with deliverables. You can model the cost of AI writing, AI-assisted product work, and AI content operations in units that mean something to a founder, marketer, operator, or product lead. Then the calculator translates those workloads into token estimates behind the scenes so the math still stays grounded in provider pricing.
That approach also makes planning conversations easier. When someone asks how much AI will cost to publish 20 blog posts per month or create 50 new SEO pages, you can answer in a monthly budget range instead of explaining why one million tokens is not actually a lot of work.
AI cost per blog post
A useful AI content cost calculator should not stop at word count. A 1,500-word blog post is not just 1,500 words of output. There is also the content brief, the SERP context, the competitor outline, internal linking instructions, editing rounds, and final revision prompts. That is why blog post cost is usually higher than a naive word-count estimate suggests.
If your team publishes frequently, the right question is cost per blog post multiplied by monthly volume. A model that is only a few cents more expensive per post can become a meaningful line item across dozens or hundreds of articles. On the other hand, a slightly pricier model can still be a bargain if it reduces editing overhead or produces better first drafts.
Use the blog-post preset when you want to estimate AI writing cost for editorial calendars, SEO programs, content agencies, or founder-led publishing. It is especially useful when comparing whether a budget model is good enough, or whether a higher-end model saves enough human review time to justify the premium.
AI cost per website page
Website page production often looks simple from the outside, but page creation usually includes more context than people expect. Messaging frameworks, brand voice, product differentiators, conversion goals, SEO targets, and revision feedback all add prompt volume. That is why a realistic cost to create website pages with AI is not just based on final page length.
For SEO teams, this page is useful as a website page cost calculator because it models publishing in batches. If you are planning 10, 25, or 100 pages in a month, you can estimate the monthly AI budget before you choose a model family. This is much easier to explain internally than talking about per-token billing.
The website-page preset is a good fit for landing pages, feature pages, industry pages, comparison pages, and template-driven SEO page programs. If your workflow includes heavier research or stricter brand review, increase the context and revision assumptions to reflect the real process.
AI cost per documentation article
Documentation and help-center content usually carry more context than marketing pages because accuracy matters. The model needs product behavior, release notes, setup instructions, troubleshooting context, and edge cases. Even when the final article is not especially long, the input side can be substantial.
This makes a documentation cost calculator particularly valuable for support-heavy products. The marginal cost of using a stronger model may be acceptable if it reduces inaccuracies, back-and-forth, or the need to rewrite technical steps manually. If your docs team is scaling, those tradeoffs matter more than raw per-token rates.
The docs preset is designed for API docs, knowledge-base articles, onboarding help, release notes, and support workflows. It is also a good approximation for internal enablement content when you want to price AI assistance at the team level.
AI cost per PRD or spec
PRDs, product briefs, technical specs, and research summaries usually need more context and more iteration than content marketing assets. Teams often paste customer feedback, bug reports, analytics notes, and engineering constraints into the prompt, then revise multiple times as decisions sharpen.
That means the right planning unit is cost per spec or cost per PRD, not just output length. A 2,000-word spec may still be cheap if the prompt is simple. A shorter document can be more expensive when it requires many rounds of structured reasoning, prioritization, and rewrite instructions.
The PRD/spec preset is useful for founders, product managers, and operators who want a fast sense of AI planning cost across a month or quarter. It can also help teams decide whether premium reasoning-heavy models belong in the planning workflow or only in narrower high-value moments.
AI cost to help ship product features
People increasingly ask a new kind of budgeting question: what does AI cost per shipped feature? That is not a token question and it is not a total engineering-cost question either. It is a workflow question. How many AI sessions does a team use while planning, coding, debugging, testing, and polishing one feature?
This calculator treats feature work as AI assistance spread across many sessions. Smaller features may need a handful of coding and debugging turns. Larger features may involve specification work, implementation help, test generation, bug fixing, migration advice, and release note drafting. The complexity presets give you a starting point without pretending the answer is universal.
If you run an AI-enabled engineering team, this framing is useful for forecasting budget, comparing model mixes, or explaining tooling cost to leadership. It is especially valuable when the user does not care about tokens at all and only wants to know how much AI support adds to the monthly software budget.
How words convert to tokens
Providers bill in tokens because that reflects how language models process text, but most planning discussions start in words. A practical planning assumption is that one English word is about 1.3 tokens. That ratio is not perfect for every language, formatting style, or code-heavy workflow, but it is a solid default for estimation.
The important part is visibility. This page shows the derived input and output tokens so you can see the math instead of treating the estimate like a black box. If you later want more precision, you can take the derived totals and jump straight into the advanced token-based calculator.
That makes this page useful both as an AI writing cost calculator for non-technical users and as a bridge into more advanced token planning for operators who eventually want tighter budget controls.
Which AI model is cheapest for content vs product work?
There is no single best model for every workload. Cheap models usually win when the job is repetitive, high-volume, and easy to review, like publishing large numbers of SEO pages or drafting straightforward support articles. Premium models often make more sense when mistakes are expensive, reasoning depth matters, or the human review loop is costly.
That is why this page compares multiple models side by side. It lets you see the spread between lower-cost options like GLM-5.1, Kimi K2.7 Code, Gemini 3 Flash, and DeepSeek V3 versus more premium options like GPT-5.5, GPT-5.5 Pro, Claude Opus 4.7, and Gemini 3.5 Flash. For some workloads the gap is small enough that quality should decide. For others the cheapest model can save thousands of dollars per month.
Use the current estimate as a scenario tool. If one model is 5x more expensive, ask whether it also removes enough human effort to be worth it. If not, the cheaper model may be the smarter operational choice even if it is not the benchmark leader.
Example monthly AI cost estimates
These example scenarios show how a human-friendly AI cost calculator becomes more useful than a token-only tool. Instead of explaining token budgets to every stakeholder, you can anchor the conversation in real monthly output.
| Workload | Derived Tokens | Cheapest In Table | GPT-5.5 |
|---|---|---|---|
| 20 blog posts / month | 927,810 in / 107,900 out | Kimi K2.7 Code — $1.46/mo | $9.01/mo |
| 50 website pages / month | 635,326 in / 118,755 out | Kimi K2.7 Code — $1.17/mo | $7.45/mo |
| 12 docs articles / month | 261,994 in / 36,881 out | Kimi K2.7 Code — $0.42/mo | $2.58/mo |
| 6 medium features / month | 107,640 in / 58,760 out | Kimi K2.7 Code — $0.43/mo | $3.01/mo |
FAQ
How do I estimate AI cost without understanding tokens?
Start with the workload you actually care about: blog posts, website pages, documentation articles, PRDs, or product features. This calculator converts those deliverables into estimated prompt and response tokens using words, revisions, and context assumptions, then applies model pricing.
How much does AI cost per blog post?
AI cost per blog post depends on draft length, revision rounds, and model choice. Lower-cost models like GLM-5.1 or Kimi K2.7 Code can be dramatically cheaper than premium frontier models for high-volume publishing, while GPT-5.5, Claude Opus 4.7, or Gemini 3.5 Flash may be worth the premium for harder briefs or more editorial polish.
How much does it cost to create website pages with AI?
Website page costs usually come from shorter output than blog posts but still include briefing, brand context, messaging constraints, and revision prompts. The right estimate is not cost per token in isolation, but cost per page at your real monthly publishing volume.
How do you estimate AI cost for documentation or help-center content?
Documentation often uses more input context than marketing copy because the model must ingest product behavior, source notes, release details, and support edge cases. That raises input token volume even if the published article is not especially long.
Can I estimate AI cost for PRDs and product specs?
Yes. PRDs and specs usually have higher context requirements and more revision loops than lightweight content. This page treats them as deliverables with larger briefs, larger context windows, and multiple rounds of refinement so you can budget for planning work, not just output length.
What does AI cost per feature really mean?
On this page, AI cost per feature means AI-assistance cost, not total engineering cost. It estimates how much model usage is involved in planning, coding, debugging, testing, and polishing a feature across many AI sessions.
How do words convert to tokens?
A practical rule of thumb is that one English word is roughly 1.3 tokens. This varies by language, formatting, and code, but it is a useful planning ratio for high-level budgeting. The calculator shows the derived token counts so you can inspect the assumptions.
Which AI models are cheapest for content production?
For raw cost efficiency, models like GLM-5.1, Kimi K2.7 Code, Gemini 3.1 Flash-Lite, and DeepSeek V3 are typically among the cheapest options in this calculator. They are often the best starting point when your workflow is high-volume and quality demands are moderate.
Which models are better for high-stakes product work?
For complex product work, many teams still evaluate premium models like GPT-5.5, Claude Opus 4.7, and Gemini 3.5 Flash because the output quality, reasoning, and iteration speed can offset a higher token bill. The best fit depends on how expensive mistakes are in your workflow.
Next steps
Use this page when the business question is about outcomes. Use the token calculator when the operations question is about prompts, completions, or request volume. For model selection, pair both with the quiz and the pricing directory.
Stay on top of pricing changes
Get notified when models update pricing, new models launch, or the cost landscape shifts.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.