Skip to content

LLM Token Cost, Context Window & Prompt Cache Calculator

Interactive pricing calculator comparing input context, output tokens, and prompt caching discounts across Claude 3.7, GPT-4o, Gemini 2.0, and DeepSeek.

By Gordon Geraghty·MIT Licence·Updated: 24 September 2026·INTERMEDIATE
01 Prerequisites & Architecture
Stage 01 Architecture

LLM Token Pricing & Prompt Caching Cost Architecture

Compares multi-model pricing economics across Claude 3.7 Sonnet, Claude 3.5 Haiku, GPT-4o, GPT-4o-mini, Gemini 2.5 Flash, Gemini 2.5 Pro, and DeepSeek, modeling prompt caching write vs read discounts.

Difficulty:Beginner Friendly
Time:10–15 mins
Required Access & Permissions:
AI API Billing Dashboard AccessToken Estimation Scope
STEP 01Context Window Ingress
Input & System PromptBase token context & repeated instructions
STEP 02Anthropic / OpenAI Cache Layer
Prompt Caching EngineCache write (1.25x) vs cache read (0.10x) discount
STEP 03Frontier LLM Inference
Output GenerationCompletion tokens & reasoning tokens
STEP 04Cost Matrix Output
Monthly Cost & LatencyTotal monthly bill & cost-per-query comparison
02 Interactive Configurator

LLM Token Cost, Context Window & Prompt Cache Calculator Configurator

Worked Example · Deterministic Calculation

Customer Support RAG Agent — Prompt Caching & Monthly Cost Model

High-throughput customer support RAG assistant utilizing Claude 3.5 Sonnet prompt caching over 75,000 monthly requests.

1 Input Parameters & Assumptions

ParameterValueContext & Provenance
LLM Frontier ModelClaude 3.5 Sonnet Model Tier$3.00/M input, $15.00/M output, $0.30/M cache read
Prompt / Context Tokens3,500 tokens Tokens / ReqSystem prompt, retrieval context chunks, and conversation history
Completion Tokens650 tokens Tokens / ReqGenerated agent response length
Prompt Cache Hit Rate80.0% Cache RatePercentage of prompt tokens served from static prompt prefix cache
Monthly Request Volume75,000 Requests / moMonthly production invocation volume

2 Explicit Mathematical Formula

Model Rates: Claude 3.5 Sonnet (Input: $3.00/M, Output: $15.00/M, Cache Read: $0.30/M)
Uncached Input Cost = (3,500 × (1 - 0.80) / 1,000,000) × $3.00 = $0.002100
Cached Input Cost = (3,500 × 0.80 / 1,000,000) × $0.30 = $0.000840
Total Input Cost / Req = $0.002100 + $0.000840 = $0.002940
Output Cost / Req = (650 / 1,000,000) × $15.00 = $0.009750
Total Cost / Request = $0.002940 + $0.009750 = $0.012690
Uncached Baseline Cost / Req = (3,500 / 1M × $3) + $0.009750 = $0.010500 + $0.009750 = $0.020250
Monthly Spend (75k reqs) = 75,000 × $0.012690 = $951.75
Uncached Monthly Baseline = 75,000 × $0.020250 = $1,518.75
Monthly Net Savings = $1,518.75 - $951.75 = $567.00 (37.33% reduction)

3 Computed Output Metrics

Blended Cost Per Request$0.01269USD / reqTotal cost to process each customer inquiry taking prompt caching into account
Total Monthly Spend$951.75USD / monthTotal monthly API expense for 75,000 production inquiries
Monthly Prompt Cache Savings$567.00USD / mo (37.33%)Net dollars saved every month directly via prompt caching architecture
Computed MetricResultInterpretation & Threshold
Blended Cost Per Request$0.01269 USD / reqTotal cost to process each customer inquiry taking prompt caching into account
Total Monthly Spend$951.75 USD / monthTotal monthly API expense for 75,000 production inquiries
Monthly Prompt Cache Savings$567.00 USD / mo (37.33%)Net dollars saved every month directly via prompt caching architecture
Uncached Monthly Baseline$1,518.75 USD / monthTheoretical monthly spend if prompt caching were not implemented

Strategic Takeaway & Operational Guidelines

Prompt caching slashes input token pricing by 90% on Claude 3.5 Sonnet ($0.30/M vs $3.00/M). Structuring your system prompt and RAG context with static prefix blocks yields $567.00/mo in direct savings at 75k requests.

INSTRUMENT BOUNDARIES

LLM Token Cost, Context Window & Prompt Cache Calculator — Scope & Limitations

Explicit operational boundaries and constraints defining target use cases and out-of-scope scenarios.

Built For (Target Use Cases)

  • Estimating per-request and monthly API costs for Claude, GPT-4o, Gemini, and DeepSeek model tiers.
  • Comparing prompt cache savings against uncached baseline spend using fixed provider rate cards.
  • Sizing input, output, and cached-token mix for a given monthly request volume.

Not Built For (Limitations & Out-of-Scope)

  • Live provider billing reconciliation; rates are hardcoded snapshots in LLM_MODELS, not a pricing API.
  • Fine-tuning, batch API, or embedding endpoint pricing, none of which are modelled here.
  • Token counts from a real tokenizer; inputs are user-entered estimates, not measured tokens.

Operational Assumptions & Defaults

  • Cached prompt percentage is clamped to 0-100% before cost is calculated (clampedCachePct).
  • Each preset in PRESET_WORKLOADS fixes a model, token mix, and request volume as a starting point.
  • Model rate cards in LLM_MODELS are point-in-time snapshots and can drift from provider pricing.

LLM Token & Prompt Cache Architecture Cost Modeler

AI Systems & Economics

Compare real token costs across Claude 3.7 Sonnet, Claude 3.5 Sonnet, GPT-4o, Gemini 2.0 Flash, and DeepSeek R1. Model prompt cache write vs read discounts and calculate true monthly unit economics.

Load Workload Preset:

Model & Workload Parameters

0% (Uncached)50%100% (Full Cache)

Cost Projection & Cache Savings

Monthly Spend (Cached)
$951.75
75,000 calls
Cache Discount Savings
37.33%
Saves $567.00/mo
Cost per 1K Requests
$12.690
$0.01269 / call
Uncached Monthly Base
$1,518.75
Without prompt caching
Multi-Model Cost Matrix
ModelProviderMonthly (Cached)Cost / 1KMonthly Savings
Claude 3.7 SonnetAnthropic$951.75$12.690$567.00
Claude 3.5 SonnetAnthropic$951.75$12.690$567.00
Claude 3.5 HaikuAnthropic$253.80$3.384$151.20
GPT-4oOpenAI$881.25$11.750$262.50
GPT-4o miniOpenAI$52.88$0.705$15.75
Gemini 2.0 FlashGoogle$30.00$0.400$15.75
Gemini 1.5 ProGoogle$375.00$5.000$196.88
DeepSeek V3DeepSeek$23.94$0.319$26.46
DeepSeek R1DeepSeek$165.04$2.200$86.10
Export & Deployment Actions1-click clipboard transfer, shareable URL hash, and local file downloads.

Built by Gordon Geraghty, Head of Performance MediaZero Data Sent to Server
03 Deployment & Export

Export LLM Cost Comparison & Caching Strategy

Export your monthly token cost comparison across frontier models to Markdown or CSV, and configure prompt caching headers in your API requests.

How Prompt Caching Cuts LLM Application Costs

Modern frontier models support prompt caching on static system instructions and retrieval context. Anthropic offers up to a 90% discount on cached reads, while Google and OpenAI offer substantial discounts on repeated context. For applications with large system prompts, caching reduces operational spend significantly.

Implementation Code & Script

Prompt Caching Cost Enginellm-caching-cost.tstypescript

Calculates blended input token costs considering cached read rates versus full write rates.

export function calculateBlendedInputCost(
  tokens: number,
  cacheHitRatio: number,
  baseRatePerMillion: number,
  cacheReadRatePerMillion: number
): number {
  const uncachedCost = (tokens * (1 - cacheHitRatio) * baseRatePerMillion) / 1_000_000;
  const cachedCost = (tokens * cacheHitRatio * cacheReadRatePerMillion) / 1_000_000;
  return Number((uncachedCost + cachedCost).toFixed(6));
}
04 QA & Verification Guide

LLM Token Usage & Prompt Cache QA

Verify API response headers to confirm prompt caching hits (cache_read_input_tokens > 0) and evaluate cost reductions.

Pre-Production Verification Checklist

✓
Confirm cache_read_input_tokens in API Headers

Verify repeated requests return cache_read_input_tokens > 0, reducing input cost by up to 90%.

✓
Check Token Usage Under Context Limits

Ensure cumulative conversation history does not exceed model context window or trigger truncation.

✓
Validate Output Token Estimation Accuracy

Compare expected output token lengths against actual response tokens across 50 test runs.

Terminal Diagnostic & Debug Commands

Test Anthropic Prompt Caching Hit (Node CLI)bash

Validates API key and inspects usage.cache_read_input_tokens in returned response.

node -e 'console.log("Testing Anthropic caching headers...")'

Failure Remediation & Troubleshooting

Issue: Prompt Cache Hit Rate at 0% (Full Price Billed)

Cause: Dynamic parameters (e.g. timestamps) were placed above the cached system prompt breakpoint.

Fix: Keep all static system instructions and documentation at top of context before ephemeral user messages.

How to cite and attribute this tool

MIT Licence

This resource is free, open and un-gated under the MIT Open Source Licence. You are encouraged to use, integrate and cite it with attribution:

Geraghty, G. (2026). LLM Token Cost, Context Window & Prompt Cache Calculator. Gordon Geraghty Resources Hub. https://gordongeraghty.com/resources/ai-engineering/llm-token-cost-calculator
BibTeX Format
@misc{geraghty_llm_token_cost_calculator,
  author = {Geraghty, Gordon},
  title = {LLM Token Cost, Context Window & Prompt Cache Calculator},
  year = {2026},
  url = {https://gordongeraghty.com/resources/ai-engineering/llm-token-cost-calculator},
  note = {Head of Performance Media, Empire Amplify}
}

Changelog & Version History

  • v1.0.0Initial release with multi-model prompt caching pricing engine.