LLM Token Cost, Context Window & Prompt Cache Calculator
Interactive pricing calculator comparing input context, output tokens, and prompt caching discounts across Claude 3.7, GPT-4o, Gemini 2.0, and DeepSeek.
LLM Token Pricing & Prompt Caching Cost Architecture
Compares multi-model pricing economics across Claude 3.7 Sonnet, Claude 3.5 Haiku, GPT-4o, GPT-4o-mini, Gemini 2.5 Flash, Gemini 2.5 Pro, and DeepSeek, modeling prompt caching write vs read discounts.
LLM Token Cost, Context Window & Prompt Cache Calculator Configurator
Customer Support RAG Agent — Prompt Caching & Monthly Cost Model
High-throughput customer support RAG assistant utilizing Claude 3.5 Sonnet prompt caching over 75,000 monthly requests.
1 Input Parameters & Assumptions
| Parameter | Value | Context & Provenance |
|---|---|---|
| LLM Frontier Model | Claude 3.5 Sonnet Model Tier | $3.00/M input, $15.00/M output, $0.30/M cache read |
| Prompt / Context Tokens | 3,500 tokens Tokens / Req | System prompt, retrieval context chunks, and conversation history |
| Completion Tokens | 650 tokens Tokens / Req | Generated agent response length |
| Prompt Cache Hit Rate | 80.0% Cache Rate | Percentage of prompt tokens served from static prompt prefix cache |
| Monthly Request Volume | 75,000 Requests / mo | Monthly production invocation volume |
2 Explicit Mathematical Formula
Model Rates: Claude 3.5 Sonnet (Input: $3.00/M, Output: $15.00/M, Cache Read: $0.30/M)
Uncached Input Cost = (3,500 × (1 - 0.80) / 1,000,000) × $3.00 = $0.002100
Cached Input Cost = (3,500 × 0.80 / 1,000,000) × $0.30 = $0.000840
Total Input Cost / Req = $0.002100 + $0.000840 = $0.002940
Output Cost / Req = (650 / 1,000,000) × $15.00 = $0.009750
Total Cost / Request = $0.002940 + $0.009750 = $0.012690
Uncached Baseline Cost / Req = (3,500 / 1M × $3) + $0.009750 = $0.010500 + $0.009750 = $0.020250
Monthly Spend (75k reqs) = 75,000 × $0.012690 = $951.75
Uncached Monthly Baseline = 75,000 × $0.020250 = $1,518.75
Monthly Net Savings = $1,518.75 - $951.75 = $567.00 (37.33% reduction)3 Computed Output Metrics
| Computed Metric | Result | Interpretation & Threshold |
|---|---|---|
| Blended Cost Per Request | $0.01269 USD / req | Total cost to process each customer inquiry taking prompt caching into account |
| Total Monthly Spend | $951.75 USD / month | Total monthly API expense for 75,000 production inquiries |
| Monthly Prompt Cache Savings | $567.00 USD / mo (37.33%) | Net dollars saved every month directly via prompt caching architecture |
| Uncached Monthly Baseline | $1,518.75 USD / month | Theoretical monthly spend if prompt caching were not implemented |
LLM Token Cost, Context Window & Prompt Cache Calculator — Scope & Limitations
Explicit operational boundaries and constraints defining target use cases and out-of-scope scenarios.
Built For (Target Use Cases)
- Estimating per-request and monthly API costs for Claude, GPT-4o, Gemini, and DeepSeek model tiers.
- Comparing prompt cache savings against uncached baseline spend using fixed provider rate cards.
- Sizing input, output, and cached-token mix for a given monthly request volume.
Not Built For (Limitations & Out-of-Scope)
- Live provider billing reconciliation; rates are hardcoded snapshots in LLM_MODELS, not a pricing API.
- Fine-tuning, batch API, or embedding endpoint pricing, none of which are modelled here.
- Token counts from a real tokenizer; inputs are user-entered estimates, not measured tokens.
Operational Assumptions & Defaults
- Cached prompt percentage is clamped to 0-100% before cost is calculated (clampedCachePct).
- Each preset in PRESET_WORKLOADS fixes a model, token mix, and request volume as a starting point.
- Model rate cards in LLM_MODELS are point-in-time snapshots and can drift from provider pricing.
LLM Token & Prompt Cache Architecture Cost Modeler
AI Systems & EconomicsCompare real token costs across Claude 3.7 Sonnet, Claude 3.5 Sonnet, GPT-4o, Gemini 2.0 Flash, and DeepSeek R1. Model prompt cache write vs read discounts and calculate true monthly unit economics.
Model & Workload Parameters
Cost Projection & Cache Savings
Multi-Model Cost Matrix
| Model | Provider | Monthly (Cached) | Cost / 1K | Monthly Savings |
|---|---|---|---|---|
| Claude 3.7 Sonnet | Anthropic | $951.75 | $12.690 | $567.00 |
| Claude 3.5 Sonnet | Anthropic | $951.75 | $12.690 | $567.00 |
| Claude 3.5 Haiku | Anthropic | $253.80 | $3.384 | $151.20 |
| GPT-4o | OpenAI | $881.25 | $11.750 | $262.50 |
| GPT-4o mini | OpenAI | $52.88 | $0.705 | $15.75 |
| Gemini 2.0 Flash | $30.00 | $0.400 | $15.75 | |
| Gemini 1.5 Pro | $375.00 | $5.000 | $196.88 | |
| DeepSeek V3 | DeepSeek | $23.94 | $0.319 | $26.46 |
| DeepSeek R1 | DeepSeek | $165.04 | $2.200 | $86.10 |
Export LLM Cost Comparison & Caching Strategy
Export your monthly token cost comparison across frontier models to Markdown or CSV, and configure prompt caching headers in your API requests.
How Prompt Caching Cuts LLM Application Costs
Modern frontier models support prompt caching on static system instructions and retrieval context. Anthropic offers up to a 90% discount on cached reads, while Google and OpenAI offer substantial discounts on repeated context. For applications with large system prompts, caching reduces operational spend significantly.
Implementation Code & Script
Calculates blended input token costs considering cached read rates versus full write rates.
export function calculateBlendedInputCost(
tokens: number,
cacheHitRatio: number,
baseRatePerMillion: number,
cacheReadRatePerMillion: number
): number {
const uncachedCost = (tokens * (1 - cacheHitRatio) * baseRatePerMillion) / 1_000_000;
const cachedCost = (tokens * cacheHitRatio * cacheReadRatePerMillion) / 1_000_000;
return Number((uncachedCost + cachedCost).toFixed(6));
}LLM Token Usage & Prompt Cache QA
Verify API response headers to confirm prompt caching hits (cache_read_input_tokens > 0) and evaluate cost reductions.
Pre-Production Verification Checklist
Verify repeated requests return cache_read_input_tokens > 0, reducing input cost by up to 90%.
Ensure cumulative conversation history does not exceed model context window or trigger truncation.
Compare expected output token lengths against actual response tokens across 50 test runs.
Terminal Diagnostic & Debug Commands
Validates API key and inspects usage.cache_read_input_tokens in returned response.
node -e 'console.log("Testing Anthropic caching headers...")'Failure Remediation & Troubleshooting
Cause: Dynamic parameters (e.g. timestamps) were placed above the cached system prompt breakpoint.
Fix: Keep all static system instructions and documentation at top of context before ephemeral user messages.
How to cite and attribute this tool
MIT LicenceThis resource is free, open and un-gated under the MIT Open Source Licence. You are encouraged to use, integrate and cite it with attribution:
@misc{geraghty_llm_token_cost_calculator,
author = {Geraghty, Gordon},
title = {LLM Token Cost, Context Window & Prompt Cache Calculator},
year = {2026},
url = {https://gordongeraghty.com/resources/ai-engineering/llm-token-cost-calculator},
note = {Head of Performance Media, Empire Amplify}
}Changelog & Version History
v1.0.0Initial release with multi-model prompt caching pricing engine.
Strategic Takeaway & Operational Guidelines
Prompt caching slashes input token pricing by 90% on Claude 3.5 Sonnet ($0.30/M vs $3.00/M). Structuring your system prompt and RAG context with static prefix blocks yields $567.00/mo in direct savings at 75k requests.