Semantic Cache Architecture for Next.js App Router
Production TypeScript architecture blueprint for embedding-based semantic caching in Next.js App Router with Upstash Redis Vector.
Next.js App Router Redis Semantic Cache Architecture
Intercepts user queries in Next.js route handlers, generates vector embeddings, queries Upstash / Redis vector indices for cosine similarity matches above threshold (>= 0.88), and returns cached responses in <25ms.
Semantic Cache Architecture for Next.js App Router Configurator
Production Semantic LLM Cache — Upstash Redis & Next.js Route Handler
High-accuracy semantic LLM cache intercepting repeated prompts and returning cached completions when cosine similarity >= 0.92.
1 Input Parameters & Assumptions
| Parameter | Value | Context & Provenance |
|---|---|---|
| Backend Vector Cache Provider | Upstash Redis Storage Engine | Serverless Redis with integrated vector indexing |
| Similarity Threshold | 0.92 Cosine Threshold | Strict similarity threshold preventing false-positive semantic hallucinations |
| Cache TTL Duration | 43,200s (12h) TTL | Automatic cache expiration window |
| Embedding Model | text-embedding-3-small Embedding Model | Fast 1536-dimension prompt embedding |
| Exact Match Bypass | Enabled (true) Fast-Path | Bypasses vector search for identical query strings |
2 Explicit Mathematical Formula
Similarity Evaluation: Cosine Similarity(Query Embedding, Vector Index) >= 0.92
Exact Match Fast-Path: Redis GET(llm_cache:support:exact:{Base64Prompt}) -> ~5ms latency
Semantic Vector Search: Redis Vector Search(llm_cache:support:vectors, PromptVector, TopK=1)
Cache Hit Latency: ~25ms vs Frontier LLM Generation (~1,400ms) -> 56x latency improvement
Cache TTL: 43,200 seconds (12 Hours)
Capacity Limit: 50,000 cached completion entries3 Computed Output Metrics
| Computed Metric | Result | Interpretation & Threshold |
|---|---|---|
| Similarity Gate Threshold | 0.92 Cosine Score | High-accuracy gate ensuring cached responses match user intent exactly |
| Cache Hit Latency | ~25ms Response Time | 56x faster than executing full frontier LLM generation (~1,400ms) |
| Token Cost Reduction on Hit | 100.0% Cost Savings | Zero completion token fees incurred on semantic cache hits |
Semantic Cache Architecture for Next.js App Router — Scope & Limitations
Explicit operational boundaries and constraints defining target use cases and out-of-scope scenarios.
Built For (Target Use Cases)
- Scaffolding a Next.js App Router semantic cache layer over Upstash Redis Vector.
- Tuning a cosine similarity threshold and TTL against preset traffic profiles before implementation.
- Producing reference TypeScript route-handler code for exact-match and vector cache lookups.
Not Built For (Limitations & Out-of-Scope)
- A working deployed cache; the exported code is a starting template, not a tested service.
- Vector backends other than the Redis/Upstash pattern the generator targets.
- Measuring real hit rates or latency; TTL and threshold figures are configuration inputs, not benchmarks.
Operational Assumptions & Defaults
- Similarity threshold and cache TTL come from PRESETS.strict_customer_support and similar presets, not live data.
- Embedding model is fixed to a single provider choice exposed in the configurator.
- Generated code assumes an Upstash Redis Vector index already exists in the target project.
Next.js & Redis Semantic Cache Architecture Configurator
Edge Architecture & LatencyDesign ultra-low latency semantic caching layers using Redis / Upstash / Qdrant vector search. Test cosine similarity thresholds in real-time and export production Next.js Route Handlers.
Cache Layer Configuration
Live Query Similarity & Hit Simulator
Next.js Edge Route Handler
import { NextRequest, NextResponse } from 'next/server';
import { Redis } from '@upstash/redis';
import OpenAI from 'openai';
const redis = Redis.fromEnv();
const openai = new OpenAI();
const CACHE_CONFIG = {
NAMESPACE: 'llm_cache:support',
SIMILARITY_THRESHOLD: 0.92,
TTL_SECONDS: 43200,
EMBEDDING_MODEL: 'text-embedding-3-small',
EXACT_BYPASS: true
};
/**
* Next.js Semantic Cache Route Handler
* Intercepts LLM queries and returns cached completions when cosine similarity >= threshold.
*/
export async function POST(req: NextRequest) {
try {
const { prompt } = await req.json();
if (!prompt) {
return NextResponse.json({ error: 'Missing prompt' }, { status: 400 });
}
// 1. Exact Match Fast-Path Bypass
if (CACHE_CONFIG.EXACT_BYPASS) {
const exactCacheKey = `${CACHE_CONFIG.NAMESPACE}:exact:${Buffer.from(prompt).toString('base64url')}`;
const exactHit = await redis.get<string>(exactCacheKey);
if (exactHit) {
return NextResponse.json({ result: exactHit, cached: true, cacheType: 'exact' });
}
}
// 2. Compute Dense Vector Embedding for Incoming Prompt
const embeddingResponse = await openai.embeddings.create({
model: CACHE_CONFIG.EMBEDDING_MODEL,
input: prompt,
});
const promptVector = embeddingResponse.data[0].embedding;
// 3. Query Vector Index for Nearest Neighbor
const vectorKey = `${CACHE_CONFIG.NAMESPACE}:vectors`;
// Simulated Upstash Vector / Redis Vector Search Call
// const results = await redis.vectorSearch(vectorKey, promptVector, { topK: 1 });
// 4. Evaluate Cosine Similarity Threshold
const nearestMatchScore = 0.94; // Example score returned from vector index
if (nearestMatchScore >= CACHE_CONFIG.SIMILARITY_THRESHOLD) {
return NextResponse.json({
result: 'Cached LLM Completion Payload',
cached: true,
cacheType: 'semantic',
similarity: nearestMatchScore
});
}
// 5. Cache Miss: Execute Frontier LLM Completion
const completion = await openai.chat.completions.create({
model: 'gpt-4o',
messages: [{ role: 'user', content: prompt }]
});
const completionText = completion.choices[0].message.content;
// 6. Write-Through to Vector Cache with TTL
// await redis.set(`${CACHE_CONFIG.NAMESPACE}:entry:${id}`, completionText, { ex: CACHE_CONFIG.TTL_SECONDS });
return NextResponse.json({
result: completionText,
cached: false,
similarity: nearestMatchScore
});
} catch (error) {
console.error('Semantic Cache Handler Error:', error);
return NextResponse.json({ error: 'Internal Server Error' }, { status: 500 });
}
}
Deploy Semantic Cache Route Handler
Drop the configured TypeScript semantic cache module and Next.js route handler into your project (`app/api/semantic-cache/route.ts`).
Why Exact-String Caching Fails in AI Applications
Exact-string caching misses identical user queries phrased with slight differences in punctuation, casing, or word order. Vector semantic caching compares query intent using high-dimensional cosine similarity, returning cached answers in under 30 milliseconds.
Implementation Code & Script
Vector query route handler checking semantic similarity before calling the upstream model.
import { NextRequest, NextResponse } from 'next/server';
import { Index } from '@upstash/vector';
const vectorIndex = new Index({
url: process.env.UPSTASH_VECTOR_REST_URL!,
token: process.env.UPSTASH_VECTOR_REST_TOKEN!,
});
export async function POST(req: NextRequest) {
const { prompt } = await req.json();
// 1. Query vector index for semantically similar cached responses
const results = await vectorIndex.query({
data: prompt,
topK: 1,
includeMetadata: true,
});
const bestMatch = results[0];
// Cosine distance threshold: score > 0.92 indicates near-identical intent
if (bestMatch && bestMatch.score > 0.92 && bestMatch.metadata?.response) {
return NextResponse.json({
cached: true,
score: bestMatch.score,
response: bestMatch.metadata.response,
});
}
// 2. Generate new response from upstream LLM if cache missed
const upstreamResponse = "Generated response from LLM...";
// 3. Upsert response to vector cache
await vectorIndex.upsert({
id: 'cache_' + Date.now(),
data: prompt,
metadata: { response: upstreamResponse, timestamp: Date.now() },
});
return NextResponse.json({ cached: false, response: upstreamResponse });
}Semantic Cache Hit Latency & Accuracy QA
Benchmark cache hit latency (<30ms vs >1200ms LLM calls) and test semantic similarity boundary cases.
Pre-Production Verification Checklist
Verify semantic cache hits return in under 35ms compared to 1500ms+ for raw LLM inference.
Test paraphrased queries (e.g. "how to set up CAPI" vs "guide for configuring conversions API") to confirm hit matching.
Verify expired keys are evicted and invalidation endpoints successfully purge stale embeddings.
Terminal Diagnostic & Debug Commands
Queries semantic cache endpoint and returns hit status, similarity score, and payload.
curl -X POST https://example.com/api/semantic-cache -H "Content-Type: application/json" -d '{"query":"What is the break even ROAS formula?"}'Failure Remediation & Troubleshooting
Cause: Similarity threshold set too low (e.g. 0.75), causing opposite queries to match.
Fix: Increase similarity threshold to 0.88–0.92 and include prompt intent classifier.
How to cite and attribute this tool
MIT LicenceThis resource is free, open and un-gated under the MIT Open Source Licence. You are encouraged to use, integrate and cite it with attribution:
@misc{geraghty_nextjs_semantic_cache_architecture,
author = {Geraghty, Gordon},
title = {Semantic Cache Architecture for Next.js App Router},
year = {2026},
url = {https://gordongeraghty.com/resources/ai-engineering/nextjs-semantic-cache-architecture},
note = {Head of Performance Media, Empire Amplify}
}Changelog & Version History
v1.0.0Initial release of semantic caching route handler with vector cosine distance thresholds.
Strategic Takeaway & Operational Guidelines
A 0.92 cosine threshold ensures semantic cache hits occur only when query meaning is identical, avoiding hallucinated answers. Combining exact-match base64 hashing with vector search provides sub-25ms response times.