When we build GenAI applications, we often focus on the LLM model and prompts.
But in production, another question becomes critical:
Why repeatedly process the same information?
Thatβs where caching strategies become extremely useful.
πΉ 1. Without caching
Imagine an application receives:
βWhat is the claim status for order-101?β
with a large system prompt, instructions, examples and context.
Every request may require the system to process a significant amount of the same information again.
For example:
Input: 1,000 tokens
Output: 500 tokens
Total processed: ~1,500 tokens
At scale, repeated processing can increase:
- π° LLM cost
- β±οΈ Latency
- π₯οΈ Model workload
- π Infrastructure requirements
πΉ 2. Prompt Caching
Idea: Reuse the already-processed portion of a repeated prompt.
Useful when applications repeatedly send:
- System instructions
- Long policies
- Few-shot examples
- Fixed RAG instructions
- Agent instructions
Instead of repeatedly processing the same prefix, the cached portion can be reused.
Best fit:
Chatbots, agents, RAG applications and applications with large static prompts.
πΉ 3. KV Cache
KV Cache works inside autoregressive generation.
During generation, the Transformer calculates Key (K) and Value (V) representations for previous tokens.
Instead of recomputing them for every newly generated token, the model keeps them in the KV cache.
Conceptually:
Without KV Cache
Previous tokens β recompute attention repeatedly β next token
With KV Cache
Previous K/V β reuse β process only what is newly needed β next token
This is particularly important for:
- Long responses
- Long conversations
- Large-context applications
- Agentic workflows
β‘ Main benefit: lower generation latency and compute.
πΉ 4. Semantic Cache
This operates at the application level.
Instead of checking only whether the new question is exactly identical, convert the question into an embedding and search for a semantically similar previous question.
Example:
Previous:
βWhere is my order?β
New:
βCan you tell me my order status?β
These aren’t identical strings, but they may have the same intent.
If the similarity is above a carefully selected threshold:
Query β Embedding β Semantic Cache β Cached response
The application can potentially avoid calling the LLM altogether.
β‘ Main benefit: avoid unnecessary model calls.
π₯ 5. So where does the saving actually happen?
Think about the three caches differently:
| Cache | Reuses | Main benefit |
|---|---|---|
| Prompt Cache | Repeated prompt/context processing | Lower input processing cost/latency, depending on provider |
| KV Cache | Previous token attention state during generation | Faster generation + lower compute |
| Semantic Cache | Previous answers/results for similar queries | Potentially eliminates an LLM call |
The important point:
These are not three versions of the same cache. They operate at different layers of the architecture.
π 6. A simple production example
Suppose an application receives 10,000 requests.
Without caching, many requests repeatedly process the same:
- System prompt
- Instructions
- Examples
- Context
- Similar questions
With appropriate caching:
Prompt Cache
β reuse repeated prompt prefixes
KV Cache
β reuse previous-token attention state during generation
Semantic Cache
β return an existing result for sufficiently similar requests
The result can be:
β Model calls
β Tokens requiring fresh processing
β Latency
β Compute
β Cost
But the exact token/cost reduction depends heavily on the model provider, pricing model, cache hit rate, prompt structure, similarity threshold and workload.
π― 7. The key architecture idea
Don’t think:
βHow do I make the LLM faster?β
Think:
βWhat work does my system keep doing repeatedly that doesn’t need to be done again?β
Then place the cache at the appropriate layer.
USER REQUEST
β
βΌ
βββββββββββββββββ
β Semantic Cacheβ
βββββββββ¬ββββββββ
β miss
βΌ
βββββββββββββββββ
β Prompt Cache β
βββββββββ¬ββββββββ
β
βΌ
LLM / Transformer
β
ββββββ΄βββββ
β KV Cacheβ
ββββββ¬βββββ
β
βΌ
RESPONSE
π‘ The takeaway
Semantic Cache asks:
βHave I already answered something essentially equivalent?β
Prompt Cache asks:
βHave I already processed this repeated prompt/context?β
KV Cache asks:
βHave I already computed attention state for these previous tokens?β
Different problems. Different layers. Different benefits.
And that’s why caching should be treated as an architectural capabilityβnot just an optimization added at the end.
#GenerativeAI #LLM #AIEngineering #AIAgents #RAG #LLMOps #PromptEngineering #SemanticCache #KVCache #PromptCaching #GenAI #MachineLearning
