When we build GenAI applications, we often focus on the LLM model and prompts.

But in production, another question becomes critical:

Why repeatedly process the same information?

That’s where caching strategies become extremely useful.

πŸ”Ή 1. Without caching

Imagine an application receives:

β€œWhat is the claim status for order-101?”

with a large system prompt, instructions, examples and context.

Every request may require the system to process a significant amount of the same information again.

For example:

Input: 1,000 tokens
Output: 500 tokens
Total processed: ~1,500 tokens

At scale, repeated processing can increase:

  • πŸ’° LLM cost
  • ⏱️ Latency
  • πŸ–₯️ Model workload
  • πŸ“ˆ Infrastructure requirements

πŸ”Ή 2. Prompt Caching

Idea: Reuse the already-processed portion of a repeated prompt.

Useful when applications repeatedly send:

  • System instructions
  • Long policies
  • Few-shot examples
  • Fixed RAG instructions
  • Agent instructions

Instead of repeatedly processing the same prefix, the cached portion can be reused.

Best fit:
Chatbots, agents, RAG applications and applications with large static prompts.


πŸ”Ή 3. KV Cache

KV Cache works inside autoregressive generation.

During generation, the Transformer calculates Key (K) and Value (V) representations for previous tokens.

Instead of recomputing them for every newly generated token, the model keeps them in the KV cache.

Conceptually:

Without KV Cache

Previous tokens β†’ recompute attention repeatedly β†’ next token

With KV Cache

Previous K/V β†’ reuse β†’ process only what is newly needed β†’ next token

This is particularly important for:

  • Long responses
  • Long conversations
  • Large-context applications
  • Agentic workflows

⚑ Main benefit: lower generation latency and compute.


πŸ”Ή 4. Semantic Cache

This operates at the application level.

Instead of checking only whether the new question is exactly identical, convert the question into an embedding and search for a semantically similar previous question.

Example:

Previous:

β€œWhere is my order?”

New:

β€œCan you tell me my order status?”

These aren’t identical strings, but they may have the same intent.

If the similarity is above a carefully selected threshold:

Query β†’ Embedding β†’ Semantic Cache β†’ Cached response

The application can potentially avoid calling the LLM altogether.

⚑ Main benefit: avoid unnecessary model calls.


πŸ”₯ 5. So where does the saving actually happen?

Think about the three caches differently:

CacheReusesMain benefit
Prompt CacheRepeated prompt/context processingLower input processing cost/latency, depending on provider
KV CachePrevious token attention state during generationFaster generation + lower compute
Semantic CachePrevious answers/results for similar queriesPotentially eliminates an LLM call

The important point:

These are not three versions of the same cache. They operate at different layers of the architecture.


πŸ“Š 6. A simple production example

Suppose an application receives 10,000 requests.

Without caching, many requests repeatedly process the same:

  • System prompt
  • Instructions
  • Examples
  • Context
  • Similar questions

With appropriate caching:

Prompt Cache
β†’ reuse repeated prompt prefixes

KV Cache
β†’ reuse previous-token attention state during generation

Semantic Cache
β†’ return an existing result for sufficiently similar requests

The result can be:

↓ Model calls
↓ Tokens requiring fresh processing
↓ Latency
↓ Compute
↓ Cost

But the exact token/cost reduction depends heavily on the model provider, pricing model, cache hit rate, prompt structure, similarity threshold and workload.


🎯 7. The key architecture idea

Don’t think:

β€œHow do I make the LLM faster?”

Think:

β€œWhat work does my system keep doing repeatedly that doesn’t need to be done again?”

Then place the cache at the appropriate layer.

                    USER REQUEST
                         β”‚
                         β–Ό
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚ Semantic Cacheβ”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚ miss
                         β–Ό
                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚ Prompt Cache  β”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
                    LLM / Transformer
                         β”‚
                    β”Œβ”€β”€β”€β”€β”΄β”€β”€β”€β”€β”
                    β”‚ KV Cacheβ”‚
                    β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜
                         β”‚
                         β–Ό
                      RESPONSE

πŸ’‘ The takeaway

Semantic Cache asks:
β€œHave I already answered something essentially equivalent?”

Prompt Cache asks:
β€œHave I already processed this repeated prompt/context?”

KV Cache asks:
β€œHave I already computed attention state for these previous tokens?”

Different problems. Different layers. Different benefits.

And that’s why caching should be treated as an architectural capabilityβ€”not just an optimization added at the end.

#GenerativeAI #LLM #AIEngineering #AIAgents #RAG #LLMOps #PromptEngineering #SemanticCache #KVCache #PromptCaching #GenAI #MachineLearning