As Large Language Models (LLMs) transition from experimental prototypes to production workhorses, the economics of inference have become a critical concern. While semantic caching—storing results based on vector similarity—has gained traction, it is not a silver bullet. Semantic matching introduces computational overhead and can sometimes miss exact matches where precision is paramount. For many enterprise use cases, deterministic strategies like Least Recently Used (LRU) and Time-To-Live (TTL) caching offer superior reliability and cost-efficiency. This post explores how to implement these strategies to slash inference costs without sacrificing service quality.
The Limitations of Purely Semantic Caching
Semantic caching relies on embedding models to determine if a previous response is "close enough" to a new query. While effective for exploratory tasks, it has two primary drawbacks for high-stakes applications: latency and cost. Computing embeddings for every incoming request adds a processing step that can negate the savings from skipping the LLM inference. Furthermore, in scenarios requiring exact data retrieval (e.g., customer support ticket lookups or financial reporting), approximate matches are unacceptable. Here, exact-key lookups via LRU or time-bound invalidation via TTL are far more appropriate.
Implementing LRU (Least Recently Used) Caching
LRU caching is ideal for repetitive query patterns. If your application frequently asks the same questions, storing the exact output keyed by the prompt allows subsequent requests to return instantly from memory or a fast in-memory store like Redis. The core principle is simple: when the cache reaches its capacity, the least recently accessed item is evicted.
In Python, you can implement a robust LRU cache using the `functools` library or integrate with Redis for distributed systems. Below is a practical example using a decorator-based approach for simplicity, which is often used in local microservices.
import time
from functools import lru_cache
# Setting maxsize to 128 means the 129th new item will evict the oldest used one
@lru_cache(maxsize=128)
def get_llm_response(query: str, model: str = "gpt-4") -> str:
"""
Simulates an LLM call. In production, this would call the API.
The decorator handles caching automatically.
"""
print(f"API Call made for: {query[:20]}...")
# Simulate network latency
time.sleep(1)
return f"AI Response to: {query}"
# First call - hits the API
print(get_llm_response("What is the capital of France?"))
# Second call - serves from cache
print(get_llm_response("What is the capital of France?"))
# Different query - hits the API again
print(get_llm_response("What is 2+2?"))
Strategic Use of TTL (Time-To-Live)
LLM outputs are often context-dependent or time-sensitive. A fact valid yesterday may be incorrect today. TTL caching addresses this by associating each cache entry with an expiration timer. This strategy is particularly useful for news summaries, real-time market analysis, or dynamic FAQs.
When implementing TTL, especially in a distributed environment using Redis, you should configure the expiration time based on the volatility of the data. For static documentation queries, a TTL of 24 hours might be sufficient. For real-time stock analysis, a TTL of 60 seconds may be necessary.
import redis
import json
# Connect to Redis
r = redis.Redis(host='localhost', port=6379, db=0)
def get_cached_llm_response(query: str, ttl_seconds=3600):
cache_key = f"llm:{query}"
# Try to get from cache
cached_response = r.get(cache_key)
if cached_response:
print("Cache Hit")
return json.loads(cached_response)
# Simulate LLM call
print("Cache Miss - Calling LLM")
# response = call_llm_api(query)
response = f"Result for: {query}"
# Store in cache with TTL
r.setex(cache_key, ttl_seconds, json.dumps(response))
return response
# This will expire after 1 hour
get_cached_llm_response("Current weather in London")
Hybrid Approach for LLMOps
The most cost-effective strategy often combines these techniques. You can use exact LRU caching for deterministic queries and semantic caching for creative or open-ended tasks. Additionally, layering TTL on top of LRU ensures that even exact matches don't serve stale data indefinitely. By carefully tuning cache sizes and expiration policies, engineering teams can reduce LLM API costs by up to 40-60% while maintaining a responsive user experience.
Conclusion
Semantic caching is a powerful tool, but it is not the only solution for optimizing LLM inference. Implementing robust LRU and TTL strategies provides a deterministic, low-latency, and highly cost-effective alternative for many production workloads. As the field of LLMOps matures, the ability to choose the right caching mechanism for the right use case will distinguish high-performing applications from those struggling with ballooning API bills.