Distributed Large Language Model (LLM) inference has transformed from a static batch processing task into a dynamic, stateful operation. While traditional web servers treat requests as independent units, modern LLMs rely heavily on the Key-Value (KV) cache to maintain context across multiple inference steps. When deployed in a distributed cluster, the placement of these requests becomes critical. If a subsequent request for a user’s conversation is routed to a different GPU node, the model must either recompute the entire KV cache from scratch or suffer severe latency penalties due to data transfer overhead. This is where stateful load balancing becomes an essential component of AI infrastructure.
The Problem with Stateless Routing
Standard load balancers, such as NGINX or HAProxy, typically operate on a stateless basis, using algorithms like Round Robin or Least Connections to distribute traffic. In the context of LLMs, this approach is suboptimal. The KV cache acts as a memory footprint that grows with the sequence length. Moving this state between GPUs is expensive. If we lose "affinity," we lose performance.
Consider a multi-turn chat application. The first prompt ("Tell me a story about a cat") is processed on GPU Node A. The KV cache for this prefix is stored in Node A's HBM memory. The user replies ("Now make it funny"). If a stateless load balancer routes this second request to GPU Node B, Node B has no knowledge of the previous context. It must either:
- Fetch the KV cache from Node A over the network (PCIe/NVLink or InfiniBand), which introduces significant bandwidth constraints.
- Recompute the prefill phase for the entire previous context, duplicating computational work.
Implementing Session Stickiness with KV Affinity
To solve this, we implement KV Cache Affinity. This strategy involves maintaining a mapping between user sessions (or unique request prefixes) and specific GPU workers. The load balancer consults this map before routing a new request. If a session is already assigned to a worker, the request is directed there. Only when a worker is saturated or unavailable does the balancer migrate the session to a new node, accepting the cost of state migration or recomputation as a necessary trade-off.
Architectural Considerations
An effective stateful balancer requires a centralized or distributed store (like Redis or etcd) to maintain the affinity map. The key design decisions involve:
- Key Definition: Using the session ID is straightforward but rigid. Using a hash of the prompt prefix offers better flexibility for semantic caching but increases computational overhead at the gate.
- Health Checks: The balancer must monitor GPU memory usage and temperature. If a worker holding a "hot" session approaches memory limits, the system must proactively evict the session or migrate it gracefully before an OOM (Out of Memory) error occurs.
- Load Shedding: When all nodes with affinity are overloaded, the system should prioritize new sessions to less utilized nodes rather than queuing indefinitely behind existing, high-affinity loads.
Code Example: Affinity-Based Routing Logic
Below is a conceptual Python implementation of a stateful load balancer that prioritizes affinity while respecting load limits.
import random
from collections import defaultdict
class StatefulLLMLoadBalancer:
def __init__(self, max_workers_per_session=1):
# Maps session_id -> worker_id
self.affinity_map = {}
# Maps worker_id -> list of active session_ids
self.worker_sessions = defaultdict(list)
# Simulated worker load (e.g., number of active requests or memory usage)
self.worker_load = defaultdict(int)
def assign_worker(self, session_id: str, current_loads: dict) -> str:
"""
Assigns a worker for a given session ID.
Prioritizes existing affinity if the worker is under the load threshold.
"""
threshold = 10 # Maximum concurrent sessions per worker before migration
# 1. Check if we have an existing affinity
if session_id in self.affinity_map:
current_worker = self.affinity_map[session_id]
# If the current worker is not overloaded, stick to it
if current_loads.get(current_worker, 0) < threshold:
return current_worker
# If overloaded, we must migrate. Remove from old worker context.
print(f"Warning: Worker {current_worker} overloaded. Migrating session {session_id}.")
self.worker_sessions[current_worker].remove(session_id)
# 2. Select a new worker
# Strategy: Least Loaded Worker among those not at capacity
available_workers = [
w for w, load in current_loads.items()
if load < threshold
]
if not available_workers:
raise Exception("System Capacity Full: No available workers under threshold.")
# Pick the worker with the lowest current load
new_worker = min(available_workers, key=lambda w: current_loads[w])
# 3. Update Affinity Maps
self.affinity_map[session_id] = new_worker
if session_id not in self.worker_sessions[new_worker]:
self.worker_sessions[new_worker].append(session_id)
return new_worker
def simulate_batch(self, sessions: list[str], initial_loads: dict):
"""Simulates routing a batch of requests."""
routing_decisions = []
for session in sessions:
worker = self.assign_worker(session, initial_loads)
initial_loads[worker] += 1
routing_decisions.append((session, worker))
return routing_decisions
# Usage Example
lb = StatefulLLMLoadBalancer()
# Simulate initial state: Worker 1 has 9 sessions, Worker 2 has 2
current_loads = {"worker_1": 9, "worker_2": 2}
# Session "abc" was previously assigned to worker_1 (simulated by pre-populating map)
lb.affinity_map["abc"] = "worker_1"
lb.worker_sessions["worker_1"].append("abc")
# New batch of incoming requests
new_requests = ["abc", "def", "ghi"]
print("Routing Decisions:")
for session, worker in lb.simulate_batch(new_requests, current_loads):
print(f"Session: {session:5s} -> Routed to: {worker}")
# Output Explanation:
# Session 'abc' should route to worker_1 if load < 10.
# Since worker_1 load is 9, it stays there. Load becomes 10.
# Session 'def' has no affinity, goes to least loaded (worker_2). Load becomes 3.
# Session 'ghi' has no affinity, goes to least loaded (worker_2). Load becomes 4.
Advanced Strategies: Semantic Affinity and Preemptive Migration
For high-scale deployments, simple session ID stickiness may not be sufficient. Advanced systems employ semantic affinity, where the load balancer hashes the initial few tokens of the prompt. If two different users start a conversation with the same system prompt (e.g., "You are a helpful coding assistant"), their KV caches for that prefix can be shared or kept on the same node. This is particularly effective for RAG (Retrieval-Augmented Generation) applications where context windows are large and static.
Furthermore, implementing preemptive migration is crucial. Instead of waiting for a worker to crash due to OOM, the balancer should monitor memory pressure. When a worker exceeds 85% memory utilization, it can signal the balancer to "cold start" new sessions elsewhere and, if possible, serialize and offload the KV caches of less active sessions to slower, larger capacity storage (like CPU RAM or NVMe SSDs) to free up GPU memory for high-priority requests.
Conclusion
Stateful load balancing is no longer a nice-to-have feature but a core requirement for efficient LLM infrastructure. By managing KV cache affinity and session stickiness, we can significantly reduce time-to-first-token (TTFT) and improve overall cluster utilization. The key lies in balancing the cost of state migration against the benefit of cache locality. As LLMs continue to grow in size and context window length, the sophistication of these routing algorithms will only increase, moving us from simple sticky sessions towards intelligent, semantic-aware resource orchestration. For developers building AI backends, integrating stateful logic early in the load balancing layer is critical to scaling effectively without incurring prohibitive computational costs.