Introduction: Beyond Standard Microservices
Large Language Model (LLM) inference systems operate under unique constraints distinct from traditional microservices. Unlike stateless web APIs, LLM endpoints rely heavily on finite GPU memory for KV cache management and often maintain long-lived connections for streaming responses. This dependency on specialized hardware makes standard chaos engineering practices insufficient. We cannot simply kill a container; we must account for the state held in VRAM, the cost of re-initializing the model, and the latency spikes caused by GPU reset cycles.
In this post, we explore how to design chaos experiments that specifically target GPU failures and network partitions. The goal is to move beyond "it comes back up" to "it comes back up with acceptable latency and no data loss for streaming clients."
Scenario 1: Simulating Sudden GPU Failure
A common failure mode in production is a GPU dropping off the PCI bus due to thermal throttling or driver crashes. In Kubernetes environments running NVIDIA drivers, this often manifests as a node being cordon-ed off or the pod crashing with a `SIGBUS` error.
To test this, we need to simulate the loss of the compute device without necessarily restarting the entire node, as the orchestrator’s reaction time can mask the application’s internal error handling.
We can use `nvidia-smi` combined with system calls to simulate this, or more practically, inject a crash into the inference server process while the GPU is under load.
#!/bin/bash
# chaos_gpu_crash.sh
# Injects a SIGBUS into the inference process to simulate GPU access failure
PID=$(pgrep -f "vllm.engine" | head -n 1)
if [ -z "$PID" ]; then
echo "Inference process not found"
exit 1
fi
echo "Simulating GPU memory access failure on PID $PID"
# SIGBUS is often raised when memory is unmapped, simulating a VRAM access error
kill -BUS $PID
**Key Metrics to Monitor:**
* **Cold Start Latency:** How long does it take for the next request to be served by a healthy replica?
* **Client Retry Behavior:** Do your load balancers or clients retry immediately? If so, are you overwhelming the remaining healthy nodes with the backlog?
* **KV Cache Loss:** Confirm that any partial generation in flight is correctly terminated or resumed from a checkpoint if supported by your framework.
Scenario 2: Network Partitions and Straggler Nodes
In distributed inference setups (e.g., using DeepSpeed or vLLM’s tensor parallelism), network partition is critical. If one node in a tensor-parallel group loses connectivity, the entire inference batch fails. However, in a disaggregated architecture (Prefill vs. Decode separation), a partition between the front-end load balancer and the inference workers causes different symptoms.
Let’s simulate a network partition where the inference backend becomes unreachable but the API gateway remains up.
# Use tc (traffic control) to introduce 100% packet loss to the inference backend pod
# This simulates a network partition without dropping the node entirely
POD_IP=$(kubectl get pod -l app=inference-backend -o jsonpath='{.items[0].status.podIP}')
# Drop all traffic destined for the inference backend
kubectl exec -it [gateway-pod] -- tc qdisc add dev eth0 root netem loss 100%
# Wait for the chaos window
sleep 30
# Restore traffic
kubectl exec -it [gateway-pod] -- tc qdisc del dev eth0 root
**Expected Behavior:**
1. The API gateway should timeout requests within the configured `proxy_read_timeout` (e.g., 5s).
2. The gateway must fail-fast to alternative replicas if available.
3. **Crucially:** For streaming (SSE) responses, the client should detect the closed connection. Test if your client-side code correctly handles a mid-stream drop and re-queues the prompt, or if it silently fails, leading to user frustration.
Implementing Automated Chaos Tests in CI/CD
Manual chaos testing is not sustainable. Integrate these scenarios into your staging pipeline using tools like LitmusChaos or Chaos Mesh.
apiVersion: litmuschaos.github.io/v1alpha1
kind: ChaosExperiment
metadata:
name: llm-gpu-chaos
spec:
appinfo:
appkind: deployment
labelselector: "app=vllm-server"
chaosConfig:
chaosKind: "pod-kill"
# Simulate a hard kill to trigger node-level GPU recovery
mode: "all"
duration: "120s"
**Validation Strategy:**
Do not just check if the pod restarts. Validate the *quality* of service.
* **P99 Latency Check:** Assert that P99 latency does not exceed 2x baseline during the failure window.
* **Token Loss Check:** For streaming, verify that the total tokens received by the client match the expected count or that a clean error is returned, rather than a truncated, corrupt response.
Conclusion
LLM inference is not just another stateless service. Its statefulness in VRAM, its high memory footprint, and its long-running streaming nature demand a specialized approach to chaos engineering. By simulating GPU failures and network partitions, you can uncover critical gaps in your retry logic, load balancing strategies, and client-side error handling. Start small: inject a single GPU crash in staging. Monitor the blast radius. Iterate. Resilience is built through intentional failure.