Model Context Protocol (MCP)

Building Resilient MCP Clients: Mastering Circuit Breakers and Retry Logic

The Model Context Protocol (MCP) represents a paradigm shift in how AI agents interact with external tools and data sources. As we deploy these agents across distributed systems, the reliability of remote MCP server connections becomes critical. A single transient network glitch or temporary server overload can cascade into a complete agent failure if not handled correctly. To build production-grade MCP applications, we must move beyond simple "try-catch" blocks and implement sophisticated resilience patterns like circuit breakers and intelligent retry logic.

Why Resilience Matters in MCP

In a standard MCP architecture, the client (often the LLM agent) communicates with a server (the tool provider) via JSON-RPC over streams. Unlike local function calls, remote MCP servers can suffer from network latency, timeouts, and resource exhaustion. Without protective mechanisms, a failing server can consume the client's entire timeout budget, leading to user-facing delays or deadlocks. Resilience patterns ensure that the client remains responsive and can gracefully degrade functionality when a specific tool becomes unavailable.

Implementing Retry Logic with Exponential Backoff

The first line of defense is retry logic. However, naive immediate retries can hammer an already struggling server. Instead, we should employ exponential backoff with jitter. This strategy increases the wait time between attempts exponentially, reducing the load on the server while allowing transient issues to resolve.


import { Client } from "@modelcontextprotocol/sdk/client";

async function retryWithBackoff<T>(
  fn: () => Promise<T>,
  maxRetries: number = 3,
  baseDelayMs: number = 100
): Promise<T> {
  let lastError: Error | null = null;

  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      return await fn();
    } catch (error) {
      lastError = error as Error;

      // Check if error is retryable (e.g., network timeout)
      if (!isRetryableError(error)) {
        throw error;
      }

      const delay = baseDelayMs * Math.pow(2, attempt);
      const jitter = Math.random() * 100;
      await new Promise(resolve => setTimeout(resolve, delay + jitter));
    }
  }
  throw new Error("Max retries exceeded") from lastError;
}

It is crucial to distinguish between retryable errors (network timeouts, 503 status codes) and non-retryable errors (400 Bad Request, syntax errors). Retrying invalid requests only wastes resources.

The Circuit Breaker Pattern

Retries alone are insufficient if the server is down. If the server is unreachable, retrying will only delay the failure. This is where the Circuit Breaker pattern shines. It monitors the health of the connection and "trips" the circuit to open, preventing further requests for a specific period. This allows the server time to recover and prevents the client from being overwhelmed by pending requests.

A typical circuit breaker has three states:

  1. Closed: Normal operation. Requests flow through. Failures are counted.
  2. Open: Requests are immediately rejected. This state lasts for a "timeout" period.
  3. Half-Open: After the timeout, a limited number of test requests are allowed. If they succeed, the circuit closes; otherwise, it opens again.

class CircuitBreaker {
  private state: "closed" | "open" | "half-open" = "closed";
  private failureCount = 0;
  private lastFailureTime = 0;

  constructor(
    private failureThreshold: number = 5,
    private resetTimeoutMs: number = 10000
  ) {}

  async execute<T>(fn: () => Promise<T>): Promise<T> {
    if (this.state === "open") {
      if (Date.now() - this.lastFailureTime > this.resetTimeoutMs) {
        this.state = "half-open";
      } else {
        throw new Error("Circuit breaker is open");
      }
    }

    try {
      const result = await fn();
      this.handleSuccess();
      return result;
    } catch (error) {
      this.handleFailure();
      throw error;
    }
  }

  private handleSuccess() {
    this.failureCount = 0;
    this.state = "closed";
  }

  private handleFailure() {
    this.failureCount++;
    this.lastFailureTime = Date.now();

    if (this.failureCount >= this.failureThreshold) {
      this.state = "open";
    }
  }
}

Combining Strategies for Maximum Resilience

The most robust MCP clients combine both patterns. The circuit breaker wraps the connection, and the retry logic wraps the individual request execution within the closed or half-open states. This ensures that we attempt to recover from transient failures without overwhelming a failing infrastructure.


const breaker = new CircuitBreaker();
const safeCall = async (request: any) => {
  return breaker.execute(() => retryWithBackoff(() => client.send(request)));
};

Conclusion

As MCP adoption grows, the complexity of distributed AI systems will increase. Implementing circuit breakers and retry logic is not just a best practice; it is a necessity for building reliable, professional-grade AI agents. By proactively managing failure states and resource consumption, you ensure that your MCP clients remain stable, responsive, and resilient in the face of inevitable network and service disruptions.

Share: