Agent Frameworks

Mastra for Serverless AI: Deploying Stateful Agents on Edge Functions with Vercel

Building scalable AI agents is a balancing act. You need the robust orchestration of a framework, the low latency of edge computing, and the cost-efficiency of serverless architectures. This is where the combination of Mastra and Vercel Edge Functions shines. Mastra provides a comprehensive toolkit for building LLM-powered applications, while Vercel’s Edge runtime ensures your agents execute globally with minimal overhead.

In this guide, we’ll explore how to build a stateful AI agent using Mastra and deploy it as a Vercel Edge Function. We will address common challenges, such as managing state in serverless environments and optimizing cold starts.

Why Mastra + Vercel Edge?

Mastra is an open-source framework designed specifically for building AI agents and workflows. It offers:

  • Agents: Encapsulated LLM capabilities with tools and memories.
  • Workflows: Deterministic, multi-step processes with branching logic.
  • Evaluators: Built-in testing and evaluation metrics for AI output.

Vercel Edge Functions run at the edge, close to your users. They provide:

  • Low Latency: Execution in milliseconds across global regions.
  • Cost Efficiency: Pay-per-request billing model.
  • Language Support: Full TypeScript and JavaScript support.

Combining these technologies allows you to build sophisticated AI agents that respond in under 100ms, regardless of where your users are located.

Setting Up Your Project

First, initialize a new Vercel project with the Edge runtime. Then, install Mastra:


npm create vercel@latest
cd your-project
npm install mastra

Create a file named edge.ts in your project root. This file will define your Edge Function entry point.

Building a Stateful Agent with Mastra

One of the challenges with Edge Functions is their ephemeral nature. State must be managed externally or within the request context. Mastra simplifies this with its Agent class, which can handle memory and context management.


import { Agent } from 'mastra';
import { generateText } from 'ai';
import { openai } from '@ai-sdk/openai';

// Define your agent with memory capabilities
const myAgent = new Agent({
  name: 'SupportAgent',
  description: 'An agent that handles customer support queries.',
  model: openai('gpt-4o'),
  tools: {
    // Example tool: lookup_order
    lookup_order: {
      description: 'Looks up an order by ID',
      execute: async ({ orderId }: { orderId: string }) => {
        // In a real app, this would query a database
        return { status: 'shipped', trackingId: '1Z999AA10123456784' };
      },
    },
  },
});

export async function GET(request: Request) {
  const { message, userId } = await request.json();

  try {
    // Mastra handles the agent loop, tool execution, and response formatting
    const result = await myAgent.generate({
      messages: [
        { role: 'user', content: message },
      ],
      // Pass userId for state management if using Mastra's memory features
      metadata: { userId },
    });

    return new Response(JSON.stringify({
      text: result.text,
      usage: result.usage,
    }), {
      headers: { 'Content-Type': 'application/json' },
    });
  } catch (error) {
    return new Response(JSON.stringify({ error: error.message }), {
      status: 500,
      headers: { 'Content-Type': 'application/json' },
    });
  }
}

Managing State in Serverless Environments

Edge Functions are stateless by design. If your agent requires memory (e.g., conversation history), you have two options:

  1. Client-Side State: Send the entire conversation history with each request. This is simple but can lead to large payloads.
  2. External Storage: Use a serverless-friendly database like Vercel KV, Upstash Redis, or a Postgres instance (e.g., Neon) to store conversation state. Mastra can integrate with these via custom tools or middleware.

For high-scale applications, using an external cache with a TTL (Time-To-Live) is recommended to keep costs low and latency high.

Optimizing Cold Starts

Cold starts can add latency to your AI agent. Here are some best practices:

  • Bundle Size: Keep your Edge Function bundle small. Only import the specific Mastra components you need.
  • Pre-warming: Use Vercel’s keepAlive feature or scheduled invocations to keep your function warm during peak hours.
  • Model Caching: If you use local models or smaller LLMs, consider caching model weights in memory (if supported by your runtime) or using Vercel’s ai SDK providers that handle connection pooling.

Deploying to Vercel

Deploying is straightforward:


vercel deploy

Ensure your vercel.json configures the Edge runtime for your function:


{
  "functions": {
    "edge.ts": {
      "runtime": "edge"
    }
  }
}

Conclusion

Deploying stateful AI agents on Vercel Edge Functions with Mastra provides a powerful combination of performance, scalability, and developer experience. By leveraging Mastra’s agent orchestration and Vercel’s global edge network, you can build responsive, cost-effective AI applications that scale effortlessly. Start small with a single agent, monitor your latency and costs, and iterate to build robust AI-powered features for your users.

Share: