Agent Frameworks

PydanticAI for Structured Data Extraction: Building Type-Safe Agents for Complex JSON Outputs

Building LLM agents that reliably return structured data is one of the most common pain points in AI engineering. While Large Language Models are excellent at creative text generation, they are prone to hallucinations and formatting errors when tasked with producing strict JSON. Traditional approaches often involve fragile regular expressions or retry loops that parse raw text, leading to unstable production systems.

Enter PydanticAI, a Python agent framework designed to make working with LLMs safe, predictable, and developer-friendly. By leveraging the power of Pydantic models, PydanticAI ensures that LLM outputs are validated against strict schemas before they are returned to your application. This post explores how to use PydanticAI to build type-safe agents capable of extracting complex, nested JSON structures from unstructured text.

Why Type Safety Matters in LLM Pipelines

In dynamic languages like Python, "duck typing" is often sufficient for prototyping. However, when an LLM output serves as input to a downstream database write, API call, or business logic engine, ambiguity is fatal. If the model returns a string where an integer is expected, or misses a required field, your application may crash or, worse, silently fail with corrupted data.

PydanticAI solves this by integrating schema validation directly into the agent's execution loop. Instead of asking the model for "some JSON," you define a Python class using Pydantic. PydanticAI automatically translates this class definition into a JSON Schema that the LLM understands. If the LLM's response doesn't match the schema, PydanticAI can automatically trigger a self-correction loop, prompting the model to fix its errors before the data is ever exposed to your application.

Defining Complex Output Structures

One of the standout features of PydanticAI is its support for complex Pydantic models. You can use nested models, enums, and custom validators to enforce business logic rules directly on the LLM's output.

Consider a scenario where we need to extract invoice details from an email. The data includes a customer name, a list of line items with calculated totals, and a date. We can define this structure with full type safety:

from pydantic import BaseModel, Field, field_validator
from datetime import date
from pydantic_ai import Agent

class LineItem(BaseModel):
    description: str = Field(..., description="Brief description of the item")
    quantity: int = Field(..., gt=0)
    price: float = Field(..., ge=0)
    
    def total(self) -> float:
        return self.quantity * self.price

class InvoiceData(BaseModel):
    customer_name: str = Field(..., description="Full legal name of the customer")
    invoice_date: date
    line_items: list[LineItem]
    total_amount: float
    
    @field_validator('total_amount')
    @classmethod
    def check_total(cls, v: float, info) -> float:
        if 'line_items' in info.data:
            calculated = sum(item.total() for item in info.data['line_items'])
            if abs(v - calculated) > 0.01:
                raise ValueError("Total amount does not match sum of line items")
        return v

agent = Agent("openai:gpt-4o", result_type=InvoiceData)

Notice the @field_validator. This isn't just for data shape; it enforces logical consistency. If the LLM hallucinates a total that doesn't match the sum of its line items, the validation fails, and PydanticAI will prompt the model to reconsider.

Executing the Agent and Handling Responses

Once the agent is defined, executing it is straightforward. The run method returns a typed object, meaning you get full IDE support and type checking in your IDE. There is no need for data['customer_name']; you can use data.customer_name.

import asyncio

email_text = """
Subject: Invoice #12345
Hi John,
Here is the invoice for October.
Item 1: Widget A x2 @ $50.00
Item 2: Widget B x1 @ $100.00
Total: $200.00
Date: 2023-10-27
"""

async def main():
    result = await agent.run(email_text)
    
    # result.data is a fully validated InvoiceData object
    print(f"Customer: {result.data.customer_name}")
    print(f"Total: {result.data.total_amount}")
    # Safe to access nested data without KeyError or AttributeError
    for item in result.data.line_items:
        print(f"- {item.description}: {item.total()}")

asyncio.run(main())

Best Practices for Complex Exports

  1. Keep Models Flat Where Possible: While nesting is supported, deeply nested structures can sometimes confuse LLMs. Flattening the schema where logically permissible often improves accuracy.
  2. Use Descriptive Field Descriptions: LLMs rely on the description parameter in Pydantic fields to understand context. Be explicit about format requirements (e.g., "ISO 8601 date").
  3. Leverage Enumerations: If a field has a fixed set of values (like status codes), use Python enum classes. This drastically reduces the chance of invalid categorical data.

Conclusion

PydanticAI represents a significant step forward in making LLM applications production-ready. By shifting the burden of data validation from post-hoc parsing to pre-eminent type checking, you eliminate a whole class of runtime errors. For developers building agents that need to interface with structured data pipelines, PydanticAI offers a clean, Pythonic, and robust solution. Embrace type safety, and your LLM agents will thank you with reliable, predictable outputs.

Share: