For developers working with Large Language Models, cost and latency are often the primary bottlenecks when processing large datasets. Whether you are generating product descriptions for an e-commerce site with millions of items, summarizing thousands of support tickets, or performing sentiment analysis on historical logs, making individual API calls for each item is inefficient and expensive.
Enter the OpenAI Batch API. This feature allows you to send large volumes of non-urgent requests to OpenAI at a significant discount. By bundling requests into a single job that is processed asynchronously, you can reduce your input and output token costs by up to 50% compared to real-time API calls. In this guide, we will explore how to implement the Batch API effectively for large-scale data processing.
Why Use the Batch API?
The primary advantage of the Batch API is cost efficiency. OpenAI offers a 50% discount on all tokens used in batch requests. However, this comes with trade-offs:
- Asynchronous Processing: You cannot expect immediate responses. Jobs typically complete within 24 hours.
- No Real-Time Interactivity: It is not suitable for chatbots or applications requiring immediate user feedback.
- Fixed Batch Size: You must prepare all requests upfront before submission.
If your use case involves offline analytics, content generation, or data labeling, the Batch API is an ideal fit.
Step 1: Prepare Your JSONL File
The first step in any batch job is creating a JSON Lines (JSONL) file. Each line in this file represents a single API request. The structure must match the standard OpenAI API request format, including a unique `custom_id` for tracking.
import json
def create_batch_file(data_points, output_filename="batch_input.jsonl"):
"""
Generates a JSONL file for the OpenAI Batch API.
Args:
data_points: List of strings to be processed.
output_filename: Name of the output JSONL file.
"""
with open(output_filename, 'w') as f:
for i, text in enumerate(data_points):
request = {
"custom_id": f"request-{i}",
"method": "POST",
"url": "/v1/chat/completions",
"body": {
"model": "gpt-4o",
"messages": [
{
"role": "user",
"content": f"Summarize the following text in one sentence: {text}"
}
],
"max_tokens": 100
}
}
f.write(json.dumps(request) + '\n')
print(f"Batch file '{output_filename}' created successfully.")
# Example usage
sample_data = [
"The quick brown fox jumps over the lazy dog.",
"Python is a high-level programming language known for its readability.",
"Machine learning is a subset of artificial intelligence."
]
create_batch_file(sample_data)
Step 2: Upload the File and Create the Batch Job
Once the JSONL file is ready, you need to upload it to OpenAI and initiate the batch job using the Python SDK.
from openai import OpenAI
import time
client = OpenAI() # API key read from environment variable
# 1. Upload the file
file_response = client.files.create(
file=open("batch_input.jsonl", "rb"),
purpose="batch"
)
file_id = file_response.id
print(f"File uploaded with ID: {file_id}")
# 2. Create the batch job
batch = client.batches.create(
input_file_id=file_id,
endpoint="/v1/chat/completions",
completion_window="24h"
)
batch_id = batch.id
print(f"Batch job created with ID: {batch_id}")
Step 3: Poll for Completion
Since the batch job is asynchronous, you must poll the status until it is completed. Depending on the size of the job, this can take several hours. For production systems, it is recommended to implement a cron job or a webhook-based notification system rather than blocking in a script.
def poll_batch_status(batch_id, client, interval=60):
"""
Polls the batch job status until it is completed or failed.
"""
while True:
batch = client.batches.retrieve(batch_id=batch_id)
status = batch.status
print(f"Current status: {status}")
if status in ["completed", "failed", "expired", "cancelled"]:
return batch
time.sleep(interval)
# Wait for completion
final_batch = poll_batch_status(batch_id, client, interval=300) # Check every 5 minutes
Step 4: Retrieve and Process Results
Once the batch is marked as `completed`, you can download the output file. This file contains the responses corresponding to each `custom_id` from the input file.
def download_and_process_results(batch_id, client):
"""
Downloads the batch output file and parses the results.
"""
# Retrieve the output file ID
output_file_id = client.batches.retrieve(batch_id).output_file_id
if output_file_id is None:
raise ValueError("No output file found. Job may have failed.")
# Retrieve the file content
file_content = client.files.content(output_file_id)
# Parse the JSONL content
results = []
for line in file_content.splitlines():
if line.strip():
data = json.loads(line)
# Extract the content from the response
response_content = data.get('response', {}).get('body', {}).get('choices', [{}])[0].get('message', {}).get('content', '')
results.append({
'id': data.get('custom_id'),
'summary': response_content
})
return results
# Process the results
results = download_and_process_results(batch_id, client)
for r in results:
print(f"{r['id']}: {r['summary']}")
Best Practices for Batch Processing
- Handle Errors Gracefully: Some requests may fail due to content policy violations or formatting issues. Always check the `error` field in the output JSONL lines.
- Chunk Large Datasets: While there is no strict limit, extremely large files can be harder to debug. Consider splitting millions of rows into chunks of 10,000-50,000 for easier management.
- Version Control Your Prompts: Since batch jobs take time to run, ensure your prompt template is stable. If you need to change the prompt, you must create a new file and a new batch job.
- Monitor Costs: Use the OpenAI Dashboard to track token usage and costs for each batch job to ensure you are maximizing savings.
Conclusion
The OpenAI Batch API is a powerful tool for developers looking to scale their AI applications without scaling their costs linearly. By leveraging asynchronous processing, you can tackle large-scale data challenges that would be prohibitively expensive or slow with real-time API calls. As LLM applications continue to evolve, mastering batch processing will be a critical skill for building efficient, cost-effective AI systems.
Start small with a test dataset, verify your results, and then scale up. The 50% discount makes the Batch API an indispensable part of any serious LLM development stack.