Try Bifrost Enterprise free for 14 days. Request access

Top 4 LLM Gateways with OpenAI Batch API and Anthropic Batch Support in 2026

Compare the top LLM gateways for the OpenAI Batch API and Anthropic Message Batches in 2026 on batch formats, provider coverage, governance, and deployment.

Top 4 LLM Gateways with OpenAI Batch API and Anthropic Batch Support in 2026

TL;DR

  • The OpenAI Batch API and Anthropic Message Batches both charge 50% of standard token prices for jobs that complete within a 24-hour window.
  • Bifrost accepts batch calls from the OpenAI, Anthropic, and AWS Bedrock SDKs and routes them to OpenAI, Anthropic, Bedrock, Gemini, and Azure batch APIs from one self-hosted gateway.
  • LiteLLM exposes an OpenAI-format /v1/batches endpoint for seven providers and forwards Anthropic Message Batches through a native pass-through route.
  • OpenRouter runs a hosted Batch API that takes inline JSON requests, pins each batch to one model and one provider, and does not accept JSONL file uploads.
  • Kong AI Gateway supports llm/v1/batches and llm/v1/files route types from Kong Gateway 3.11; Cloudflare AI Gateway was left out because its docs do not cover provider batch APIs.

The OpenAI Batch API processes large groups of requests asynchronously at a 50% discount with a 24-hour completion window, and Anthropic's Message Batches API applies the same discount to batches of up to 100,000 requests. Routing those jobs through an LLM gateway keeps batch traffic under the same keys, budgets, and logs as real-time traffic, but gateway support for provider batch APIs varies widely. Bifrost, the open-source AI gateway written in Go and built by Maxim AI, is the best choice for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability, including batch inference across OpenAI, Anthropic, and Bedrock. This guide compares four gateways with verified batch API support and explains how to pick one.

What Is a Batch API?

A batch API is an asynchronous inference endpoint that accepts a large group of model requests as one job, processes them within a fixed completion window, and returns every result together at a reduced per-token price. Batch APIs suit evaluation runs, classification, data extraction, and embedding backfills, where no user is waiting on the response.

The table summarizes the provider batch APIs these gateways route to, using figures from the OpenAI Batch API guide, the Anthropic batch processing docs, and the Amazon Bedrock batch inference guide.

Attribute OpenAI Batch API Anthropic Message Batches API Amazon Bedrock batch inference
Input format JSONL file uploaded through the Files API Inline requests array JSONL files in an Amazon S3 bucket
Price 50% of synchronous pricing 50% of standard API prices Set by Bedrock pricing per model
Per-batch limit 50,000 requests or 200 MB 100,000 requests or 256 MB Set by AWS service quotas
Completion window 24 hours 24 hours; most batches finish in under 1 hour Job-based
Results Output file and error file Results available for 29 days Output files written to S3

Bifrost tracks which providers expose batch endpoints in its provider capability matrix, which lists Batch API support for OpenAI, Anthropic, Azure, Bedrock, and Gemini.

How the OpenAI Batch API Works

The OpenAI Batch API works in four steps: upload a JSONL file of requests through the Files API, create a batch that points at the file, poll until the batch reaches a terminal status, and download the output and error files. Each JSONL line carries a custom_id, an HTTP method, a target endpoint, and a request body.

OpenAI Batch API lifecycle: a JSONL file is uploaded, a batch is created with a 24-hour window, then results land in output and error files

Figure 1: The batch lifecycle is asynchronous end to end, so the client polls or is notified rather than holding a connection open.

As Figure 1 shows, results map back to inputs only through custom_id.

Supported targets are /v1/responses, /v1/chat/completions, /v1/embeddings, /v1/completions, /v1/moderations, and image endpoints. OpenAI batch API pricing is 50% of the synchronous rate, and batch rate limits come from a separate pool, so a large batch does not consume the per-model limits that real-time traffic depends on. An organization can create up to 2,000 batches per hour, and requests that miss the window land in the error file with a batch_expired code. Related tactics appear in this guide to managing OpenAI rate limits at scale.

How the Anthropic Batch API differs

The Anthropic batch API, also called the Claude batch API, takes requests inline: the create call carries a requests array of custom_id and params objects, and no file upload is involved. A batch holds up to 100,000 requests or 256 MB, most batches complete within an hour, and results stay downloadable for 29 days. Prompt caching discounts stack with the batch discount, although cache hits are best effort because batch requests run concurrently.

Where Bedrock batch inference fits

Bedrock batch inference reads JSONL input in InvokeModel or Converse format from an S3 bucket and writes responses back to S3. A gateway therefore has to handle S3 storage configuration as well as the job API, which makes Bedrock the least supported of the three among the gateways here. Bifrost handles both through its Bedrock SDK files and batch integration.

Why Run Batch Inference Through an LLM Gateway

Batch inference through an LLM gateway gives every batch producer one credential, one budget model, and one log, regardless of which provider batch API runs the job. Without one, each pipeline holds raw provider keys and moving a job from OpenAI to Bedrock means rewriting the batch client.

Three batch producers submit jobs to an LLM gateway that applies keys, budgets, and logging before forwarding to OpenAI, Anthropic, and Bedrock batch APIs

Figure 2: The gateway gives every batch producer one credential, one budget model, and one log, whichever provider batch API runs the job.

A primer on what an LLM gateway does for enterprise AI covers the general model; for batch jobs, four functions matter most:

  • Credential isolation: pipelines authenticate with a gateway key, and provider keys stay inside the gateway.
  • Spend control: batch jobs count against the same budgets as interactive traffic, at batch token rates.
  • Format translation: one client format reaches providers whose batch APIs use different payload shapes.
  • Audit trail: batch calls appear in the same logs as every other call.

For a comparison beyond batch support, see this production-ready comparison of the top LLM gateways, and the Bifrost governance overview for policy design.

Key Criteria for Evaluating LLM Gateways for Batch Workloads

The criteria that separate LLM gateways for batch workloads are which batch API surfaces they expose, which providers they reach, whether they translate formats across providers, how they govern and price batch calls, and where they run. Batch support listed on a feature page often means one provider in one format.

Criterion Why it matters What to check
Batch API surfaces Existing code uses the OpenAI SDK, the Anthropic SDK, or boto3 Support for /v1/batches, /v1/messages/batches, and Bedrock model invocation jobs
Provider coverage Batch discounts differ by provider and model Which providers are documented for batch, not only for chat
Cross-provider routing Moving a job between providers should not require a new client Whether one SDK can target several provider batch APIs
Governance Batch jobs can spend a month of budget in one submission Keys, model allowlists, budgets, and rate limits on batch calls
Cost accounting Batch tokens are billed at different rates from real-time tokens Whether spend tracking uses batch pricing
Deployment model Batch inputs often contain production data Self-hosted, in-VPC, or hosted only

Budget design matters more for batch traffic, because one call can enqueue tens of thousands of requests. Bifrost applies budgets and rate limits at the virtual key, team, and customer levels, and its cost calculation accounts for batch pricing; the hierarchy is explained in this guide to LLM budget management with virtual keys.

Cloudflare AI Gateway was evaluated and excluded because its documentation does not cover OpenAI or Anthropic batch endpoints.

LLM Gateways with Batch API Support Compared at a Glance

LLM gateways with batch API support differ most in format and provider reach. Bifrost and LiteLLM speak the OpenAI file-based batch format and reach several providers, OpenRouter uses its own inline batch format on a hosted service, and Kong adds batch route types to Kong Gateway.

Capability Bifrost LiteLLM OpenRouter Kong AI Gateway
OpenAI-format batches and files Yes, through the OpenAI SDK Yes, /v1/batches and /v1/files No; inline JSON requests, no JSONL upload Yes, llm/v1/batches and llm/v1/files (3.11+)
Anthropic Message Batches Yes, Anthropic SDK beta.messages.batches Native pass-through route /v1/messages request shape inside an OpenRouter batch Native format via llm_format: anthropic
Bedrock batch inference Yes, boto3 jobs with S3 files Yes Not published Not published
Providers documented for batch OpenAI, Anthropic, Azure, Bedrock, Gemini OpenAI, Azure, Vertex AI, Bedrock, Mistral, vLLM, xAI Models with :batch endpoints OpenAI, Azure OpenAI, Anthropic, Gemini, Vertex AI
Cross-provider batch from one SDK Yes, per-request provider selection Model-based routing with managed files (beta) One provider per batch, cheapest by default Not published
Batch governance Virtual keys, budgets, rate limits Batch rate limits from input-file tokens, model allowlists Workspace-scoped batches, BYOK Plugin-based; analytics, logging, cost calculation
Deployment Self-hosted, in-VPC, on-prem Self-hosted proxy Hosted only On-prem or Konnect

The LLM gateway buyer's guide covers criteria beyond batch support.

1. Bifrost

The Bifrost gateway supports the OpenAI, Anthropic, and AWS Bedrock batch interfaces natively and routes batch jobs across providers from any of them. Teams pick the provider per request, and virtual keys govern every files and batch call.

OpenAI, Anthropic, and Bedrock SDK clients send batch calls to Bifrost, which applies virtual keys and routes them to five provider batch APIs

Figure 3: Teams keep the SDK they already use and pick the provider per request, while virtual keys govern every batch call in one place.

Bifrost connects to 25+ providers and 10,000+ models through one OpenAI-compatible API and adds 11 microseconds of overhead per request at 5,000 RPS in published benchmarks.

Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.

OpenAI SDK: files and batches across providers

The OpenAI SDK files and batch integration supports file upload, list, retrieve, delete, and content download, plus batch create, list, retrieve, and cancel. The target provider is set with extra_body on POST requests or extra_query on GET requests:

batch = client.batches.create(
    input_file_id="file-abc123",
    endpoint="/v1/chat/completions",
    completion_window="24h",
    extra_body={"provider": "openai"},  # or "bedrock", "anthropic", "gemini"
)

Bifrost converts OpenAI-style JSONL to Bedrock and Gemini batch formats internally. Bedrock jobs add an S3 storage_config and output_s3_uri; Anthropic jobs pass inline requests.

Anthropic SDK: Message Batches with cross-provider routing

The Anthropic SDK batch integration exposes beta.messages.batches for create, list, retrieve, cancel, and results. Setting the x-model-provider header to openai or gemini sends the same inline batch to those providers. The Anthropic SDK sends a Bifrost virtual key in x-api-key, so batch calls are governed like any other request. Bedrock batches, which need S3 file input, run through the Bedrock SDK path instead.

Bedrock SDK: S3 files and model invocation jobs

Through boto3, Bifrost serves an S3-compatible files endpoint and the Bedrock batch methods: create_model_invocation_job, list_model_invocation_jobs, get_model_invocation_job, and stop_model_invocation_job. The x-model-provider header also routes these jobs to OpenAI or Gemini.

Governance and cost accounting for batch jobs

Virtual keys are deny-by-default for providers and carry model allowlists, budgets, and rate limits. The Model Catalog stores separate batch input and output token prices, so spend recorded for a batch job reflects the discounted rate rather than real-time pricing.

Async inference and webhooks for smaller deferred jobs

Bifrost async inference accepts a normal request on a /v1/async/ path, returns a job_id with 202 Accepted, and stores the result for a default TTL of 3,600 seconds. Jobs submitted with a virtual key can only be polled with that key, and governance and cost tracking still run. Signed webhooks in the Standard Webhooks format notify a receiver when an async job completes or fails, with at-least-once delivery and retries. Webhooks apply to async jobs, not provider batch jobs.

Enterprise deployment

Batch inputs often contain production records, so deployment location matters. Bifrost runs inside a private cloud through in-VPC deployments and as a high-availability cluster with gossip-based state sync through clustering.

Signed audit logs record administrative activity: who changed what, when, and which resource was affected. The Bifrost Enterprise tier adds these controls on top of the open-source gateway.

2. LiteLLM

LiteLLM is a Python SDK and self-hosted proxy with a /batches endpoint that covers both batches and files in OpenAI format. Its documentation lists provider-native batch support for OpenAI, Azure OpenAI, Google Vertex AI, Amazon Bedrock, Mistral, vLLM, and xAI, with Anthropic Message Batches available through a native pass-through route.

LiteLLM's batch features as documented:

  • OpenAI-format batches: clients upload a JSONL file to /v1/files and create the job on /v1/batches.
  • Anthropic pass-through: /anthropic/v1/messages/batches forwards Message Batches in native format with no translation, with cost tracking on that route.
  • Batch rate limiting: limits are enforced at POST /v1/batches from the input file, and model allowlists are checked per line.
  • Guardrails on batch input: guardrails configured on the proxy apply to the records inside an uploaded batch file.
  • Cost tracking: batch cost tracking on the /batches endpoint is listed as an Enterprise feature.

Best for: Python-centric teams that want a self-hosted proxy with broad OpenAI-format batch coverage, including Vertex AI and Mistral, and accept that Anthropic batches use a separate pass-through route.

Teams comparing self-hosted options can review these open-source LLM gateways for self-hosted deployments, and teams moving off LiteLLM can see this roundup of LiteLLM alternatives.

3. OpenRouter

The OpenRouter batch API is a hosted endpoint at /api/v1/batches that accepts an inline JSON requests array instead of an uploaded JSONL file. Each batch targets one endpoint shape (chat completions, Responses, Anthropic Messages, or embeddings), one model, and one provider, and OpenRouter handles JSONL persistence internally.

Key behaviors from the OpenRouter Batch API quickstart:

  • Provider selection: OpenRouter picks the cheapest eligible :batch endpoint for the model by default, and provider.only pins specific providers.
  • No fallbacks: allow_fallbacks and other sync preferences are rejected.
  • Pricing: batch requests are typically billed at 50% of standard per-token pricing, and BYOK batches bill inference to the provider key.
  • Results: completed batches return results inline in the GET /batches/{id} response, with no separate download endpoint.

Best for: teams already on OpenRouter's hosted service that want discounted batch pricing without managing provider accounts, and do not need the OpenAI file-based batch format or self-hosting.

Teams evaluating self-hosted replacements can compare options in this guide to the best OpenRouter alternative in 2026.

4. Kong AI Gateway

Kong AI Gateway adds batch support through the AI Proxy and AI Proxy Advanced plugins, which support llm/v1/batches and llm/v1/files route types in OpenAI format from Kong Gateway 3.11. Kong also passes Anthropic Message Batches and Gemini and Vertex AI batches through in native provider formats.

Kong's batch capabilities as documented:

  • OpenAI-format routes: llm/v1/batches and llm/v1/files support create, retrieve, and delete operations, with documented examples for OpenAI and Azure OpenAI.
  • Anthropic batches: with llm_format set to anthropic, requests reach /v1/messages/batches; Kong notes Anthropic batch processing is supported in native SDK format only.
  • Observability in native mode: analytics, logging, and cost calculation still apply when payloads are not translated.
  • Deployment: the plugins run on self-managed Kong Gateway or in Konnect, and the AI Proxy Advanced batch example is marked for the AI Gateway Enterprise tier.

Best for: organizations already running Kong Gateway for API management that want to add batch routes for OpenAI and Azure OpenAI as plugin configuration.

Teams weighing a move away from a plugin-based model can review these Kong AI Gateway alternatives.

Batch API, Async Inference, or Real-Time Calls: Which to Use

Use a provider batch API when a workload has thousands of requests and can wait up to 24 hours, async inference when a single request can be deferred but should finish in minutes, and real-time calls when a user is waiting. The batch discount only pays off at volume.

Decision flow sending interactive work to real-time calls, single deferrable jobs to async inference, and large deferrable jobs to a provider batch API

Figure 4: The batch discount only pays off when the workload is large and can tolerate a completion window of up to 24 hours.

Figure 4 reduces the choice to two questions. Workloads between batch and real time, such as a long report a user checks later, fit Bifrost async inference with webhook delivery better than a batch job. Batch pricing is one of several cost levers; this guide on how to cut LLM API and token costs covers routing and caching, and this comparison of AI gateways with semantic caching covers repeated real-time queries.

Frequently Asked Questions

What is batch API in OpenAI?

The OpenAI Batch API is an asynchronous endpoint for sending a large group of requests in one job. Requests go into an uploaded JSONL file, the batch runs within a 24-hour window, and results return in an output file. Batch usage costs 50% less than synchronous calls and uses a separate rate limit pool.

How much does the OpenAI Batch API cost?

OpenAI charges 50% of the synchronous price for the same model and endpoint. If a batch expires, unfinished requests are cancelled and tokens from completed requests are still billed. Gateways that track batch spend should record the batch rate; Bifrost stores separate batch input and output token prices in its Model Catalog pricing data for this reason.

What is the difference between a bulk API and a batch API?

A bulk API usually performs many write operations in one synchronous call, such as inserting thousands of records at once. An LLM batch API is asynchronous: it accepts many inference requests, queues them, and returns results later within a completion window, typically at a discount.

Does the Anthropic batch API support prompt caching?

Yes. The Anthropic Message Batches API supports prompt caching, and the caching discount stacks with the 50% batch discount. Because batch requests are processed asynchronously and concurrently, cache hits are best effort, and Anthropic reports typical hit rates between 30% and 98% depending on traffic patterns.

Does OpenRouter have a batch API?

Yes. OpenRouter offers a hosted Batch API at /api/v1/batches that accepts inline JSON requests for chat completions, Responses, Anthropic Messages, or embeddings. Each batch runs one model on one provider, the only completion window is 24 hours, and pricing is typically 50% of standard rates. It does not accept OpenAI-style JSONL file uploads.

Does Cloudflare AI Gateway support the OpenAI Batch API?

Cloudflare AI Gateway documentation does not describe support for OpenAI or Anthropic batch endpoints. Its documented endpoints cover chat completions, Responses, and Anthropic Messages. Cloudflare offers a separate Asynchronous Batch API for Workers AI models.

Can an LLM gateway route a batch job to a different provider?

Yes, if the gateway translates batch formats. Bifrost accepts an OpenAI-format batch and sends it to OpenAI, Anthropic, Bedrock, or Gemini based on a per-request provider setting, converting JSONL to Bedrock and Gemini formats internally and sending Anthropic jobs as inline requests. Virtual keys still apply provider and model allowlists, so a team can permit batch traffic to approved providers only.

Try Bifrost for OpenAI Batch API Workloads

The OpenAI Batch API and Anthropic Message Batches cut deferred inference costs in half, and an LLM gateway keeps that traffic under the same keys, budgets, and logs as real-time calls. Bifrost supports batch jobs from the OpenAI, Anthropic, and Bedrock SDKs, routes them across five provider batch APIs, and adds async inference with signed webhooks for smaller deferred jobs. Explore the Bifrost resources hub for deployment guides, or book a demo to see how the Bifrost AI gateway handles batch inference at enterprise scale.