OpenAI vs Claude API: Contracts, Cost Accounting, and Evaluation

Compare OpenAI Responses and Claude Messages using official API references, with a repeatable evaluation method and cost accounting. No unsupported performance ranking.

This article compares API contracts and selection methods. It does not rank models by quality, speed, or price. This page contains no matched provider responses, invoices, or latency logs from a shared dataset, so it cannot establish a cheaper, faster, or better front-end/back-end provider.

Originally published in May 2026; API documentation was checked on September 29, 2026 using the official sources below. Requests are templates with model IDs to replace; no paid API calls were made for this revision. The English and Chinese versions cover the same core method and conclusions.

1) Define the comparison

Record exact model IDs, date, endpoint, SDK version, account limits, and reasoning/output budgets. Context windows, modalities, rate limits, and prices vary with the selected model and account. Product experiences in Codex or Claude Code are not measurements of the underlying APIs.

2) Minimal requests: Responses and Messages

The OpenAI example uses the Responses API. Do not transplant Chat Completions messages, response_format, or legacy max_tokens fields into this request. Replace each placeholder with an accessible model that supports the selected parameters.

OpenAI Responses API

POST https://api.openai.com/v1/responses
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json
{
  "model": "REPLACE_WITH_AVAILABLE_OPENAI_MODEL_ID",
  "instructions": "You are a helpful assistant.",
  "input": "Hello!",
  "max_output_tokens": 1024,
  "store": false
}

Anthropic Messages API

POST https://api.anthropic.com/v1/messages
x-api-key: YOUR_API_KEY
anthropic-version: 2023-06-01
Content-Type: application/json
{
  "model": "REPLACE_WITH_AVAILABLE_CLAUDE_MODEL_ID",
  "system": "You are a helpful assistant.",
  "messages": [
    {
      "role": "user",
      "content": "Hello!"
    }
  ],
  "max_tokens": 1024
}

Messages uses a top-level system field and requires max_tokens. The two output limits have different names; check how the selected model accounts for reasoning tokens. Equal numbers do not guarantee equal reasoning budgets. Keep API keys on the server, out of public web pages; an API key is not a JWT to decode.

3) Parse responses and validate structured output

Do not assume the first array item is text. Raw Responses JSON contains output items; collect output_text blocks from message content. Messages content also contains typed blocks: handle text, tool_use, and other supported types separately. Streaming needs its own event parser and completion handling.

Both providers offer structured output, subject to model support. In Responses, a user-facing JSON schema goes in text.format with type: json_schema. Claude JSON output uses output_config.format, while strict tool inputs use strict: true. XML tags or a plain “return JSON” instruction are not substitutes for schema constraints.

Handle refusals, truncation, timeouts, and tool calls, then validate the business meaning of each field. Schema validity does not establish factual correctness. This page provides no measured evidence that one provider is inherently more reliable.

4) Run a comparison on your own workload

  1. Select de-identified business tasks and save inputs, acceptance criteria, and failure definitions. Keep a holdout set that was not used to tune prompts. An initial 30–50 tasks can reveal issues; that count alone does not guarantee statistical power.
  2. Use the same tasks and schemas, documenting necessary API differences. Fix model versions, concurrency, and budgets; repeat tasks and alternate provider order. Separate cold-cache and warm-cache runs.
  3. Record time to first token, total latency, task success, schema validity, manual corrections, request IDs, and usage. Count timeouts, 429s, refusals, and tool failures separately, including failed requests in cost.
  4. Publish sample size, repetitions, median/P95 latency, and uncertainty around success rates. Without these records, observations cannot support a general ranking.

First require acceptable quality and latency, then compare total cost per accepted task. This article supplies the method; it does not publish a completed matched API experiment.

5) Cost: reconcile against billed usage

Use the model price in effect on the call date and response usage: uncached input + cache reads/writes where applicable + output + additional tool/service fees. Apply only the batch or service-tier rates that actually apply. Do not assume discounts stack. Different tokenizers may bill different token counts for the same text.

Cost per accepted task = actual cost of all attempts / accepted tasks
Report cache hit rate, retries, and manual correction time separately

Fixed price tables and plan rankings without retained source evidence have been removed. Reconcile estimates with API invoices. Check consumer subscription and API-credit entitlements separately instead of combining their budgets.

6) What the linked project records show

The links below retain the AgentHub screenshots and delivery records. They show outputs for particular prompts, environments, and manual checks. They do not establish a “Codex for back-end, Claude Code for front-end” rule or replace matched API cost/latency measurements.

The reusable conclusion is to define provider adapters and acceptance criteria, then evaluate the same workload. Choose from those records rather than a brand-level ranking.

Official sources and verification scope

Checked: 2026-09-29. These sources support API fields and structured-output behavior; they are not performance or pricing measurements.

Frequently asked questions

Which API did this article measure as better?

Neither. This is a documentation-based contract comparison and evaluation method, without matched raw responses, latency logs, or invoices sufficient for a ranking.

Does Claude require Tool Use or XML for JSON output?

No. Its structured-output documentation also provides output_config.format; verify model support. XML prompts alone are not schema constraints.

How should costs be compared?

Use the same acceptance criteria, include failed attempts and retries in billed cost, then divide by accepted tasks. Record date, models, caching, and concurrency.