A precision glass-and-metal translator separating layered signals into visible and opaque output channels
All posts

How to set reasoning effort across providers: OpenAI, Anthropic and Gemini

Compare OpenAI reasoning effort, Anthropic thinking and Gemini thinking controls. Configure AIVAX requests and handle summaries, state and token costs.

To set reasoning effort across OpenAI, Anthropic and Gemini, choose the control supported by the exact model and endpoint: OpenAI uses reasoning_effort in Chat Completions or reasoning.effort in Responses; Anthropic uses thinking configuration and output_config.effort; Gemini uses thinking levels or a model-specific thinking budget. In AIVAX, send reasoning_effort to the chat-completions endpoint, but do not assume one value means the same amount of work across providers.

The setting controls only part of the integration. A reasoning response can also contain readable summaries, opaque state, signatures and usage. The final answer, the explanation you may display, and the state needed for the next tool turn are different data. Keep them separate.

Which reasoning parameter does each provider accept?

Check both the model and the API surface before copying a request. A field supported by one endpoint can be unknown on another, even when both serve the same model.

  • OpenAI Chat Completions: use reasoning_effort. OpenAI Responses: use the nested reasoning.effort field. The OpenAI reasoning guide lists model-dependent levels that can include none, minimal, low, medium, high, xhigh and max. This is not a universal enum: some models cannot turn reasoning off, and defaults differ. Requesting a summary is a separate choice from setting effort.
  • Anthropic Messages: use output_config.effort on compatible models. Effort affects text, tool calls and thinking, not just the hidden thinking budget. Current effort documentation distinguishes low, medium, high, xhigh and max, with model-specific support. For models using adaptive thinking, follow that model's thinking configuration; some enable it by default. Manual extended thinking uses thinking.type: "enabled" and budget_tokens only on models that still support that mode. Do not carry an older budget configuration into a newer model unchanged.
  • Google Gemini: distinguish the endpoints. The Interactions thinking guide uses generation_config.thinking_level. The GenerateContent guide documents generationConfig.thinkingConfig.thinkingLevel and, for budget-based models, thinkingBudget. Accepted levels and the ability to disable thinking depend on the model. A thinking level is not a portable numeric token budget.

These are native provider contracts, not extra fields to paste into an AIVAX request. A provider-compatible gateway may translate them differently. Start with a supported setting, measure task success and latency, and change one control at a time.

How do I set reasoning_effort in AIVAX?

Send this JSON body to POST https://inference.aivax.net/v1/chat/completions with Content-Type: application/json and Authorization: Bearer YOUR_AIVAX_API_KEY. Use a private key on a trusted backend. Replace <REASONING_CAPABLE_MODEL_ID> with a currently available model whose provider supports medium effort; the placeholder is not a runnable model ID.

{
    "model": "<REASONING_CAPABLE_MODEL_ID>",
    "messages": [
        {
            "role": "user",
            "content": "Compare two retry policies for a payment service. Identify duplicate-charge risks and recommend safeguards."
        }
    ],
    "reasoning_effort": "medium",
    "stream": false
}

To find a candidate, inspect GET /v1/models: use data[].id and check _details.provider_model.capabilities for Thinking, along with deprecation and provider availability metadata. That capability identifies reasoning support, not the accepted effort values or a guarantee of visible reasoning. Availability still depends on your account and the selected provider.

The request's non-null reasoning_effort overrides the gateway setting. If omitted, the gateway's configured value remains in effect. AIVAX treats effort as a provider-specific string rather than enforcing a universal list of levels. Acceptance by the gateway therefore does not prove that every selected model supports the value.

On the Chat Completions upstream path, AIVAX sends reasoning_effort. On its Responses transport, it converts that field into reasoning: { "effort": "medium", "summary": "auto" }. A top-level reasoning object in the public AIVAX request is not a substitute for reasoning_effort.

The current implementation uses OpenAI-compatible upstream transports; it does not itself implement native Anthropic budget_tokens or Gemini thinkingBudget mappings. When a managed provider serves those model families, its compatibility layer determines the translation. Check the selected model/provider combination instead of treating a common field name as identical behavior.

Use the AIVAX inference reference and AI Gateway configuration guide alongside the request. When changing an alias or model, repeat the compatibility checks described in pinning a model ID or using an alias.

Where do reasoning summaries appear in responses and streams?

In AIVAX's standard chat-completion envelope, read the final answer from choices[0].message.content and available reasoning text from choices[0].message.reasoning. With stream: true and the default rendering mode, assemble answer fragments from choices[0].delta.content and reasoning fragments from choices[0].delta.reasoning separately. Reasoning deltas are omitted when there is no text to emit.

Do not require reasoning_content or reasoning_details in that public response. AIVAX recognizes those fields on internal upstream paths, but its public chat-completion objects expose reasoning, not a lossless copy of every upstream reasoning field. The Responses translator can place both reasoning text and summaries in that channel. Its content should not be described as the model's complete raw chain of thought.

Native streams use different event types. OpenAI Responses distinguishes reasoning-summary deltas; Anthropic separates thinking and signature deltas inside content blocks; Gemini Interactions separates thought steps from answer text. A parser written for one native stream should not be reused unchanged for another.

Treat reasoning as a separate display and retention decision. A useful developer summary may contain user data or intermediate hypotheses that do not belong in an end-user answer. If the answer is JSON, parse only the completed answer channel, not reasoning fragments; see fixing invalid JSON with structured outputs.

What reasoning state must survive a tool call?

A readable summary is not a replacement for provider state. OpenAI Responses can use encrypted reasoning items for stateless continuity. Anthropic thinking blocks can carry signatures or redacted content. Gemini returns thought signatures that must remain attached to the appropriate parts when replayed.

For native integrations, follow the provider's replay rules and preserve required blocks unchanged. Do not concatenate signed parts, turn opaque state into display text, or assume it can move to another provider.

AIVAX's Responses translator preserves structured reasoning details internally so compatible upstream items can be reconstructed during its workflow. That internal support does not establish a public, lossless replay contract: the chat-completions response described above does not return all those details. If your application manages the next tool turn itself and requires opaque provider state, verify that the chosen API path exposes and accepts it before relying on stateless replay.

Test a complete sequence, not just whether a thinking panel appears:

  1. Generate reasoning followed by an answer without tools.
  2. Generate a tool call, return its result, and verify that the next answer succeeds.
  3. Repeat with multiple tool calls and an interrupted stream.
  4. Check what survives when only a summary or opaque state is returned.
  5. Repeat after a model or transport change; a successful first turn does not prove replay compatibility.

How are reasoning tokens billed?

Visible summary length is not a reliable measure of reasoning cost. OpenAI bills reasoning tokens as output tokens. Anthropic counts current-turn thinking as output; retained prior-turn thinking can also contribute to input usage according to the model's preservation rules. Google documents response pricing as output plus thinking tokens, even when only a summary is visible.

In AIVAX, reasoning contributes through the normalized output-usage accounting; this path does not apply a separate reasoning surcharge. Its public chat-completion usage reports aggregate token counts and cost rather than a promised completion_tokens_details.reasoning_tokens breakdown. Do not add an estimated reasoning count to completion_tokens yourself: upstream conventions differ, and the output total may already include it.

Measure the returned cost, time to first answer text, total latency and task outcome together. Lower effort can be useful for routine requests; higher effort may be justified by harder planning or tool use. Neither is a universal quality guarantee. A small output cap can also truncate an answer after the model has already spent tokens reasoning, so distinguish effort tuning from an output limit.

Frequently asked questions

Why does the API reject reasoning_effort or reasoning?

Check the endpoint first: OpenAI Chat Completions and Responses use different parameter shapes. Then check the exact model's supported values. For AIVAX chat completions, use reasoning_effort, not a top-level reasoning object, and confirm that the selected upstream supports the value.

Does higher reasoning effort always produce a better answer?

No. It allows a different allocation of reasoning work, not a guaranteed improvement on every task. Compare representative requests at supported effort levels and keep the lowest setting that meets your quality requirements. Include tool completion and recovery, not just the first answer, in that comparison.

Do I pay for reasoning tokens when no reasoning text is returned?

Yes, when the provider generates billable reasoning tokens. Hiding or omitting a summary does not make the underlying work free. Use usage and cost data rather than the length of visible reasoning, and keep provider-specific token breakdowns separate from totals that already include them.