Model Gateway
Your application shouldn't be tied to one way of running a model.
One OpenAI-compatible interface to integrated models and their providers. Keep the request familiar. Choose how the work gets served.
POST /v1/chat/completions“Summarize this incident report.”
Eligible providers.
Configured levels.
Keep the model. Change the preference.
Price, throughput or quality. Make the trade-off explicit.
Provider routing chooses among currently available providers compatible with your request. It does not replace the model you asked for.
Balance the competing priorities.
Balances price, speed and quality across eligible providers.
"routing_preset": "Balanced"Choose the lowest applicable token price.
Selects the provider with the lowest applicable input and output token price—not a guaranteed lowest total application cost.
"routing_preset": "Cheapest"Favor the highest available throughput.
Prioritizes provider throughput. This is not a fixed latency promise or a guarantee of the shortest end-to-end request.
"routing_preset": "Fastest"Use the provider AIVAX ranks highest for quality.
Chooses by AIVAX's provider quality ranking, without optimizing for price or speed. It does not select a different model.
"routing_preset": "Quality"Explanations of routing behavior, not a live route simulation. No provider prices, speed measurements or rankings are fabricated here.
A different question: which model?
Not every request needs the same model.
Configure a model for each complexity level in an AI Gateway. The complexity router classifies the latest user request as low, medium or high, then selects the model configured for that level.
Provider routing chooses where a model runs. Complexity routing can change which model runs. Keep those decisions distinct.
Configure model routingConceptual mapping, not a performance tier or benchmark. When available, X-Model-Routed-Complexity reports the selected complexity level on the HTTP response.
A familiar call, with a deliberate choice
Change one request. Or save the behavior for every client.
Use an integrated model tag for a direct call. Use an AI Gateway ID or private-key slug when the model configuration, instructions, knowledge and tools should be shared.
Save the preference in your AI Gateway as parameters.routingOption. A request-level routing_preset overrides it for that call only. It requires a private API key and is an AIVAX extension to the OpenAI-compatible request body.
From an incident to a summary.
Your backend submits the report once. AIVAX selects an eligible provider for the requested integrated model using the preference you supply.
See the request body
Send to POST /v1/chat/completions with Bearer authentication. Replace the model placeholder with an available integrated model tag or a gateway using one.
{
"model": "YOUR_INTEGRATED_MODEL_OR_GATEWAY_ID",
"messages": [{
"role": "user",
"content": "Summarize this incident report: checkout requests timed out after the database pool reached capacity."
}],
"routing_preset": "Cheapest"
}This example shows request structure. No inference is executed on this page.
From the model catalog
Recent releases. One place to compare them.
Loading current catalog... Groups identify model families, not the number of serving providers.
Compare all modelsLatest dated releases among available models, ordered by the catalog's release date—not the date AIVAX added them. Loaded from the public catalog when you visit; this is not a latency guarantee.