AI infrastructure. Memory. Reason. Act.
AIVAX brings models, retrieval and tools into one OpenAI-compatible operating layer, so behavior can evolve without losing the trail of what happened.
Knowledge matters when its path into an answer remains visible.
Approximate cumulative activity since AIVAX launched.
- Token volume
- —
- Token volume (USD)
- —
- RAG searches
- —
- RAG indexes
- —
Loading AIVAX platform activity.
Latest from AIVAX
View all postsThe platform at a glance
Everything you need to put AI to work.
Connect models and tools. Give them knowledge. Understand and manage what happens next.
Act.
From a request to something done.
-
AI gateways
Models, instructions, skills and tools behind one endpoint.
-
Inference & model routing
An OpenAI-compatible API, model choice and your own provider keys.
-
Tools & MCP connections
Connect APIs, MCP servers and model-callable functions. MCP servers can also bring their published instructions and skills into your gateway.
-
Workers & hooks
Adjust context and tool results with your own request-time logic.
-
Web search & research
Find sources, fetch pages and bring fresh information into answers.
-
Media descriptions
Analyze images, audio, videos and files through dedicated media inference.
-
Realtime voice
Build live, two-way voice conversations with your assistant.
-
WhatsApp & Telegram
Connect AI gateways to conversations on your messaging channels.
-
Shell sandbox
Give agents an isolated environment to execute commands and work with files.
-
Image generation
Create or edit images from prompts and reference images.
-
Audio generation
Turn text into spoken audio with text-to-speech models.
-
Audio transcription
Convert recordings and spoken messages into text.
-
Internet research MCP
Bring AIVAX web search and page fetching to any compatible MCP client.
-
Inference MCP
Expose an AIVAX model or AI gateway as a tool for other agents.
-
Media Generation MCP
Discover models and prices, then generate images or MP3 speech from an MCP client.
Memory.
The right knowledge, in the right context.
-
Knowledge collections
Organize documents into searchable knowledge bases.
-
Import & automatic updates
Ingest documents in bulk and reindex content when it changes.
-
Media to knowledge
Turn videos, PDFs and other supported media into searchable knowledge.
-
Semantic search
Find relevant passages by meaning, not just matching words.
-
Reranking & Reflex
Prioritize relevant results, or rank documents without a collection.
-
Grounded answers & citations
Answer from your content and trace results back to their sources.
-
Agentic RAG MCP
Let MCP-compatible agents search your collections on demand.
-
Persistent memories
Keep user context available across conversations.
-
Reusable skills
Write instructions once, reuse them across gateways and load them on demand.
-
Skill learning
Turn a demonstration video into reusable instructions with Teach Skill.
-
RAG in AI gateways
Attach collections to a gateway and retrieve knowledge within the inference request.
-
RAG rules & strategies
Set result limits, score thresholds, rerankers, query rewriting and agent-driven search.
-
AIVAX-managed chat sessions
Keep conversation history on AIVAX for embedded chats and messaging integrations.
Reason.
See clearly. Decide. Stay in control.
-
Decisions & classification
Route requests, assign labels and score content against your criteria.
-
Structured JSON & tool calls
Call tools before returning schema-validated JSON, with repair attempts when needed.
-
Agentic tests
Simulate conversations and evaluate how your gateway behaves.
-
Logs & observability
Inspect messages, tool calls, errors and usage when logging is enabled.
-
Batch workflows
Process records asynchronously and track each result.
-
Usage & costs
Follow consumption, check model pricing and manage your balance.
-
Moderation & guardrails
Screen gateway input and apply shared policies before inference.
-
Account & resource control
Manage keys, limits and platform resources from the console or API.
-
OCR to text or JSON
Extract readable text or schema-shaped data from images and documents.
-
Multimodal pre-processing
Convert media into text for models that cannot natively read images, audio, videos or files.
-
Context truncation
Set context limits, trim older messages and control how much tool history stays in view.
-
Response rendering
Render reasoning, tool activity and answers together as blocks in your chat timeline.
-
Provider routing preferences
Prioritize price, speed or quality among eligible providers for the same model.
-
Text segmentation
Split long text into useful sections for indexing and downstream processing.
AIVAX Toys
Try a capability before you write any code.
Free, single-purpose demos that run in your browser and show the API call behind every step. The first one reads a PDF, scanned or not, and hands back Markdown and chunks ready for a RAG collection.
No account needed. Up to 10 MB per PDF and 20 requests an hour per IP (5 a minute); reading and splitting count separately.
-
01 / Upload
handbook.pdf Text or scanned pages, up to 10 MB.
-
02 / Read
# Leave policy ## Annual leave Employees accrue 1.5 days of paid leave per month. ## Sick leave A medical note is required after three days.POST /api/v1/web/fetchFetch & OCR return Markdown. -
03 / Split
docid: handbook.pdf#1# Leave policydocid: handbook.pdf#2## Annual leave — Employees accrue 1.5 days...docid: handbook.pdf#3## Sick leave — A medical note is required...POST /api/v1/generations/segmentChunks download as JSONL.
Everything the client doesn't have to carry.
Compose knowledge, model behavior and tools behind one OpenAI-compatible call. Each capability stays reusable, inspectable and shared.
RAG
02 / Reason Shape how context becomes a reliable answer.Gateway capabilities
03 / Act Connect models and tools to systems outside the prompt.Tools + providers
A shared path matters when every application can rely on it.
Field guide
Five ways into the system.
Start with the question in front of you: knowledge, execution, evaluation, model choice or cost.