OpenAI Evals shutdown: where do multi-turn agent tests go before November 30?
Manually recreate OpenAI Evals datasets and graders in Promptfoo, then choose between its multi-turn tests and AIVAX gateway conversations.
Read postAIVAX / Blog
Field notes on AI infrastructure, retrieval, model operations and the systems that keep every request visible.
27 published field notes
Manually recreate OpenAI Evals datasets and graders in Promptfoo, then choose between its multi-turn tests and AIVAX gateway conversations.
Read post
Measure what MCP tool definitions cost, then cut it with allowlists, deferred tool search, or a shell. What each option needs and what AIVAX supports.
Read post
Convert a PDF to Markdown with OCR, split it into chunks for retrieval, and export JSONL for a RAG collection. Code and a free demo included.
Read post
OpenAI shuts down GPT-4, o1, o3-mini and o4-mini on October 23, 2026. Find hidden model IDs and check replacements for parameters, output and cost.
Read post
Configure AIVAX's hosted Media Generation MCP, discover models and prices, and generate images or MP3 speech with explicit tool and privacy limits.
Read post
Deduplicate the same tool operation; changed arguments create a new one. Require an explicit conflict policy to allow only one write per request.
Read post
Before embedding, sample real pages, write the facts you expect, and grade each generated document for wrong numbers, lost table headers, and missing facts.
Read post
SEP-2640 lets MCP servers offer discoverable workflow skills as resources. The host still decides what to load, verify, approve, and execute.
Read post
The expanded Free plan includes separate daily allowances for RAG embeddings, Reflex reranking, Julia-1 semantic decisions, and Fetch/OCR extraction. Here is what each covers—and what remains metered.
Read post
Check truncation, refusals and schema support first. In AIVAX, distinguish native response_format from response_schema validation and bounded healing.
Read post
A lower token rate does not tell you what a completed task costs. AIVAX's Complexity Router selects a model and reasoning effort for each request so teams can measure quality and spend at the task level.
Read post
Invoices arrive as scanned PDFs, prices live in rendered pages, and evidence sits in spreadsheets. AIVAX Fetch extracts readable text — and optionally typed JSON — from URLs and files through one endpoint, metered in processing units instead of model tokens.
Read post
Support triage usually means prompting a chat model and parsing its prose back into fields your code can use. AIVAX's Decisions API takes the message once and returns urgency, department, and frustration as typed answers — billed from the same account balance as inference, RAG, voice, and images.
Read post
A server-side research workflow should separate source discovery from source reading, preserve the URLs and extracted content as evidence, and evaluate each stage independently.
Read post
Coding agents lose context between sessions and guess at the world outside the repo. AIVAX's Collections and Web Utilities MCP servers give them writable semantic memory plus fetch and search — with scoped credentials, bounded retrieval, and explicit write control.
Read post
Secure persistent agent memory with scoped writes, expiry, provenance, and review. Understand AIVAX externalUserId boundaries and deletion limits.
Read post
Bulk AI work fails at the boundary between the queue and the provider: rate limits, balance, validation, and overload. AIVAX Batch answers with bounded admission, per-item validation, and failure-shaped retries.
Read post
Agentic Tests already scores every turn of a simulated conversation. Applied to real production traffic, the same trajectory signal — score, at-risk state, persistent loss — tells you when a live conversation is drifting before the user gives up.
Read post
MCP 2026-07-28 removes protocol sessions and transport replay, moving durable state, retries, long-running work, and compatibility into explicit application contracts.
Read post
A vector database handles vector search. RAG also needs source preparation, updates and context assembly. Compare what to build and what to manage.
Read post
Review MCP tools, credentials and imported instructions before connecting. Apply authorization and approval rules, with AIVAX configuration examples.
Read post
Map Assistants threads, runs and file search to Responses after the shutdown. Recover saved history and decide which state your application should own.
Read post
Compare OpenAI reasoning effort, Anthropic thinking and Gemini thinking controls. Configure AIVAX requests and handle summaries, state and token costs.
Read post
Use a separate moderation step to decide whether input may proceed before generation. System prompts guide behavior; they do not enforce permissions.
Read post
Define an observable goal, simulate a user, and judge the full conversation. Inspect weak turns, recovery, outcomes, and costs with AIVAX Agentic Tests.
Read post
Test a reranker when relevant documents rank too low. Fix missing candidates first, then measure answer quality, latency, and cost.
Read post
Reflex combines semantic relevance, lexical evidence and account-scoped cache reuse for recurring-document retrieval.
Read post