AIVAX / Blog

Notes from the operating layer.

Field notes on AI infrastructure, retrieval, model operations and the systems that keep every request visible.

27 published field notes

Latest blog posts

  1. Separate blue paper pieces beside a continuous folded yellow strip, with blue threads connecting the paper shapes

    AIVAX Labs

    OpenAI Evals shutdown: where do multi-turn agent tests go before November 30?

    Manually recreate OpenAI Evals datasets and graders in Promptfoo, then choose between its multi-turn tests and AIVAX gateway conversations.

    Read post
  2. A dense cluster of tiny pink glass seedpods narrows toward three larger coral pods on a rippled sage-green surface.

    AIVAX Labs

    Too many MCP tools: how to reduce tool-definition tokens in your agent's context

    Measure what MCP tool definitions cost, then cut it with allowlists, deferred tool search, or a shell. What each option needs and what AIVAX supports.

    Read post
  3. Five cut bamboo sections stacked beside curled paper in a misty bamboo grove

    AIVAX Labs

    How to convert a PDF to Markdown for RAG and split it into chunks

    Convert a PDF to Markdown with OCR, split it into chunks for retrieval, and export JSONL for a RAG collection. Code and a free demo included.

    Read post
  4. A row of weathered wooden stakes with pale bands stands in wet tidal sand at dusk, a lantern hanging from the nearest one, while driftwood and a small rowboat drift along curved currents

    AIVAX Labs

    OpenAI GPT-4, o1, o3-mini and o4-mini shutdown on October 23, 2026: what breaks and what to check

    OpenAI shuts down GPT-4, o1, o3-mini and o4-mini on October 23, 2026. Find hidden model IDs and check replacements for parameters, output and cost.

    Read post
  5. A red fabric ribbon opens into a broad sail and deep pleated folds against a charcoal background

    AIVAX Labs

    How to generate images and speech from Claude Code, Cursor or VS Code with MCP

    Configure AIVAX's hosted Media Generation MCP, discover models and prices, and generate images or MP3 speech with explicit tool and privacy limits.

    Read post
  6. Two dark teal glacial streams converge into one channel between textured white and pale blue ice banks, seen from directly above

    AIVAX Labs

    How to make LLM tool calls idempotent so a retry doesn't create a duplicate ticket

    Deduplicate the same tool operation; changed arguments create a new one. Require an explicit conflict policy to allow only one write per request.

    Read post
  7. Sunlit open-air concrete room with blank paper sheets clipped to slatted walls, diagonal bands of yellow light across them, and a blue ring on a stand in the foreground

    AIVAX Labs

    How to check PDF and scan ingestion quality before you embed for RAG

    Before embedding, sample real pages, write the facts you expect, and grade each generated document for wrong numbers, lost table headers, and missing facts.

    Read post
  8. An open paper folio connected by red thread to folded blue paper pieces on an ivory field

    AIVAX Labs

    How MCP skills carry a workflow across the server boundary

    SEP-2640 lets MCP servers offer discoverable workflow skills as resources. The host still decides what to load, verify, approve, and execute.

    Read post
  9. Four distinct botanical forms grow from one pale seed pod in a sunlit greenhouse

    AIVAX Labs

    Build a retrieval pipeline on AIVAX Free

    The expanded Free plan includes separate daily allowances for RAG embeddings, Reflex reranking, Julia-1 semantic decisions, and Fetch/OCR extraction. Here is what each covers—and what remains metered.

    Read post
  10. Rough cracked clay vessels on a wooden bench passing under a glowing brass caliper gauge, emerging as smooth measured vessels on a ledger grid

    AIVAX Labs

    How to fix invalid JSON from an LLM API with structured outputs

    Check truncation, refusals and schema support first. In AIVAX, distinguish native response_format from response_schema validation and bounded healing.

    Read post
  11. Three branching channels of water pass through separate gates before converging into a single measuring basin

    AIVAX Labs

    Route by complexity, price by the task

    A lower token rate does not tell you what a completed task costs. AIVAX's Complexity Router selects a model and reasoning effort for each request so teams can measure quality and spend at the task level.

    Read post
  12. A ceramic water mill turning a mixed stream of paper sheets, photo prints, and ledger pages into a single clear channel of plain text slips

    AIVAX Labs

    Fetch turns the messy web into text your code can use

    Invoices arrive as scanned PDFs, prices live in rendered pages, and evidence sits in spreadsheets. AIVAX Fetch extracts readable text — and optionally typed JSON — from URLs and files through one endpoint, metered in processing units instead of model tokens.

    Read post
  13. A brass sorting machine on a wooden desk sliding three blank paper tickets down chutes into wooden trays

    AIVAX Labs

    Stop parsing chat output: turn support messages into typed decisions

    Support triage usually means prompting a chat model and parsing its prose back into fields your code can use. AIVAX's Decisions API takes the message once and returns urgency, department, and frustration as typed answers — billed from the same account balance as inference, RAG, voice, and images.

    Read post
  14. An ice-white evidence table where pale-blue conduits connect many translucent search cards to three open source documents under glass mounts

    AIVAX Labs

    Research is a pipeline: search discovers, fetch reads

    A server-side research workflow should separate source discovery from source reading, preserve the URLs and extracted content as evidence, and evaluate each stage independently.

    Read post
  15. A glasshouse conservatory at dawn where labeled seedling trays feed a long irrigation channel that runs past open gardening manuals toward the glasshouse doors

    AIVAX Labs

    Give your coding agent a memory it can read and a world it can fetch

    Coding agents lose context between sessions and guess at the world outside the repo. AIVAX's Collections and Web Utilities MCP servers give them writable semantic memory plus fetch and search — with scoped credentials, bounded retrieval, and explicit write control.

    Read post
  16. A weathered wall of card-catalog drawers with blank brass label plates, one drawer half-open with a slipped card tied by a single red thread to a small brass desk lamp

    AIVAX Labs

    How to protect LLM agent memory from poisoning

    Secure persistent agent memory with scoped writes, expiry, provenance, and review. Understand AIVAX externalUserId boundaries and deletion limits.

    Read post
  17. A stone sluice gate with three channels meters a river of record-cards into calm terraced basins while a gatekeeper works a lever

    AIVAX Labs

    Batch is an admission-control problem, not a queue

    Bulk AI work fails at the boundary between the queue and the provider: rate limits, balance, validation, and overload. AIVAX Batch answers with bounded admission, per-item validation, and failure-shaped retries.

    Read post
  18. Terraced night gardens crossed by thin diagram paths with node dots, one ember-red path veering away past a small stone watchtower

    AIVAX Labs

    From test to monitor: the judge trajectory as a production signal

    Agentic Tests already scores every turn of a simulated conversation. Applied to real production traffic, the same trajectory signal — score, at-risk state, persistent loss — tells you when a live conversation is drifting before the user gives up.

    Read post
  19. Blue request cards, each carrying a yellow state token, travel independently toward three red server blocks in a textured screen print

    AIVAX Labs

    MCP went stateless. Your application did not

    MCP 2026-07-28 removes protocol sessions and transport replay, moving durable state, retries, long-running work, and compatibility into explicit application contracts.

    Read post
  20. Botanical flat lay in cream, pink, and green with pressed leaves, handmade papers, bowls, and a vessel of grains

    AIVAX Labs

    RAG vs vector database: is vector search enough?

    A vector database handles vector search. RAG also needs source preparation, updates and context assembly. Compare what to build and what to manage.

    Read post
  21. Translucent capsules carried on deep-blue currents pass through gated openings in a porous white boundary into calm terraced channels

    AIVAX Labs

    How to safely connect third-party MCP servers

    Review MCP tools, credentials and imported instructions before connecting. Apply authorization and approval rules, with AIVAX configuration examples.

    Read post
  22. A scaffold bridge carries seed pods, glass vessels and archives from a root-bound grove monolith across a gorge into an open modular garden

    AIVAX Labs

    Migrate OpenAI Assistants threads, runs and file search to the Responses API

    Map Assistants threads, runs and file search to Responses after the shutdown. Recover saved history and decide which state your application should own.

    Read post
  23. A precision glass-and-metal translator separating layered signals into visible and opaque output channels

    AIVAX Labs

    How to set reasoning effort across providers: OpenAI, Anthropic and Gemini

    Compare OpenAI reasoning effort, Anthropic thinking and Gemini thinking controls. Configure AIVAX requests and handle summaries, state and token costs.

    Read post
  24. A monumental stone gateway filtering irregular colored fragments before they enter a calm central chamber

    AIVAX Labs

    LLM input moderation: system prompt or separate moderation step?

    Use a separate moderation step to decide whether input may proceed before generation. System prompts guide behavior; they do not enforce permissions.

    Read post
  25. A tactile tabletop evaluation with three parallel conversation tracks converging on a final scored outcome

    AIVAX Labs

    How to test multi-turn AI agent conversations with a simulated user and an LLM judge

    Define an observable goal, simulate a user, and judge the full conversation. Inspect weak turns, recovery, outcomes, and costs with AIVAX Agentic Tests.

    Read post
  26. A broad field of document cards narrowing into a candidate set and then an ordered shortlist

    AIVAX Labs

    Do I Need a Reranker for RAG? Semantic Search vs. Reranking

    Test a reranker when relevant documents rank too low. Fix missing candidates first, then measure answer quality, latency, and cost.

    Read post
  27. Smooth stones arranged in decreasing size on warm sand, representing recurring document retrieval

    AIVAX Labs

    Reflex: retrieval built for recurring documents

    Reflex combines semantic relevance, lexical evidence and account-scoped cache reuse for recurring-document retrieval.

    Read post