<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>AIVAX Blog</title>
    <link>https://aivax.net/blog/</link>
    <description>Field notes on AI infrastructure, retrieval and model operations.</description>
    <language>en</language>
    <atom:link href="https://aivax.net/blog/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>OpenAI Evals shutdown: where do multi-turn agent tests go before November 30?</title>
      <link>https://aivax.net/blog/openai-evals-shutdown-multi-turn-agent-tests/</link>
      <guid isPermaLink="true">https://aivax.net/blog/openai-evals-shutdown-multi-turn-agent-tests/</guid>
      <pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate>
      <description>Manually recreate OpenAI Evals datasets and graders in Promptfoo, then choose between its multi-turn tests and AIVAX gateway conversations.</description>
    </item>
    <item>
      <title>Too many MCP tools: how to reduce tool-definition tokens in your agent&apos;s context</title>
      <link>https://aivax.net/blog/reduce-mcp-tool-definition-tokens-too-many-tools/</link>
      <guid isPermaLink="true">https://aivax.net/blog/reduce-mcp-tool-definition-tokens-too-many-tools/</guid>
      <pubDate>Thu, 08 Oct 2026 00:00:00 GMT</pubDate>
      <description>Measure what MCP tool definitions cost, then cut it with allowlists, deferred tool search, or a shell. What each option needs and what AIVAX supports.</description>
    </item>
    <item>
      <title>How to convert a PDF to Markdown for RAG and split it into chunks</title>
      <link>https://aivax.net/blog/pdf-to-markdown-for-rag-chunks-jsonl/</link>
      <guid isPermaLink="true">https://aivax.net/blog/pdf-to-markdown-for-rag-chunks-jsonl/</guid>
      <pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate>
      <description>Convert a PDF to Markdown with OCR, split it into chunks for retrieval, and export JSONL for a RAG collection. Code and a free demo included.</description>
    </item>
    <item>
      <title>OpenAI GPT-4, o1, o3-mini and o4-mini shutdown on October 23, 2026: what breaks and what to check</title>
      <link>https://aivax.net/blog/pin-llm-model-id-or-use-alias-model-deprecations/</link>
      <guid isPermaLink="true">https://aivax.net/blog/pin-llm-model-id-or-use-alias-model-deprecations/</guid>
      <pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate>
      <description>OpenAI shuts down GPT-4, o1, o3-mini and o4-mini on October 23, 2026. Find hidden model IDs and check replacements for parameters, output and cost.</description>
    </item>
    <item>
      <title>How to generate images and speech from Claude Code, Cursor or VS Code with MCP</title>
      <link>https://aivax.net/blog/generate-images-and-speech-from-claude-code-with-mcp/</link>
      <guid isPermaLink="true">https://aivax.net/blog/generate-images-and-speech-from-claude-code-with-mcp/</guid>
      <pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate>
      <description>Configure AIVAX&apos;s hosted Media Generation MCP, discover models and prices, and generate images or MP3 speech with explicit tool and privacy limits.</description>
    </item>
    <item>
      <title>How to make LLM tool calls idempotent so a retry doesn&apos;t create a duplicate ticket</title>
      <link>https://aivax.net/blog/make-llm-tool-calls-idempotent-protocol-functions/</link>
      <guid isPermaLink="true">https://aivax.net/blog/make-llm-tool-calls-idempotent-protocol-functions/</guid>
      <pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
      <description>Deduplicate the same tool operation; changed arguments create a new one. Require an explicit conflict policy to allow only one write per request.</description>
    </item>
    <item>
      <title>How to check PDF and scan ingestion quality before you embed for RAG</title>
      <link>https://aivax.net/blog/check-pdf-ingestion-quality-before-embedding-for-rag/</link>
      <guid isPermaLink="true">https://aivax.net/blog/check-pdf-ingestion-quality-before-embedding-for-rag/</guid>
      <pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
      <description>Before embedding, sample real pages, write the facts you expect, and grade each generated document for wrong numbers, lost table headers, and missing facts.</description>
    </item>
    <item>
      <title>How MCP skills carry a workflow across the server boundary</title>
      <link>https://aivax.net/blog/mcp-skills-extension-workflows/</link>
      <guid isPermaLink="true">https://aivax.net/blog/mcp-skills-extension-workflows/</guid>
      <pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate>
      <description>SEP-2640 lets MCP servers offer discoverable workflow skills as resources. The host still decides what to load, verify, approve, and execute.</description>
    </item>
    <item>
      <title>Build a retrieval pipeline on AIVAX Free</title>
      <link>https://aivax.net/blog/free-tier-for-rag-reflex-decisions-and-ocr/</link>
      <guid isPermaLink="true">https://aivax.net/blog/free-tier-for-rag-reflex-decisions-and-ocr/</guid>
      <pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate>
      <description>The expanded Free plan includes separate daily allowances for RAG embeddings, Reflex reranking, Julia-1 semantic decisions, and Fetch/OCR extraction. Here is what each covers—and what remains metered.</description>
    </item>
    <item>
      <title>How to fix invalid JSON from an LLM API with structured outputs</title>
      <link>https://aivax.net/blog/structured-output-healing-boundary/</link>
      <guid isPermaLink="true">https://aivax.net/blog/structured-output-healing-boundary/</guid>
      <pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate>
      <description>Check truncation, refusals and schema support first. In AIVAX, distinguish native response_format from response_schema validation and bounded healing.</description>
    </item>
    <item>
      <title>Route by complexity, price by the task</title>
      <link>https://aivax.net/blog/route-by-complexity-price-by-task/</link>
      <guid isPermaLink="true">https://aivax.net/blog/route-by-complexity-price-by-task/</guid>
      <pubDate>Tue, 22 Sep 2026 00:00:00 GMT</pubDate>
      <description>A lower token rate does not tell you what a completed task costs. AIVAX&apos;s Complexity Router selects a model and reasoning effort for each request so teams can measure quality and spend at the task level.</description>
    </item>
    <item>
      <title>Fetch turns the messy web into text your code can use</title>
      <link>https://aivax.net/blog/fetch-messy-web-to-text/</link>
      <guid isPermaLink="true">https://aivax.net/blog/fetch-messy-web-to-text/</guid>
      <pubDate>Mon, 21 Sep 2026 00:00:00 GMT</pubDate>
      <description>Invoices arrive as scanned PDFs, prices live in rendered pages, and evidence sits in spreadsheets. AIVAX Fetch extracts readable text — and optionally typed JSON — from URLs and files through one endpoint, metered in processing units instead of model tokens.</description>
    </item>
    <item>
      <title>Stop parsing chat output: turn support messages into typed decisions</title>
      <link>https://aivax.net/blog/support-triage-typed-decisions/</link>
      <guid isPermaLink="true">https://aivax.net/blog/support-triage-typed-decisions/</guid>
      <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
      <description>Support triage usually means prompting a chat model and parsing its prose back into fields your code can use. AIVAX&apos;s Decisions API takes the message once and returns urgency, department, and frustration as typed answers — billed from the same account balance as inference, RAG, voice, and images.</description>
    </item>
    <item>
      <title>Research is a pipeline: search discovers, fetch reads</title>
      <link>https://aivax.net/blog/research-is-a-pipeline-search-discovers-fetch-reads/</link>
      <guid isPermaLink="true">https://aivax.net/blog/research-is-a-pipeline-search-discovers-fetch-reads/</guid>
      <pubDate>Sat, 12 Sep 2026 00:00:00 GMT</pubDate>
      <description>A server-side research workflow should separate source discovery from source reading, preserve the URLs and extracted content as evidence, and evaluate each stage independently.</description>
    </item>
    <item>
      <title>Give your coding agent a memory it can read and a world it can fetch</title>
      <link>https://aivax.net/blog/upgrade-your-harness-with-aivax-mcps/</link>
      <guid isPermaLink="true">https://aivax.net/blog/upgrade-your-harness-with-aivax-mcps/</guid>
      <pubDate>Wed, 09 Sep 2026 00:00:00 GMT</pubDate>
      <description>Coding agents lose context between sessions and guess at the world outside the repo. AIVAX&apos;s Collections and Web Utilities MCP servers give them writable semantic memory plus fetch and search — with scoped credentials, bounded retrieval, and explicit write control.</description>
    </item>
    <item>
      <title>How to protect LLM agent memory from poisoning</title>
      <link>https://aivax.net/blog/persistent-memory-is-a-write-path/</link>
      <guid isPermaLink="true">https://aivax.net/blog/persistent-memory-is-a-write-path/</guid>
      <pubDate>Mon, 07 Sep 2026 00:00:00 GMT</pubDate>
      <description>Secure persistent agent memory with scoped writes, expiry, provenance, and review. Understand AIVAX externalUserId boundaries and deletion limits.</description>
    </item>
    <item>
      <title>Batch is an admission-control problem, not a queue</title>
      <link>https://aivax.net/blog/batch-is-an-admission-control-problem-not-a-queue/</link>
      <guid isPermaLink="true">https://aivax.net/blog/batch-is-an-admission-control-problem-not-a-queue/</guid>
      <pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate>
      <description>Bulk AI work fails at the boundary between the queue and the provider: rate limits, balance, validation, and overload. AIVAX Batch answers with bounded admission, per-item validation, and failure-shaped retries.</description>
    </item>
    <item>
      <title>From test to monitor: the judge trajectory as a production signal</title>
      <link>https://aivax.net/blog/judge-trajectory-production-signal/</link>
      <guid isPermaLink="true">https://aivax.net/blog/judge-trajectory-production-signal/</guid>
      <pubDate>Fri, 04 Sep 2026 00:00:00 GMT</pubDate>
      <description>Agentic Tests already scores every turn of a simulated conversation. Applied to real production traffic, the same trajectory signal — score, at-risk state, persistent loss — tells you when a live conversation is drifting before the user gives up.</description>
    </item>
    <item>
      <title>MCP went stateless. Your application did not</title>
      <link>https://aivax.net/blog/mcp-stateless-core-moves-session-state-into-your-application/</link>
      <guid isPermaLink="true">https://aivax.net/blog/mcp-stateless-core-moves-session-state-into-your-application/</guid>
      <pubDate>Wed, 02 Sep 2026 00:00:00 GMT</pubDate>
      <description>MCP 2026-07-28 removes protocol sessions and transport replay, moving durable state, retries, long-running work, and compatibility into explicit application contracts.</description>
    </item>
    <item>
      <title>RAG vs vector database: is vector search enough?</title>
      <link>https://aivax.net/blog/a-vector-database-is-not-a-rag-system/</link>
      <guid isPermaLink="true">https://aivax.net/blog/a-vector-database-is-not-a-rag-system/</guid>
      <pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate>
      <description>A vector database handles vector search. RAG also needs source preparation, updates and context assembly. Compare what to build and what to manage.</description>
    </item>
    <item>
      <title>How to safely connect third-party MCP servers</title>
      <link>https://aivax.net/blog/mcp-is-a-trust-boundary-not-just-a-tool-catalog/</link>
      <guid isPermaLink="true">https://aivax.net/blog/mcp-is-a-trust-boundary-not-just-a-tool-catalog/</guid>
      <pubDate>Fri, 28 Aug 2026 00:00:00 GMT</pubDate>
      <description>Review MCP tools, credentials and imported instructions before connecting. Apply authorization and approval rules, with AIVAX configuration examples.</description>
    </item>
    <item>
      <title>Migrate OpenAI Assistants threads, runs and file search to the Responses API</title>
      <link>https://aivax.net/blog/migrating-from-openai-assistants-without-rebuilding-the-same-coupling/</link>
      <guid isPermaLink="true">https://aivax.net/blog/migrating-from-openai-assistants-without-rebuilding-the-same-coupling/</guid>
      <pubDate>Thu, 27 Aug 2026 00:00:00 GMT</pubDate>
      <description>Map Assistants threads, runs and file search to Responses after the shutdown. Recover saved history and decide which state your application should own.</description>
    </item>
    <item>
      <title>How to set reasoning effort across providers: OpenAI, Anthropic and Gemini</title>
      <link>https://aivax.net/blog/reasoning-is-a-protocol-not-just-a-model-setting/</link>
      <guid isPermaLink="true">https://aivax.net/blog/reasoning-is-a-protocol-not-just-a-model-setting/</guid>
      <pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate>
      <description>Compare OpenAI reasoning effort, Anthropic thinking and Gemini thinking controls. Configure AIVAX requests and handle summaries, state and token costs.</description>
    </item>
    <item>
      <title>LLM input moderation: system prompt or separate moderation step?</title>
      <link>https://aivax.net/blog/aivax-gateway-moderation/</link>
      <guid isPermaLink="true">https://aivax.net/blog/aivax-gateway-moderation/</guid>
      <pubDate>Mon, 24 Aug 2026 00:00:00 GMT</pubDate>
      <description>Use a separate moderation step to decide whether input may proceed before generation. System prompts guide behavior; they do not enforce permissions.</description>
    </item>
    <item>
      <title>How to test multi-turn AI agent conversations with a simulated user and an LLM judge</title>
      <link>https://aivax.net/blog/introducing-agentic-tests/</link>
      <guid isPermaLink="true">https://aivax.net/blog/introducing-agentic-tests/</guid>
      <pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate>
      <description>Define an observable goal, simulate a user, and judge the full conversation. Inspect weak turns, recovery, outcomes, and costs with AIVAX Agentic Tests.</description>
    </item>
    <item>
      <title>Do I Need a Reranker for RAG? Semantic Search vs. Reranking</title>
      <link>https://aivax.net/blog/semantic-search-vs-reranking/</link>
      <guid isPermaLink="true">https://aivax.net/blog/semantic-search-vs-reranking/</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
      <description>Test a reranker when relevant documents rank too low. Fix missing candidates first, then measure answer quality, latency, and cost.</description>
    </item>
    <item>
      <title>Reflex: retrieval built for recurring documents</title>
      <link>https://aivax.net/blog/reflex-retrieval-built-for-recurring-documents/</link>
      <guid isPermaLink="true">https://aivax.net/blog/reflex-retrieval-built-for-recurring-documents/</guid>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
      <description>Reflex combines semantic relevance, lexical evidence and account-scoped cache reuse for recurring-document retrieval.</description>
    </item>
  </channel>
</rss>
