RAG vs vector database: is vector search enough?
A vector database handles vector search. RAG also needs source preparation, updates and context assembly. Compare what to build and what to manage.
A vector database stores and searches embeddings; RAG is the pipeline that retrieves relevant source material, assembles it as context, and uses a language model to generate an answer. A vector database alone is not enough for a complete RAG application. For changing files, you also need extraction, chunking, update handling, and answer evaluation; reranking is an optional step when useful candidates rank too low.
Choose based on which ingestion and maintenance tasks your team already operates. If your application already produces clean chunks and reliable change events, a vector database may be the only retrieval infrastructure you need to add. If it starts with PDFs, scans, or changing manuals, compare who prepares that material, keeps it current, and checks the evidence that reaches the model.
What does a vector database actually do?
At its core, a vector database stores vectors and associated data, builds indexes over those vectors, and executes similarity queries. pgvector, for example, adds vector similarity search to Postgres and supports exact search plus approximate indexes such as HNSW and IVFFlat. It gives an application a capable vector storage and query layer inside a familiar database.
That layer does not inherently know that a row came from page 37 of a revised employee handbook. It does not decide whether the page header is content or noise, whether a table should remain intact, whether a section needs to overlap its neighbor, or whether a deleted source file should remove twelve derived chunks. The application must either arrive with those decisions made or connect another service that makes them.
The unit mismatch explains much of the confusion. Source systems contain files, pages, tickets, records, media, permissions, and revision histories. Vector indexes contain searchable units with identifiers, vectors, text or payloads, and metadata. A production ingestion path has to translate one into the other without losing provenance, access rules, or update semantics.
If your application already produces clean, stable chunks, a vector database may be exactly the right abstraction. You keep control over parsing, embedding, deployment, and data movement while the database handles indexing and search. The mistake is treating that choice as if the surrounding responsibilities disappeared.
What must happen between a source file and a RAG answer?
“Managed RAG” is not a perfectly standardized product category. Services expose different connectors, parsers, stores, models, retrieval methods, and generation integrations. The useful definition is operational: the service coordinates a larger part of the lifecycle that turns source content into retrievable context and keeps that lifecycle running.
The diagram stops at selected context; a complete RAG application then sends that evidence to a language model and checks whether the answer is supported. Reranking is optional, and integrated database products can cover more than the highlighted index-and-search core.
For a changing corpus, the source-to-context path must run again when content or permissions change. Define how the application detects failed imports, retries safely, removes obsolete chunks, and traces retrieved evidence to a source revision. Buying managed retrieval only transfers the responsibilities its contract actually covers.
Do hybrid search and integrated reranking close the gap?
Some database products cover more than vector storage. Pinecone accepts precomputed vectors, while an index with integrated embedding can accept text records and convert them using its associated hosted model. Pinecone also documents hybrid search combining keyword and semantic signals and integrated reranking with hosted models.
Postgres can cover a similar retrieval boundary without a separate vector database service. The pgvector README's hybrid-search section describes combining pgvector with Postgres full-text search, using reciprocal rank fusion or a cross-encoder to combine results. That is a retrieval design you assemble, not an automatic source-file ingestion workflow.
Those features remove real integration work: with clean text records, stable identifiers, correct metadata, and an update stream, an application can use integrated embeddings to avoid a separate embedding client, combine semantic and lexical evidence through hybrid search, and rerank results without another deployment boundary.
They do not automatically connect to the system of record or determine how a PDF, wiki tree, database row, or video becomes those clean records. Nor do they establish deletion propagation, document-level access policy, source revision tracking, dead-letter handling, or evaluation criteria unless another documented feature supplies them. Apply the same scrutiny to managed RAG products because some still leave parsing or synchronization to the customer.
What should you compare in a managed RAG service?
| Task | Vector database | Managed RAG: what to verify |
|---|---|---|
| Extract and chunk | Usually receives prepared vectors or text records. | Supported file types, extraction fidelity, chunk controls, and reviewable output. |
| Embed and retrieve | Stores and searches vectors; may embed text or support hybrid search. | Embedding choices, retrieval methods, filters, and access boundaries. |
| Update and delete | Changes indexed records when instructed. | Source synchronization, stable IDs, deletion propagation, and retry behavior. |
| Rerank and generate | May rerank candidates; model context and generation can remain external. | Optional reranking, context assembly, source attribution, and generation integration. |
| Evaluate and rebuild | Database health does not establish answer quality. | Ingestion status, retrieval evidence, model-change handling, and rebuild cost. |
Product boundaries vary; verify each row rather than treating the service label as a guarantee. OpenAI File Search, for example, is a hosted tool that searches previously uploaded files through semantic and keyword search in vector stores. It supports a documented set of file types and metadata filtering. That is a larger managed surface than a bare vector index, but applications still decide which files to upload, how to assign metadata, when to replace them, and whether retrieved evidence is acceptable for the task.
Managed operation does not delegate product judgment. Your team still decides which identities may retrieve each document, which metadata is safe to expose, how much latency and ingestion cost the application can accept, and what evidence qualifies an answer as supported. It must also test chunking and retrieval with domain-specific questions, because a service can execute a configured policy reliably while that policy remains wrong for the corpus. Evaluation belongs outside the provider's health dashboard: measure source coverage, candidate recall, ranking quality, citation support, and access-control behavior against cases that represent the application.
Cost needs the same end-to-end view. Storage price alone omits extraction, embedding, repeated synchronization, query volume, reranking, and full-corpus rebuilds. Compare the bill across the lifecycle you will operate instead of extrapolating from an isolated demo query. Even when the provider runs the machinery, your team defines acceptable results.
How do you find what is missing from your RAG pipeline?
A nearest-neighbor query can be correct while the answer is wrong. The index may faithfully return a chunk that lost its heading during extraction, carries stale permissions, or omits the exception stated in the next paragraph. Retrieval cannot reconstruct information removed upstream.
Start before embeddings. Check whether extracted text preserves headings, table headers, numbers, and exceptions. Then inspect whether each chunk makes sense on its own. There is no universal chunk size that settles those questions. Use a small sample with known facts, as described in how to check PDF ingestion quality before embedding for RAG, before comparing retrieval models.
Synchronization can leave stale or contradictory chunks in the index. Stable identifiers must connect derived chunks to source revisions so updates and deletions do not leave contradictory evidence behind. Reindexing must distinguish a metadata-only change from text that needs a new embedding. If the embedding model changes, old and new vectors may not belong in the same index space. Do not assume a managed service automatically migrates them. Ask whether it pins model versions, coordinates reindexing, supports a staged rebuild, and charges for the work.
The retrieval handoff has its own failure mode: a reranker only sees candidates selected by the first stage. Our guide to semantic search and reranking explains how to diagnose whether the relevant document is missing or merely ordered too low. Managed reranking does not compensate for poor candidate recall, just as a managed index does not compensate for missing source content.
Observability must therefore cross stage boundaries. Useful evidence includes source revision, extraction and chunking status, embedding model, index timestamp, query and filters, retrieved candidates, reranked order, final context, latency, and cost. A green database health check covers only part of that chain.
How does AIVAX separate preparation, indexing, and retrieval?
AIVAX illustrates why product-specific boundaries matter. AIVAX Collections store persistent documents for semantic search. Direct and JSONL imports receive prepared document text with a stable name: name in the single-document API or docid in JSONL. When text changes, the document is queued for reindexing; a metadata-only update does not reindex the text.
Direct imports do not automatically chunk long documents. The RAG preparation guidance warns that oversized text can be truncated for embedding, so prepare focused documents before importing.
Choose the preparation path separately:
- Already prepared text: import it directly into a collection.
- Long text that needs boundaries: Text Segmentation returns segments to your application without creating embeddings or storing the submitted document. Review them before import.
- Source PDFs, images, audio, or video: the dashboard's Media Injector generates focused collection documents and queues them for indexing. This is generated knowledge, not a promise of verbatim transcription or complete coverage. Use reviewed text and direct import when exact wording matters.
Once documents are indexed, AIVAX Semantic Search searches one or more collections. Configurable reranking can adjust the order of retrieved candidates, while the standalone reranking API reorders candidate strings supplied by the application. Neither reranking path searches for text absent from its input.
Collections can also be attached to an AI Gateway so retrieved documents enter the model context. That connects retrieval to generation; it does not remove the need to evaluate source coverage and answer support.
When should you choose a vector database or managed RAG?
Choose the smallest managed boundary that removes work your team cannot reliably operate:
- Use a vector database or pgvector when you already own clean chunks, stable IDs, embeddings, access rules, and update events. You retain control, but also own ingestion failures and the generation layer.
- Evaluate managed RAG when file preparation and maintenance are the hard part. Require a documented import path, inspectable results, update and deletion behavior, and a way to retrieve context or pass it to a model. Verify which steps remain yours.
- Add a reranker to an existing retriever when the right evidence is present but ordered poorly. Compare ranking quality against the additional latency and cost before adopting it.
Resolve four questions that the feature list and responsibility matrix cannot answer for your workload:
- When the embedding model or version changes, who coordinates a staged reindex, validates compatibility, and accounts for rebuild cost?
- Can every retrieved chunk be traced to its exact source revision and the access policy enforced for that request?
- Can telemetry connect ingestion to retrieved and reranked context, and can your evaluation measure coverage, recall, ranking quality, citation support, and access-control behavior?
- During partial failure, model deprecation, or a full reindex, which retries, rollback paths, and availability guarantees apply?
Record these answers in the architecture decision and assign each remaining operation to a team before selecting a product.
FAQ
Do I need a vector database for RAG?
No. RAG requires retrieval, not a particular storage engine. A system can retrieve passages with full-text search and pass them to a language model. Vector search is an option for matching meaning beyond exact wording; test it against the questions and documents your application actually handles.
Can I do RAG with Postgres and pgvector?
Yes. Use pgvector to store embeddings and retrieve candidates, optionally combine it with Postgres full-text search, then pass selected passages to a language model. The pgvector README documents exact search, HNSW and IVFFlat indexes, and hybrid-search examples. You still need to implement or integrate source preparation, synchronization, context assembly, and generation.
When do I need a reranker?
Consider one when relevant documents reach the candidate set but rank too low to enter the final context. If they are absent, investigate ingestion, chunking, filters, and retrieval coverage first. A reranker cannot recover text outside its input. Validate the change with representative queries and known relevant documents.
Does managed RAG automatically chunk every upload?
Do not assume so. The import path determines that behavior. In AIVAX, direct document and JSONL imports expect prepared text; Text Segmentation returns segments for review, while Media Injector generates collection documents from supported source files. Choose the path deliberately and inspect its output.
References
- pgvector: vector similarity search and hybrid search for Postgres
- Pinecone: upsert records and integrated embeddings
- Pinecone: hybrid search
- Pinecone: rerank results
- OpenAI: File Search
- AIVAX Collections and Documents
- AIVAX Media Injector
- AIVAX Semantic Search
- AIVAX Text Segmentation
- AIVAX Best Practices for RAG
- AIVAX Rerankers