Why Vector Search Became the Backbone of Modern AI Pipelines
Generative models are excellent at producing plausible output and terrible at remembering specifics. A language model can write a fluent summary but cannot tell you what your team decided in last week's planning doc. A video model can render a convincing street scene but cannot match the exact lighting of the clip you shot on Tuesday. Retrieval is how teams close that gap, and retrieval at scale almost always means vectors.
A vector database stores numeric representations of content — embeddings — and finds the nearest neighbours to a query embedding in milliseconds. Instead of matching keywords, it matches meaning. A search for "how do I reduce render time" lands next to a document titled "cutting encode latency" even though the two share no words at all.
That single capability now powers three very different workloads: assistants that ground language models in private documents, media pipelines that search and assemble footage by visual similarity, and financial systems that need to react to anomalies or sentiment shifts quickly. Open source engines matter here because they can be self-hosted, audited, and tuned — exactly what regulated and latency-sensitive teams require.
What an Open Source Vector Database Actually Does
A vector database is three layers stacked together: the embedding contract, the index, and a query layer with filtering. Treating them separately prevents most early architecture mistakes.
Embeddings, distance metrics, and index types
An embedding is a fixed-length array of floats — commonly 384, 768, 1024, or 1536 dimensions — produced by a sentence transformer, a CLIP-style image encoder, or a multimodal model. The model you pick defines what "similar" means. Two documents on the same topic become neighbours; two frames with the same colour grading become neighbours; a query in Portuguese can retrieve a German document if the model is multilingual.
Distance metrics follow from how the model was trained. Cosine similarity is the safe default for normalised text embeddings. Dot product works when magnitude carries information. Euclidean distance shows up often in image and audio pipelines. Pairing the wrong metric with a model is a silent accuracy killer: results still return, they are just quietly worse.
Indexes trade recall for speed and memory. HNSW builds a navigable graph and delivers strong recall at high throughput, at the cost of RAM. IVF with product quantization clusters and compresses vectors, which suits billion-scale collections on modest hardware. DiskANN keeps most of the index on SSD for large, cost-sensitive deployments. Scalar and binary quantization can shrink memory by 4x to 32x with a small recall penalty that is often recovered by reranking top candidates in full precision.
Where traditional databases stop
If your collection fits in a few million vectors and traffic is modest, an extension such as pgvector inside PostgreSQL is usually the right call: one system to operate, transactional consistency, and familiar SQL joins against business data. Purpose-built engines earn their place when you need horizontal sharding, per-tenant isolation, sub-50 ms p95 latency under heavy concurrency, or index types a relational engine does not expose. The honest answer for most teams is to start boring and migrate only when a measured bottleneck says so.
Choosing an Engine: A Practical Decision Framework
| Engine | Strongest fit | Watch out for |
|---|---|---|
| pgvector / pgvectorscale | Teams already on PostgreSQL; up to low tens of millions of vectors | Horizontal scale and very high QPS need extra work |
| Qdrant | Filter-heavy workloads, per-tenant isolation, Rust-level performance | Fewer managed options than the biggest clouds |
| Milvus | Billion-scale collections, multiple index types, GPU indexes | Operationally heavy; needs a real platform team |
| Weaviate | Built-in hybrid search and modules for vectorising content | Module abstraction can obscure tuning knobs |
| Chroma | Prototypes, notebooks, small local apps | Not built for large concurrent production loads |
| LanceDB | Embedded, columnar, file-based; great for local media indexes | Smaller ecosystem around advanced filtering |
| FAISS | Research, benchmarks, custom index experimentation | A library, not a service — you build the rest |
Three questions cut through most of the comparison. First, where does the data already live? If it is in Postgres and the volume is reasonable, avoiding a second datastore is worth more than a marginal latency win. Second, how complex are your filters? Metadata filtering — by user, date range, language, rights status, shot type — is where many engines diverge sharply in real-world performance. Third, who operates it? A self-hosted cluster you cannot monitor is worse than a simpler index you understand.
Run a bake-off with your own data. Load 100k to 1M real vectors, replay 200 real queries, and measure recall@10 against brute-force ground truth alongside p95 latency and memory per million vectors. That afternoon of work beats any benchmark chart.
Building a Retrieval Pipeline Step by Step
Chunking, metadata, and stable IDs
Chunking decides your ceiling. Too large and each chunk mixes topics, diluting the embedding. Too small and you lose the context that made the passage meaningful. For prose, 300–800 tokens with 10–20% overlap is a reasonable starting point; for transcripts, split on speaker turns and keep timestamps in metadata. For video, chunk at the shot level rather than the frame level, and store the frame range alongside the vector.
Metadata is not optional. Every record should carry a stable ID, a source pointer, a timestamp, a language, and whatever access-control field your product needs. Without a rights or tenant field in metadata, you cannot filter safely later without a painful reindex.
Hybrid search and reranking
Pure vector search struggles with exact strings: product codes, ticker symbols, names, error codes. Hybrid search combines a lexical score (BM25 or similar) with the vector score and fuses them, often with reciprocal rank fusion. A reranker — a cross-encoder that reads the query and candidate together — then reorders the top 50 to 100 results. In practice, hybrid plus reranking is the single biggest quality improvement most teams can make, larger than switching engines.
Evaluate before you scale
Build a small golden set: 100 to 300 queries with known correct answers. Track recall@k, mean reciprocal rank, and answer correctness downstream. Then change one variable at a time — chunk size, embedding model, fusion weights. Without this set, every tuning decision is a guess, and regressions arrive silently when you swap a model.
Applying Vector Search to AI Video Workflows
Video is where semantic retrieval becomes genuinely transformative, because editorial decisions that used to require scrubbing through hours of footage become search queries.
Shot and scene retrieval
Encode every shot with an image or video encoder and store the resulting vector with metadata: duration, resolution, camera movement, timecode, rights status, and a transcript excerpt. A director can then ask for "wide shots at golden hour with no people" or "close-ups where the subject is off-centre" and get ranked candidates in seconds. Storyboard matching works the same way: encode the reference image, find the nearest existing shots, and hand editors a shortlist instead of a password-protected folder of raw files.
Style consistency through frame embeddings
Generative video pipelines drift. Colour, grain, and lighting shift between clips, and consistency is what separates a usable sequence from an obvious AI collage. Frame embeddings give you a measurable handle on that drift: compute a style vector for your reference clip, compute vectors for each generated clip, and flag anything whose distance exceeds a threshold. You can then apply a colour transform, regenerate, or simply reject the outlier before it reaches the timeline.
Asset search, deduplication, and quality control
Beyond creative search, embeddings are a maintenance tool. Near-duplicate detection removes the same clip uploaded three times under different filenames. Anomaly detection surfaces corrupted frames, black frames, or shots that do not match their own transcript. Logo and product detection becomes a similarity query against a small reference set rather than a hand-built detector. All of these run on the same index you already maintain for search, which is why a single well-structured vector store tends to pay for itself across several features.
Vector Search in Finance and Other Regulated Workloads
Sentiment and news retrieval
Instead of scoring a fixed keyword list, encode news, filings, and research notes, then retrieve the passages semantically closest to your current positions or watchlist. A query like "supply chain disruption affecting semiconductor margins" surfaces relevant coverage across languages and phrasing, and the retrieved set becomes grounded input for a summarisation model. The advantage over keyword screening is coverage: paraphrases, jargon, and translations all land in the same neighbourhood.
Anomaly and fraud detection
Embeddings of transactions, device fingerprints, or session behaviour make unusual patterns easy to find. Build a baseline neighbourhood of normal behaviour per account, then flag events whose nearest neighbours are far away or whose neighbourhood composition has shifted. This complements rules-based systems: rules catch known patterns, vectors catch the ones nobody wrote a rule for yet. The output is a ranked queue for human review rather than an automatic verdict, which keeps the workflow defensible.
Recommendation and personalisation
Product or content embeddings make recommendations a nearest-neighbour query with filters for eligibility, risk profile, and jurisdiction. Because the index supports metadata filtering, a single query can enforce suitability rules and still return semantically relevant results. That combination — semantic ranking plus hard constraints — is difficult to achieve cleanly with collaborative filtering alone.
Auditability and data sovereignty
Self-hosted engines keep sensitive records inside your own perimeter, and open source indexes can be reproduced, versioned, and explained. Store the embedding model name and version with every record, log queries and returned IDs, and you can reconstruct why a system surfaced a given document months later. That traceability is often the difference between a pilot that clears compliance review and one that stalls.
Performance, Cost, and Scaling Considerations
Memory dominates cost for most deployments. A million 1536-dimension float32 vectors occupy roughly 6 GB before index overhead, so quantization decisions have an immediate budget impact. Binary quantization with float32 reranking frequently cuts memory by an order of magnitude while keeping recall within a couple of percentage points, which is more than acceptable when a reranker sits downstream.
Latency budgets should be set end to end, not per component. If a request includes an embedding call, a vector search, a reranker, and a generation step, the vector search is usually the cheapest of the four. Optimising it from 40 ms to 20 ms feels satisfying but rarely changes user-perceived speed; caching embeddings for repeated queries or prefetching during typing does.
Scaling strategy depends on shape. Read-heavy search scales with replicas and caching. Write-heavy ingestion scales with sharding and batch upserts — never insert vectors one at a time in a loop if you can batch a thousand. Multi-tenant products usually want either a collection per tenant or a strict partition key, and you should decide which before you have customers, not after.
Common Mistakes and How to Avoid Them
Changing the embedding model without reindexing. Vectors from different models are not comparable. Version your embeddings and treat a model swap as a full migration, ideally with a shadow index you can compare against.
Ignoring filters until production. Filtering by tenant, date, or rights status changes query plans dramatically. Test filtered recall from day one.
Storing raw text only in the vector store. Keep the canonical content in your primary database and store a reference plus metadata in the index. Vector stores are search engines, not systems of record.
Tuning on vibes. Without a golden query set, nobody can prove an improvement. Assemble one before you tune anything.
Over-engineering the first version. A single-node instance with a sensible schema and hybrid search will outperform a poorly understood distributed cluster almost every time.
FAQ
Do I need a vector database, or can I just use a search engine? If your queries are keyword-driven and content is text only, a traditional search engine with semantic reranking may be enough. Once you need cross-modal search, fine-grained metadata filters, or millions of high-dimensional vectors, a purpose-built index pays off.
How many vectors before pgvector stops being enough? Many teams run comfortably into the low tens of millions with proper indexing and hardware. The breaking point is usually concurrency and filtered-query latency, not raw count.
Which embedding model should I start with? Pick a well-supported multilingual text model for documents and a CLIP-style or modern image-text model for visual content. Compare two candidates on your own golden set before committing.
How often should I re-embed? Re-embed when you change models, when the underlying content changes materially, or on a scheduled cadence for fast-moving corpora. Incremental re-embedding of changed records is usually sufficient.
Can one index serve both video and text? Yes, if you use a shared embedding space or store parallel vectors per item. Multimodal encoders make this practical, and metadata filters keep the two workloads from interfering.
What does a realistic pilot look like? One use case, one index, real data, a golden query set, and a measured baseline. Expand to other features only after the first one shows a clear quality or speed win.
A Thirty-Day Adoption Plan
Week one: pick a single high-value use case and gather 1,000 to 10,000 real records. Week two: stand up one engine — pgvector if you already run Postgres, Qdrant or LanceDB if you want a purpose-built or embedded option — and build a baseline index with a chunking strategy you can articulate. Week three: add hybrid search and a reranker, then measure recall against a small golden set. Week four: put a thin interface in front of it — an internal search page, an editor-facing shot finder, or an analyst queue — and let real users break it.
The teams that get the most from vector search are rarely the ones with the most elaborate architecture. They are the ones who versioned their embeddings, wrote down their evaluation criteria, and shipped one narrowly scoped feature to real users before expanding.


