Retrieval and vector

Elasticsearch

Lexical search engine that also stores vectors.

made by
Elastic

Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Elastic - we just build with this.

What Elasticsearch is

Elasticsearch is a distributed search engine built on Lucene. It does BM25 keyword search, structured filtering, and aggregations, and it also stores dense vectors with approximate nearest-neighbour search, so hybrid retrieval can happen in one query. OpenSearch is the AWS-maintained fork of the pre-2021 codebase.

How we use it

When a client already runs Elasticsearch we index into it rather than adding a vector database beside it, and combine BM25 with kNN so exact strings - part numbers, statute references, error codes - still hit. Metadata filters are applied as hard constraints rather than as post-filters, which is what keeps permission-scoped retrieval correct. Analyzers get tuned per field, because tokenization decides what BM25 can ever match.

Where it is the wrong choice

It is heavy to operate. Heap tuning, shard sizing, and cluster upgrades are real work for a corpus that would fit inside Postgres, and for a few million chunks pgvector next to the source data is far less to run.

Building something on Elasticsearch?

Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherElasticsearch is even the right call for it.

Start a project

Elasticsearch and Elastic are trademarks of their respective owners, used here to say what we work with.