Today, we’re announcing metadata pre-filtering for Amazon S3 Vectors , which delivers higher recall on filtered queries by evaluating your metadata filter before the similarity search. You can filter on attributes such as tenant, category, status, or time, and pre-filtering adds prefix matching with $startsWith for paths, URLs, and hierarchical keys. Each vector carries up to 2 KB of filterable metadata, and a single query supports up to 100 filter constraints.
There is no additional cost, no re-ingestion, and no change to your queries. Most applications never search a whole index. They search the part of it that belongs to a particular user, account, or category, and they express that scope as a metadata filter.
Semantic search, retrieval-augmented generation (RAG), and agentic applications all need the same thing from a filtered query: a similarity search that covers the vectors matching the filter, and returns the closest of them. With pre-filtering, a filtered query returns more of the relevant matches your index contains, giving you higher recall on filtered searches.
Common use cases Pre-filtering applies wherever results have to be both relevant and correctly scoped: Legal and professional services : A law firm or e-discovery platform searches documents scoped to a single client, and with $startsWith narrows further by matter number, folder path, or document ID prefix. A single client is a small share of a firm-wide archive, and filters this narrow are where pre-filtering improves recall most.
Financial services : An investment research platform searches analyst notes, filings, and call transcripts scoped by issuer, document type, and publication date. Media and entertainment : A streaming service filters by content rating and regional licensing before the semantic search, finding similar titles restricted to G and PG content licensed in one territory.
Agentic applications : An agent working within a user’s session filters on fields such as owner, document set, and timestamp so its searches cover the material relevant to the task at hand. Higher recall means more of that material reaches the agent, which improves task reliability How pre-filtering works Each vector in an S3 Vectors index can carry application-defined metadata, and a query can filter on those fields.
Every vector index has an index mode. On an index whose index mode is ENHANCED , S3 Vectors resolves your filter first, then searches only the vectors that match. On an index whose index mode is CLASSIC , S3 Vectors performs the vector search and filter evaluation in tandem, validating each candidate vector against your filter as it searches.
Existing indexes use CLASSIC until you update them. Consider a support knowledge base of 8 million tickets, where an agent searches one customer’s history for a recurring error. If that customer accounts for 400 of those tickets, resolving customer_id first means the similarity search runs across all 400 of them, so the agent sees that customer’s prior occurrences.
Before the index was updated, the same query drew its candidates from the full 8 million, and the result set contained fewer of that customer’s matching tickets. On highly selective filters, pre-filtering returns up to 5x more of the matching vectors than the same query returned before on CLASSIC indexes. Getting started Before you start, make sure your IAM policy grants permissions for the new actions.
You can get started in three steps. The walkthrough below builds a small product-catalog index and runs a selective filter against it, the same pattern you would use for a multi-tenant RAG store or a document search scoped to one client. First, create a vector index: aws s3vectors create-index \ --index-name product-catalog \ --vector-bucket-name my-vector-bucket \ --dimension 1536 \ --distance-metric cosine The dimension must match the output size of your embedding model, and distance-metric should match how that model was trained (cosine is common for text embeddings).
Second, write vectors with the PutVectors API, attaching up to 2 KB of filterable metadata to each vector: aws s3vectors put-vectors \ --index-name product-catalog \ --vector-bucket-name my-vector-bucket \ --vectors '[{ "key": "doc-001", "data": {"float32": [0. 1, 0. 2, 0.
3, ...] }, "metadata": { "tenant_id": "t-10428", "category": "legal", "created_date": "2026-03-15", "active": true } }]' Each vector carries the attributes your application filters on. In this example, tenant_id scopes results to a single customer, category narrows by document type, created_date records when the document was created, and active is a boolean flag.
By default every metadata field is filterable, so you can query on any of them without declaring a schema up front. Third, run a filtered similarity query with the QueryVectors API. The filter uses a compact JSON syntax where a bare key-value pair is an equality match, and operators such as $and , $or , and $gt combine or refine conditions.
Pass --return-metadata so the query returns each vector’s metadata: aws s3vectors query-vectors \ --index-name product-catalog \ --vector-bucket-name my-vector-bucket \ --query-vector '{"float32": [0. 1, 0. 2, 0.
3, ...] }' \ --top-k 50 \ --return-metadata \ --filter '{"$and": [ {"tenant_id": "t-10428"}, {"category": "legal"}, {"active": true} ]}' The expected result is a single vector, doc-001 , the only one matching all three filter conditions (tenant_id, category, and active): { "vectors": [ { "distance": 0. 9717477560043335, "key": "doc-001", "metadata": { "tenant_id": "t-10428", "category": "legal", "created_date": "2026-03-15", "active": true } } ], "distanceMetric": "cosine" } S3 Vectors first narrows the search space to vectors matching all three filter conditions, then returns the 50 most similar vectors from that subset.
Because the filter is applied before the search, those results are drawn from across all the vectors that match it.
Originally published at aws.amazon.com

