Search & AI / 011

Metadata Strategy for Vector Search

Search & AI · Search & Retrieval · Engineering notes

An architecture note on the metadata needed to filter, cite, refresh and permission-check indexed AEM content in a vector-search system.

Mohammed Boudoun · Published

AEM · Search · AI · Architecture exploration

In this engineering note

PROBLEM

Define a durable metadata contract for source paths, fragment identity, model, locale, version and access scope before indexing documents. This is an engineering notes brief, not a claim of production deployment or measured outcomes.

Metadata must travel with each retrieval unit; see Chunking AEM Content for RAG.

Apply it to Hybrid Retrieval with Elasticsearch.

AEM CONTENT SOURCE

Identify the Content Fragment models and fields that are authoritative for this use case. Record how localization, publication state, references and source revisions are represented; do not treat every authored field as suitable for retrieval.

INGESTION

Specify a repeatable extraction contract and trigger: initial backfill, publish or update, unpublish, deletion, retry and reconciliation. This outline does not claim a deployed connector or measured ingestion behavior.

CHUNKING

Compare whole-fragment retrieval with field-aware or semantic chunks. Preserve enough parent and section context for useful results, and make chunk boundaries deterministic so updates can replace prior records.

METADATA

Carry stable source identifiers, AEM paths, fragment model, locale, content version and applicable access scope alongside each chunk. Decide which values are filterable and which are safe to expose as citations.

EMBEDDINGS

Select an embedding model only after defining language coverage, data handling constraints, dimensions, refresh policy and evaluation criteria. Re-embedding is a versioned data migration, not an invisible model swap.

VECTOR STORE

Choose a store against workload, scale, filtering, hybrid query support, operational ownership, backup, residency and cost requirements. The named architecture is a question to evaluate, not a production result.

HYBRID RETRIEVAL

Test lexical, semantic and hybrid candidate generation on a representative query set. Define filters and ranking explicitly, then compare relevance and latency instead of assuming vector similarity alone is sufficient.

LLM

Keep generation behind a replaceable boundary. Select a model and hosting arrangement based on enterprise data policy, latency, context limits and the ability to return grounded answers.

GROUNDING

Return source-linked evidence with the answer and define behavior when retrieval is weak or conflicting. Citations should resolve to an allowed AEM source and make freshness inspectable.

PERMISSIONS

Carry identity and authorization decisions through retrieval. A content update or unpublish must not leave a stale index entry available, and filtering must not rely on the language model to enforce access.

EVALUATION

Build a labeled set of realistic questions and expected sources. Evaluate retrieval relevance, citation correctness, groundedness, access control, freshness, latency and cost before making quality claims.

FAILURE MODES

Plan for stale or deleted source content, duplicate chunks, incomplete backfills, embedding-version drift, irrelevant retrieval, permission leakage and model answers unsupported by retrieved evidence.

TRADEOFFS

Document the balance between authoring convenience, retrieval quality, freshness, explainability, operating cost and implementation complexity. Validate the tradeoffs in a bounded PoC before promoting a design.

References