Search & AI / 010

Chunking AEM Content for RAG

Search & AI · Search & Retrieval · Exploration

A practical exploration of splitting structured AEM content for retrieval while preserving component context, citations and update boundaries.

Mohammed Boudoun · Published

AEM · Search · AI · Architecture exploration

In this engineering note

PROBLEM

Test how fragment fields, headings and semantic boundaries can become retrieval units without losing the context needed to cite or update the source. This is an exploration brief, not a claim of production deployment or measured outcomes.

Chunking choices depend on the source contract in GraphQL vs Sling Model Exporter.

See the end-to-end flow in Building a RAG Pipeline on Top of AEM Content.

AEM CONTENT SOURCE

Identify the Content Fragment models and fields that are authoritative for this use case. Record how localization, publication state, references and source revisions are represented; do not treat every authored field as suitable for retrieval.

INGESTION

Specify a repeatable extraction contract and trigger: initial backfill, publish or update, unpublish, deletion, retry and reconciliation. This outline does not claim a deployed connector or measured ingestion behavior.

CHUNKING

Compare whole-fragment retrieval with field-aware or semantic chunks. Preserve enough parent and section context for useful results, and make chunk boundaries deterministic so updates can replace prior records.

METADATA

Carry stable source identifiers, AEM paths, fragment model, locale, content version and applicable access scope alongside each chunk. Decide which values are filterable and which are safe to expose as citations.

EMBEDDINGS

Select an embedding model only after defining language coverage, data handling constraints, dimensions, refresh policy and evaluation criteria. Re-embedding is a versioned data migration, not an invisible model swap.

VECTOR STORE

Choose a store against workload, scale, filtering, hybrid query support, operational ownership, backup, residency and cost requirements. The named architecture is a question to evaluate, not a production result.

HYBRID RETRIEVAL

Test lexical, semantic and hybrid candidate generation on a representative query set. Define filters and ranking explicitly, then compare relevance and latency instead of assuming vector similarity alone is sufficient.

LLM

Keep generation behind a replaceable boundary. Select a model and hosting arrangement based on enterprise data policy, latency, context limits and the ability to return grounded answers.

GROUNDING

Return source-linked evidence with the answer and define behavior when retrieval is weak or conflicting. Citations should resolve to an allowed AEM source and make freshness inspectable.

PERMISSIONS

Carry identity and authorization decisions through retrieval. A content update or unpublish must not leave a stale index entry available, and filtering must not rely on the language model to enforce access.

EVALUATION

Build a labeled set of realistic questions and expected sources. Evaluate retrieval relevance, citation correctness, groundedness, access control, freshness, latency and cost before making quality claims.

FAILURE MODES

Plan for stale or deleted source content, duplicate chunks, incomplete backfills, embedding-version drift, irrelevant retrieval, permission leakage and model answers unsupported by retrieved evidence.

TRADEOFFS

Document the balance between authoring convenience, retrieval quality, freshness, explainability, operating cost and implementation complexity. Validate the tradeoffs in a bounded PoC before promoting a design.

References