Structured AEM content offers fields and provenance that can help build retrievable passages. This proposed PoC studies whether that structure supports useful, source-backed answers; it does not assume an LLM will preserve the source's meaning or access restrictions.
AEM + RAG Services on Kubernetes.
PoC design / architecture exploration — not a production deployment
Explore a proof-of-concept architecture that connects governed AEM Content Fragments to retrieval and an enterprise assistant. Kubernetes supplies workload boundaries and rollout controls; the PoC must still establish permission safety, content freshness and answer quality.
Representative engineering design, not a verified client delivery record. Implementation steps and validation checks describe the proposed approach; no measured results are claimed.
- Content Fragments
- GraphQL / API
- Ingestion
- Retrieval index
- RAG service
- LLM endpoint
- Assistant
Conceptual dependency flow. AEM and the LLM endpoint remain external; Kubernetes orchestrates ingestion, retrieval and RAG service workloads.
Use governed content as a source, not a guarantee
Orchestration does not establish trustworthy retrieval
A technically healthy RAG service can retrieve stale content, omit the relevant passage or disclose information the caller cannot access. Those are application and governance failures that Pod health checks cannot detect.
Separate ingestion from interactive answering
An ingestion workload fetches approved content through GraphQL or APIs, transforms it into versioned passages and updates a retrieval index. Interactive services authenticate callers, retrieve permitted evidence, build a bounded prompt and call an external LLM endpoint before returning an answer with source references.
Carry identity and provenance through the pipeline
Attach source URL, content identifier, revision and access metadata to passages. Enforce authorization before content reaches the model, propagate deletions and permission changes, and treat retrieved text as untrusted data rather than executable instructions.
- Use stable identifiers and resumable ingestion so retries do not duplicate passages.
- Separate service identities and network access for ingestion, retrieval and model calls.
- Version chunking, embedding inputs, prompt templates and model configuration for reproducible evaluation.
Evaluate behavior before promoting a service revision
Run a fixed question set against the candidate retrieval and prompt configuration, then inspect citations, abstentions and permission failures. Deploy compatible service and index revisions together, with a rollback plan that accounts for their shared data contract.
- Approved content
- Ingest
- Index
- Evaluate
- Deploy candidate
- Review answers
Evaluation is proposed work. No answer-quality benchmark or production result is asserted.
Scale ingestion and answering independently
Batch ingestion and interactive requests have different latency and concurrency needs. Bound model-call budgets and retries, record retrieval timing and source revisions, and avoid storing sensitive prompts or answers in general-purpose logs.
Freshness, isolation and platform overhead
A service boundary permits independent scaling and releases but introduces more failure points and operational cost. Before choosing Kubernetes, compare a simpler managed execution model; a small PoC may not justify a cluster operating model.
Define evidence before considering production use
The PoC's output should be an evaluation record and a decision about suitability, not a claim that deployment establishes trustworthy AI. Production adoption would require explicit security, content and operational acceptance.
- Ask the same question as differently authorized users and inspect retrieved passages as well as final answers.
- Delete or restrict an AEM fragment and verify it no longer reaches prompts after the defined propagation window.
- Test unsupported questions, conflicting sources and injected instructions inside retrieved content.
- Simulate LLM timeouts, ingestion interruption and incompatible index revisions; verify bounded failure and recovery.
Treat answer quality as a release concern
Orchestration supports repeatability, but the decisive artifact is the evaluation evidence. A credible AI architecture makes permission handling, provenance and failure behavior as reviewable as its deployment manifests.
Technical references
These sources document product behavior. The design and validation approach above are engineering proposals, not claims made by the vendors.
Building or modernizing an AEM platform?
This representative case study explores technical trade-offs and architectural decisions for a specific engineering scenario. If you are planning a similar migration, modernization, or integration, let's discuss the engineering approach.
Start a conversation