Retrieval-augmented generation, or RAG, improves a generative AI application by retrieving relevant enterprise information and placing that information into the model’s context before generation. The architecture sounds simple, but production quality depends on a chain of decisions: which documents are allowed, how text is parsed and chunked, which embedding model is used, how retrieval is filtered and ranked, how prompts represent evidence, and how answers are evaluated for groundedness.
The Databricks Generative AI Engineer Associate certification remains focused on designing, building, deploying, and governing GenAI solutions, including RAG, AI Search, Model Serving, MLflow, and Unity Catalog. A useful RAG design should be judged as an information system, not only as a prompt.
Start with the knowledge boundary
Define which data sources the application is allowed to use and which users are allowed to retrieve from them. Enterprise knowledge bases often include public policies, internal procedures, customer records, source code, and regulated documents with very different access requirements. Mixing them into one unrestricted corpus creates a governance problem before the first embedding is generated.
Use Unity Catalog permissions and data ownership to establish the retrieval boundary. The index should not become a way to bypass the access model that protects the source data.
Prepare documents for retrieval, not just storage
Raw documents contain headers, navigation, repeated footers, tables, images, and boilerplate that can reduce retrieval quality. Normalize text while preserving structure that helps answer questions. Metadata such as source, owner, publication date, department, and security classification can be as important as the text because it enables filtering and traceability.
Chunking should follow semantic structure where possible. Very large chunks may retrieve too much irrelevant material, while tiny chunks can lose context. Test chunk size and overlap with representative questions rather than selecting values by habit.
Choose embeddings for the actual corpus
Embeddings place queries and documents into a semantic vector space so similar meanings can be retrieved even when exact words differ. The embedding model should fit the languages, technical vocabulary, and latency requirements of the application. A model that works well for general English may underperform on source code, legal text, or multilingual support content.
Changing the embedding model usually requires rebuilding the index because query and document vectors must be comparable. Treat embedding-model choice as a versioned system dependency rather than an invisible implementation detail.
Use AI Search with metadata filters
Databricks AI Search, formerly Vector Search, supports similarity-based retrieval for RAG and semantic search. Pure similarity is often insufficient in enterprise systems. Apply metadata filters for tenant, region, document status, product, or authorization boundary so semantically similar but inappropriate documents do not enter the prompt.
Filters can also improve quality by narrowing the candidate set before ranking. If the question concerns a specific product version, retrieving older documentation simply because it uses similar language may produce a confident but outdated answer.
Separate retrieval quality from generation quality
If the model receives the wrong context, changing the prompt may not fix the system. Evaluate whether relevant documents were retrieved before evaluating the final answer. Retrieval metrics and inspection help distinguish index problems from generation problems.
A useful debugging sequence is: did the correct document exist, was it indexed correctly, did the query retrieve it, did ranking place it high enough, did the prompt preserve it, and did the model use it faithfully? This prevents teams from treating every bad answer as an LLM problem.
Design prompts to expose evidence clearly
Retrieved content should be structured so the model can distinguish instructions from evidence and separate one source from another. Include source identifiers or citations when the application needs traceability. Instruct the model how to respond when evidence is insufficient instead of encouraging it to fill gaps from unsupported assumptions.
Prompt design matters, but the system should not depend on one enormous prompt containing every policy. Retrieval exists to supply the smallest useful context for the current request while keeping the application’s operating instructions stable.
Protect against unsafe retrieved content
Documents can contain instructions, malicious text, stale policies, or sensitive values, which is why broader AI security risks belong in the RAG threat model. Treat retrieved content as data, not trusted system instructions. Separate system-level behavior from retrieved passages and enforce access control before retrieval results reach the model.
The concerns in protecting PII in AI systems are especially relevant because RAG can expose information through natural-language answers even when users never see the original table or document.
Evaluate with representative question sets
Build evaluation datasets that include normal questions, ambiguous requests, missing-answer cases, adversarial wording, permission boundaries, and difficult retrieval examples. Measure relevance, groundedness, correctness, safety, and refusal behavior separately so one high aggregate score does not hide a serious weakness.
Use the same evaluation set when comparing chunking strategies, embedding models, prompts, or ranking changes. This turns RAG tuning into controlled experimentation rather than subjective demo review.
Operate RAG as a living data product
Documents change, permissions change, embeddings age, and user questions reveal new failure modes. Monitor retrieval traces, answer quality, latency, empty-result rates, and feedback in production. Re-index when source content or embedding strategy changes, and remove deprecated information from retrieval paths rather than leaving it available indefinitely.
The wider Databricks certification ecosystem connects this AI architecture to the same governance and data-engineering foundations used elsewhere on the platform. A strong RAG system is not a chatbot glued to a vector index; it is a governed retrieval pipeline with measurable evidence quality, explicit access boundaries, and a repeatable evaluation loop.
Document freshness should be part of retrieval design. An index that contains both current and superseded policies can return semantically relevant but operationally wrong context. Store effective dates and status metadata, then filter or rank so current authoritative content wins unless the user explicitly asks for historical information.
Chunk identifiers should remain stable enough to support debugging and re-indexing. When chunking logic changes, record the version so evaluators can tell whether an answer difference came from the source document, the chunker, the embedding model, or the retriever. Without versioned preprocessing, retrieval experiments are difficult to reproduce.
Hybrid retrieval can combine semantic similarity with keyword or metadata signals. Exact identifiers, product codes, legal clauses, and error messages are often better served by lexical matching than pure embeddings. Evaluate retrieval methods against the corpus instead of assuming one approach is best for every question.
Reranking can improve precision when the first-stage retriever returns a broad candidate set. The trade-off is added latency and cost. Measure whether reranking materially improves hard questions and use it where the quality gain justifies the extra step.
Context windows are finite even when modern models accept large prompts. More retrieved text is not automatically better. Irrelevant context can distract the model, increase latency, and make citations less meaningful. Select the smallest evidence set that covers the question and keep chunk provenance intact.
Answer citations should point to sources users can actually access. If the model cites a restricted document that the user cannot open, the application creates confusion and may reveal metadata about protected content. Authorization should therefore shape both retrieval and the evidence shown in the response.
Queries themselves can contain sensitive data. Logs and traces may capture employee names, account numbers, or confidential business questions. Apply retention and redaction controls to observability data so improving the RAG system does not create a less-governed copy of user information.
Evaluation should include “no answer” cases. A trustworthy system needs to recognize when the corpus lacks sufficient evidence rather than generating plausible text from the model’s general knowledge. Measure refusal or uncertainty behavior explicitly because it is a core part of groundedness.
User feedback is most valuable when it is structured. A simple negative rating does not reveal whether retrieval was wrong, the answer ignored evidence, the source was outdated, or the wording was unclear. Capture reason codes and feed recurring failure types back into evaluation datasets.
Latency budgets should be broken down by component. Embedding the query, retrieving candidates, reranking, generating the answer, and post-processing citations each consume time. Optimize the slow component instead of making broad changes that may hurt quality elsewhere.
Cost should be evaluated per successful answer, not only per model call. Larger contexts, reranking models, frequent index refreshes, and expensive generation models can all increase spend. A cheaper system that returns weak evidence may cost more operationally if users must repeat queries or escalate to humans.
RAG quality is ultimately a property of the whole chain. Reliable systems govern the knowledge base, preserve metadata, retrieve deliberately, evaluate components separately, trace production behavior, and update indexes as the organization changes. The LLM is one participant in that architecture rather than the sole source of intelligence.
Index refresh strategy depends on source change frequency and answer freshness requirements. Some knowledge bases can rebuild nightly; others need near-real-time synchronization. Choose a refresh model that exposes new and deleted content quickly enough for the business risk, and monitor index lag separately from source freshness.
Deletion propagation is especially important. If a document is removed for legal, privacy, or correctness reasons, ensure the corresponding chunks and embeddings are also removed from retrieval. A deleted source that remains searchable through an index is still effectively available.
Embedding pipelines should preserve language and document metadata. Multilingual corpora may need language-aware chunking or embeddings, and users may expect results in the language of the question. Evaluate language pairs explicitly rather than assuming performance transfers from English tests.
Structured data can complement document retrieval. Some questions are better answered by SQL or a governed tool call than by embedding rows into a vector index. Architect the application so it can choose the right evidence path instead of forcing every source into RAG.
RAG applications should distinguish factual retrieval from workflow execution. Retrieving a policy that says a refund is allowed is not the same as authorizing and issuing a refund. Keep action-taking tools behind separate permissions and confirmation logic.
Evaluation datasets should include permission-sensitive pairs: two users asking the same question but entitled to different information. This tests whether identity filters are enforced before generation rather than merely checking answer quality for one privileged account.
Use source authority as a ranking signal when the corpus contains duplicated information. Official policy may need to outrank an old wiki or discussion thread even when all are semantically similar. Metadata can encode trust level so retrieval favors the intended source of truth.
Observability should capture retrieval candidates, final selected context, model response, and latency without exposing more sensitive text than necessary. This level of traceability makes it possible to reproduce poor answers while still respecting privacy controls.
When users ask follow-up questions, conversation history becomes another source of context. Limit how much prior dialogue is carried forward and separate user statements from authoritative retrieved evidence. A mistaken claim earlier in the conversation should not silently become trusted knowledge.
A production RAG system also needs capacity planning. Index size, query volume, embedding throughput, model concurrency, and refresh jobs can all become bottlenecks. Test peak patterns rather than only interactive demos so retrieval quality remains stable under real load.
Evaluation should also test citation fidelity. If the answer claims that one source supports a statement, verify that the cited chunk actually contains the evidence. A system that retrieves good documents but attaches citations to the wrong claim can create false confidence.
Knowledge-base ownership matters after launch. Each source should have someone responsible for freshness and retirement. When no owner exists, outdated material tends to remain indexed indefinitely because the AI team cannot determine whether it is still authoritative.
Use controlled experiments for chunking and retrieval changes. Compare candidate configurations on the same evaluation dataset and inspect both aggregate results and hard examples. This prevents teams from adopting a new embedding or chunk size based on a handful of impressive demonstrations.
RAG reliability ultimately depends on maintaining the evidence pipeline with the same discipline as any other production data product. Sources, transformations, indexes, permissions, evaluations, and monitoring all need owners and versioned change processes.
Retrieval testing should include documents that are semantically similar but intentionally wrong for the question, such as an outdated policy or another product version. These adversarial near-matches reveal whether metadata filters and ranking are strong enough to select authoritative evidence.
When RAG supports high-impact decisions, require users to see enough source context to verify the answer themselves. The application should accelerate access to evidence, not hide the evidence behind an unreviewable model response.
Version the retrieval stack as one releaseable unit where practical: preprocessing logic, embedding choice, index configuration, ranking, prompt template, and model. A quality regression can then be traced to a defined release instead of a mixture of independently changed components whose combined behavior is impossible to reproduce.
Do not treat evaluation as a pre-launch event. New documents, changed permissions, model updates, and evolving user behavior can all change answer quality after deployment. Continuous sampling and regression testing keep the retrieval system aligned with the knowledge base it is supposed to represent.
Production readiness also includes fallback behavior. If retrieval is unavailable or returns no trustworthy evidence, the application should fail safely, explain the limitation, and avoid converting an infrastructure problem into an unsupported answer.
Safe fallback behavior is part of RAG quality, not an optional operational feature added after launch.