INSIGHTS
AI & Data

Hybrid Retrieval and Grounded Answers in AI-103

Implement authorized AI-103 RAG using keyword, vector, hybrid search, semantic ranking, source provenance, and refusal when evidence is missing.

In this article
  1. Begin with a search question and an authorization boundary
  2. Understand lexical, vector and hybrid matching
  3. Use semantic reranking as a second-stage decision
  4. Keep retrieval metadata alongside the chunk
  5. Build prompts that distinguish source facts from instructions
  6. Diagnose retrieval failures separately from generation failures
  7. Evaluate source quality before scoring answer quality
  8. Measure the full latency and cost path
  9. Implement an explicit no-answer path

A user asks, “What is the latest isolation procedure for machine X-41?” The knowledge index contains a revised procedure, an older checklist and a training slide with similar vocabulary. A generative model can produce a polished response from any of them. The application is only reliable when retrieval selects current, authorized evidence and generation stays within what that evidence supports. That is the distinction between a search-enabled chatbot and a defensible retrieval-augmented generation system.

AI-103 places RAG inside its generative and information-extraction domains. Engineers need to design a query path, not just an ingestion job. They should understand why full-text, vector and hybrid search return different results, how reranking changes ordering, and how citations and missing-evidence behavior affect an application’s safety.

Begin with a search question and an authorization boundary

Define the source scope before sending a query. If the application serves employees in different regions, some manuals may be visible only to a subset. Tenant identity, region and document classification should constrain retrieval through enforceable filters. Never rely on a system prompt asking a model not to quote restricted documents after the application has already supplied them.

Describe what the user actually needs. An exact equipment code benefits from lexical matching; a question about “intermittent coolant leakage after maintenance” may require semantic similarity because no document uses those words exactly. If the query contains both an identifier and a symptom, losing either component can change the correct result. Protect exact terms while allowing paraphrase-aware retrieval for the descriptive part.

Suppose a support assistant serves two enterprise customers whose documents mention the same product code. A query for a device fault returns the highest-scoring passage—but it belongs to the other customer’s private support case. Relevance scoring has done its job while authorization has failed. Each indexed chunk needs enough tenant and document-access metadata for the backend to construct a security-filtered retrieval query. The model should never receive a forbidden chunk and then be trusted not to quote it.

Test the data boundary with paired fixtures: both tenants use the same phrase, only one tenant’s documents contain a specific serial number, and a privileged operator has additional service-wide read rights. A test must prove that ordinary users cannot see the other tenant’s chunk even when it ranks first in an unfiltered benchmark. It should also prove that the operator sees only what the intended business operation permits. A single shared service key used by the retrieval layer must not silently become a substitute for end-user authorization.

Hybrid retrieval complicates the quality decision usefully. A precise serial number often benefits from lexical matching; a paraphrased complaint may rely on vector similarity. Semantic ranking can reorder a candidate set, but it cannot recover a source excluded by a filter or absent from the indexed corpus. Measure expected documents per query, query type and audience; then run generation on that exact retrieved context and evaluate whether the final answer correctly expresses uncertainty.

When no authorized evidence exists, the assistant should say the documents available to this user do not establish the answer. It may offer a permitted escalation path, but it must not fill the gap with a generic manual from another customer. This behavior protects confidentiality and makes the product’s evidence boundary understandable to the end user.

Understand lexical, vector and hybrid matching

Full-text retrieval can match words, phrases and structured fields using lexical ranking. Vector retrieval compares embeddings representing conceptual similarity. Hybrid search evaluates lexical and vector queries in one request and merges their ranked results. This helps recover both exact codes and paraphrased descriptions, but it does not guarantee the top document is authoritative or current.

Microsoft’s hybrid-search overview describes how Azure AI Search combines full-text and vector results using reciprocal rank fusion. An RRF score is a merged-rank signal; it is not an absolute probability that the answer is correct. Compare returned passages and sources rather than imposing an arbitrary “confidence greater than 0.9” rule on a fusion score that was never calibrated as confidence.

Use semantic reranking as a second-stage decision

A query may return a broad candidate set from keyword and vector paths. Semantic ranking then evaluates the textual relevance of candidates and can reorder them. This is especially useful when a document has the right terms but the surrounding paragraph is about a different workflow. The index must provide text fields that let the reranker make that comparison; it cannot rescue missing or unauthorized source records.

Plan the candidate set thoughtfully. A tight first stage may exclude the correct passage before semantic ranker sees it; an overly broad set adds latency and cost. Build a small labeled evaluation set with expected relevant documents and compare recall and ranking quality under each configuration. Microsoft’s semantic ranking guidance explains the relationship between initial ranking and the reranking stage.

Keep retrieval metadata alongside the chunk

A passage needs a source identity, revision, section or page reference, and authorization context. An answer that cites only a generated chunk ID is difficult for a user to verify. Store enough mapping to open the original document through the same access-controlled system the application uses. A citation should lead to the precise relevant passage or an accessible document with a meaningful location marker.

Timestamp handling matters. A search result from a historical runbook can be genuinely relevant and dangerously obsolete. Use effective-date and status filters when the task asks for current instructions. If the app is asked to compare versions, retrieve both deliberately and label them. Do not let the model treat version history as one undifferentiated pile of evidence.

Build prompts that distinguish source facts from instructions

When retrieved material reaches the model, identify it as quoted evidence. Ask the model to answer only within the approved task scope and cite supported claims. A document may contain a sentence that resembles a system instruction, but the retrieval layer does not have the authority to redefine the agent’s role or tool permissions.

Prompt wording helps but cannot enforce a security boundary. The service that retrieves documents should already have filtered access; the tool that performs actions should validate its arguments independently. For unsupported questions, the correct output may be “the supplied sources do not establish that fact.” Requiring a citation does not solve unsupported answers if the model fabricates a plausible-looking reference, so citations also need deterministic validation against the actual result set.

Diagnose retrieval failures separately from generation failures

Suppose a user reports a wrong warranty date. First inspect which documents retrieval returned. If the current contract was absent, troubleshoot ingestion, query terms, filters, embeddings or reranking. If the correct contract was present and the model selected the wrong line, investigate generation prompt, context ordering or answer validation. If the result cited another customer’s contract, the highest-priority fault is authorization, not an ordinary relevance defect.

This separation leads to better experiments. Change one retrieval setting and measure relevant-document recall, then hold retrieval constant while comparing generation configurations. Combining a model upgrade with reindexing and a new prompt makes it impossible to know why a change helped or regressed. Ingestion and query-time access filtering therefore need their own tests: a perfectly grounded answer can still be unsafe if its evidence came from a document the user was not allowed to retrieve.

Evaluate source quality before scoring answer quality

A RAG evaluation set needs more than question-and-answer pairs. Record which documents and passages are relevant, which are prohibited, and which are intentionally outdated. This permits retrieval-level measures such as relevant-document ranking and missing-evidence rates. At the response level, compare factual correctness, groundedness, completeness and the faithfulness of citations to retrieved passages.

Microsoft provides RAG evaluators for retrieval and grounded response quality. Their metrics should be interpreted in context: a high response relevance score cannot excuse a data leak, and a high retrieval relevance score cannot prove that the generated answer accurately cited a value. Include adversarial and “no available evidence” cases so the system learns not to fill every gap with fluent text.

Measure the full latency and cost path

RAG latency includes query reformulation, embedding, search, optional reranking and model generation. Additional tool calls and retries can dominate an agent workflow. Measure those stages separately so tuning does not focus on generation when index retrieval or network round trips are responsible for the delay.

Not every query deserves the same number of retrieval steps. A precise part number may be resolved with a restricted lexical query, while a complex cross-document comparison benefits from a richer pipeline. Set limits on retrieval fan-out and the size of evidence passed to the model. Balance costs by successful, grounded task completion rather than by the number of tokens alone.

Implement an explicit no-answer path

Imagine a query asks whether an unpublished regulatory draft is now mandatory. The index contains only an older consultation note. The system should not claim certainty just because the retrieved note mentions the same subject. It should explain that the approved sources do not establish the current mandate, identify the outdated source and avoid converting an inference into official advice.

An AI-103 lab can make this visible: index a prior and current policy, remove the current record, and compare what the assistant does under an unsupported request. Then restore it and test exact-source attribution. A dependable RAG system succeeds not only by answering supported questions but also by recognizing when its evidence is missing, conflicting or outside the user’s authorization boundary.

Filed under AI & Data