A retrieval-augmented generation system can only answer from information that it successfully ingested, parsed, indexed and made available to the right users. The model may be excellent at explanation, but if the search index has lost a table heading or mixed two document versions into one chunk, generation starts from damaged evidence. That is why an AI-103 developer needs to understand the ingestion pipeline as a data product with schema, security, quality and freshness requirements.
Consider an internal safety manual that includes PDF diagrams, tables of machine limits and revised procedures. A weak pipeline converts each file to a single huge string, embeds it and calls the job done. A better pipeline preserves sections, page and version identifiers, equipment IDs, access classification and stable source pointers. Retrieval quality then has a chance to reflect the original documents rather than artifacts of a careless conversion.
Define the source system and freshness contract
Identify where documents originate, who is allowed to publish them and how revisions are represented. A file share, blob container, ticket store and wiki have different update and deletion behavior. If a document is withdrawn, the retrieval index must stop serving it within a defined window. A nightly full reload may be adequate for marketing copy and unacceptable for a revoked safety instruction.
Give each source a durable ID and version or update timestamp. Decide whether index documents are replaced, soft-deleted or retained as historical records. If multiple versions remain searchable, add a field that lets the application filter to the effective version. Logging the source version that supported an answer is essential for investigating why an agent gave a statement that was correct last month and wrong today.
Parse structure before choosing chunk sizes
Raw text extraction often destroys the clues that make a passage interpretable. A table may rely on column headings, a legal clause may depend on a preceding definition, and a diagram caption may describe an image that OCR never reads. Preserve logical paragraphs, headings, tables, document identifiers, and page boundaries where the source permits. Some content requires layout-aware processing rather than ordinary PDF-to-text conversion.
Chunking is a compromise between context and precision. Too-small chunks may detach a maintenance limit from its unit or scope. Too-large chunks waste context and bury the answer among unrelated sections. Experiment with section-aware boundaries and limited overlap, then use representative questions to test whether retrieved chunks contain enough evidence to answer correctly. A numeric chunk length is not a universal best practice; its suitability depends on document structure and the model’s retrieval task.
Design index fields for both content and controls
An Azure AI Search index may contain a searchable text field, vector fields, filterable attributes, and identifiers used to retrieve or cite the original source. Fields used for authorization and version selection must be filterable in the way the application intends. Do not depend on text similarity to enforce access control; a highly similar passage remains prohibited when its tenant or access group does not match the current caller.
Useful fields might include document_id, section_id, source_uri, effective_date, tenant_id, access_groups, language, content_type, normalized body text and one or more embeddings. Plan update behavior before selecting key fields: a change in a document’s title should not make its old content impossible to delete. Model your metadata and content separately enough that operators can tell whether a failed answer came from parsing, indexing, filtering or retrieval.
Choose embeddings with a clear vector contract
A text embedding encodes semantic information as a numeric vector that can be compared with query vectors. The index’s vector dimension and similarity configuration must match the embedding model and query strategy. Changing embedding families without rebuilding or migrating the corresponding vector fields can make similarity scores meaningless or queries fail.
Store which embedding configuration produced each indexed record. Decide whether the application supports one index version at a time or a staged rebuild with a switch-over. Validate domain terminology: abbreviations, part numbers and proper names may be better preserved by lexical search than semantic embeddings alone. A robust design does not pretend that every identifier becomes easier to find after vectorization. Microsoft’s Azure AI Search RAG overview explains the ingestion and retrieval roles of chunking, embeddings and indexing.
Use enrichment only where it improves evidence
A document may contain scanned text, a photograph, a table and a handwritten annotation. OCR can make text searchable; layout analysis can preserve relationships; language processing may produce entities or labels. Those steps have costs and can introduce errors. Apply them based on document type and downstream questions, not because every enrichment is available.
Suppose a part number is printed beneath a photograph. If the parser loses the label-to-image association, the index can return the right page while leaving the model unable to identify the part. Inspect parsed output with document owners and retain source coordinates or sections where feasible. For mixed media, Azure Content Understanding may provide structured extraction; its output still needs validation against the underlying source.
Protect ingestion from tenant leakage and malformed input
Authorization should apply at data preparation and query time. A multi-tenant ingestion job needs strict source routing and an index design that prevents one customer’s documents from becoming another customer’s search hits. Source file names can contain untrusted strings; document content can include prompt-injection text; metadata can be incomplete or misleading. Do not treat any of these inputs as configuration instructions for the ingestion job.
Validate types, required fields, file formats and tenant ownership before indexing. Quarantine malformed documents with a reason that an operator can investigate. Avoid silently indexing an object without its access classification, even if that means temporarily omitting it from search. Treat private source references carefully: a citation field must not become a public download link that bypasses the document’s normal authorization system.
Separate the indexer’s health from the index’s usefulness
An indexing process may report success because it processed all files, even though OCR confidence is low, chunks are empty or relevant records have lost their source identifiers. Collect quality signals: number of documents, chunks per document, parse failures, missing metadata, duplicate versions, index lag and searchability of known test passages. Include a small set of canary documents with predictable section content to verify the full path.
Run a retrieval test after significant ingestion changes. Query by an exact equipment code, a conceptual symptom, and a phrase that appears only in a recently updated procedure. Check that the returned result includes correct tenant scope, effective version, source URI and readable passage. These tests are more informative than simply observing that the index contains the expected number of rows.
Plan update and deletion workflows explicitly
Production knowledge changes. A manual is superseded, a contract is terminated, and a customer requests deletion of a record. Determine how those changes propagate through any blob mirror, enrichment cache, embeddings, search index and secondary backup or evaluation dataset. Retention and data deletion requirements may conflict with the convenience of preserving every historical chunk.
A safe reindex plan may build a new versioned index, validate it, then change an alias or application configuration once it passes. That reduces the risk of serving a partially rebuilt corpus. But it also requires a process to retire old indexes and remove unauthorized historical copies. Index aliases solve cutover mechanics; they do not solve data-governance policy by themselves.
Imagine an employee handbook that changes its remote-work eligibility section on Monday, while an internal assistant is still answering from Friday’s version. All services may show green health metrics: the crawler ran, the model responded and Azure AI Search returned results. The error is semantic freshness. Assign each indexed chunk a stable source document ID, source version, effective date and ingestion timestamp, then test that the retrieval layer selects the current authorized version rather than simply the most semantically similar text.
A practical reindexing design separates three operations. A new document should appear in search only after its extracted text and permissions are validated. An updated document should replace or version the previous chunks atomically enough that readers do not receive mutually contradictory policy versions. A deleted or access-revoked document must cease appearing in user queries even when old embeddings still exist in a search index. Do not rely on periodic rebuilds as the sole mechanism for emergency permission revocation; enforce current access metadata during query execution or through another authoritative security layer.
Build a test collection with one document whose content changes, one that is deleted and one whose permitted audience changes. Run identical questions as two users with different roles. Check not just the returned answer but the exact document IDs and policy versions sent into the model’s context. If an answer cites an out-of-date or unauthorized chunk, troubleshoot ingestion and authorization before changing the prompt or model. A more fluent answer generator cannot repair a stale or incorrectly permissioned evidence source.
Operational dashboards should therefore distinguish crawler completion, document parse failures, embedding errors, index-document counts, ACL sync lag and query relevance on a labeled benchmark set. Each metric answers a different question. Index availability is necessary, but correctness depends on the source-to-answer chain remaining current and within the user’s permission boundary.
Diagnose missing answers from ingestion evidence
When a model says it cannot find an instruction, do not immediately increase context length or change the generation prompt. Confirm that the source file was eligible, processed, parsed, chunked and uploaded. Search for a distinctive phrase in the raw index. If it exists, inspect filters, query formulation and ranking. If it does not, troubleshoot the document pipeline. This sequence keeps ingestion defects from being misclassified as hallucination or model incompetence.
For an AI-103 lab, deliberately corrupt a source metadata field, simulate a withdrawn document and change a table layout. Build tests that detect each problem and restore the correct state. An ingest pipeline is production ready when it delivers authorized, current and attributable evidence reliably—not simply when a demo query retrieves something that looks relevant.