Scanning a document and extracting its text are not the same as understanding the record. A purchase order may contain a supplier address, a handwritten delivery note, a table of quantities and a printed total. OCR can recover many characters, but an application must still preserve layout, identify which number belongs to which field and validate the relationship between rows and totals. An AI assistant that confidently reports an incorrect invoice amount can cause far more damage than one that admits the scan is unreadable.
Microsoft AI-103 includes document extraction using OCR, layout analysis, structured output and Content Understanding. A practical developer should know when ordinary OCR is enough, when a specialized analyzer is justified, and where human verification belongs. The right design starts with the information needed by the downstream business process, not with a generic promise to turn every PDF into JSON.
Distinguish characters, layout and field meaning
OCR converts visible text in a scan into machine-readable characters. Layout analysis identifies regions such as pages, paragraphs, tables, headings and key-value relationships. Field extraction maps those elements into a business schema, such as invoice_date, supplier_id, line_items and amount_due. These operations build on one another but fail in different ways.
A scanned invoice can have accurate character recognition yet a wrong field mapping if it contains both a purchase-order number and an invoice number. A table might preserve all digits while swapping quantity and unit-price columns. Review each layer independently. A model’s polished response cannot compensate for a missing unit or row association in the source representation.
Select analyzers by the document problem
Prebuilt analyzers can be useful for common document forms when their supported schema and performance fit the task. Custom analyzers help when an organization has unusual document layouts or domain-specific fields. A multimodal LLM may assist with messy, unstructured documents that a rigid form parser struggles with, but it can also hallucinate a plausible value from surrounding context. The application’s risk tolerance should determine whether generative extraction is acceptable.
Microsoft’s Content Understanding documentation describes document analyzers that combine text, layout and structured extraction. The analyzer reference distinguishes general, domain-specific and custom analyzers. Confirm whether a chosen feature is generally available or preview in the API version being deployed; the answer can change over time.
Design the output schema before prompts and forms
Business systems need stable field names and data types. Define which fields are required, which may be null, which values repeat, and how monetary units or dates are represented. For line items, include item code, description, quantity, unit price, tax and line total if relevant. Keep a pointer back to the page or region that supplied each value whenever the extraction service supports it.
An output schema is a contract, not proof. A JSON object may pass syntax and type checks while holding the wrong supplier ID. Validate checksums or identifier formats, compare extracted suppliers with master data, and recalculate totals. If two documents disagree, return a conflict state rather than silently selecting the value the model finds most persuasive.
Prepare scanned inputs responsibly
Documents arrive as searchable PDFs, raster scans, photos and office files with embedded tables or images. Each requires different parsing. A photo may be skewed or shadowed; a fax may have missing characters; a digital PDF may already contain structured text that should not be degraded by unnecessary OCR. Choose preprocessing based on the input rather than applying every enhancement to every file.
Retain a secure source identifier, document version and processing timestamp. For noisy scans, route uncertain regions to a review queue with the original crop or page visible to the reviewer. If the model cannot distinguish a six from an eight in an amount, it should not invent the difference from common invoice patterns. Source fidelity is more important than completing every field.
Preserve tables and cross-field relationships
Tables are difficult because a single row may span pages, merged cells can shift columns, and footnotes can change how totals are interpreted. Do not flatten tables into an unordered sequence of words before creating the business record. Keep row grouping and column headings where possible. The application should be able to show which printed row became which structured line item.
Cross-field consistency checks often catch errors that individual confidence scores miss. If the unit prices and quantities do not support the extracted total, flag it. If the invoice date is after the due date and the business process expects otherwise, review the record. Some anomalies are legitimate, so validation should not rewrite values automatically; it should identify cases that need evidence or a person to resolve.
An invoice may contain two rows with the same part number but different quantities and discounts. A naive OCR pipeline extracts the words correctly and then flattens the table into one paragraph. The field extractor later associates a line-item price with the wrong row and produces a mathematically plausible total. The failure happened at layout reconstruction, not character recognition. Preserve row and column structure, page boundaries and links between a field and its source region wherever the analyzer provides them.
Build a verification step that recalculates totals from extracted line items, compares tax and currency consistently, and rejects contradictory values. A confidence score can prioritize human review, but it is not proof that a payment amount is correct. When the model returns one total and the arithmetic disagrees, retain both observed values and mark the document as unresolved. Never silently make a choice because one answer looks more fluent in JSON.
Document inputs also need a controlled lifecycle. A page photographed at an angle may require orientation correction; a low-resolution screenshot may not contain enough pixels to read small digits. Compression can irreversibly remove the evidence, and aggressive preprocessing may distort characters. Test real acquisition conditions before committing to one normalization path, and retain a permitted source reference for auditors to revisit.
Measure accuracy by field and document class, not only by the proportion of jobs returning some output. Invoice totals, bank details and tax identifiers deserve stricter thresholds than a noncritical memo title. Track manual-review time and the kinds of errors reviewers catch. If expanding the model’s ability to guess increases throughput but also misroutes payments, the system has optimized the wrong outcome.
Add classification before extraction where documents vary
A mixed intake queue may include invoices, credit notes, receipts, contracts and unrelated correspondence. Applying an invoice schema to every file can produce nonsensical but plausible values. First classify or route the document type with an explicit unknown category. Then apply the analyzer and validation rules that correspond to the classified type.
Classification should be testable and reversible. If a credit note is misclassified as an invoice, the downstream workflow must not issue a payment. Retain type evidence and allow a reviewer to correct the route. For bundles, splitting at the right page boundary may be as important as field extraction; a multi-document PDF should not combine a supplier from one form with a total from another.
Protect documents and extracted records
Invoices and identity documents can contain bank details, addresses, personal numbers or confidential commercial terms. Apply storage authorization, retention, encryption and tenant separation to the source and every derived artifact. An extraction system that produces structured JSON may make sensitive data easier to search and leak; that increases the importance of least privilege, not reduces it.
Do not put full source text or private documents into debugging traces by default. Use safe source IDs and field-level error codes, and control who can view document previews during review. If a business process allows the extractor to create a draft record, separate that permission from authority to approve payment or release customer data. Extraction is a data interpretation capability, not an authorization grant.
Measure field accuracy and review workload together
Track exact match for identifiers, normalized date and currency correctness, table-row fidelity, missing-value behavior and error categories. A system that extracts 97 percent of simple fields but fails on handwritten total amounts may still create an unacceptable review burden. Evaluate by document type, source channel and input quality so performance is not hidden by a blended average.
Set a review threshold based on business risk and observed error rates rather than treating an uncalibrated model confidence score as a guarantee. A low-value, reversible classification may allow a lower threshold; a payment instruction requires stronger evidence and approval. Count the number of records sent to review and how often reviewers correct the model. Those corrections should inform evaluation and preprocessing improvements.
Build a controlled invoice intake lab
For AI-103, assemble a small set of approved synthetic invoices with clean, rotated, handwritten and multi-page variants. Extract header fields and line items, validate arithmetic, link every value to its source, and require review for conflicting totals. Add a credit note and a non-invoice to test classification rather than forcing extraction regardless of document type.
Finally, simulate a withdrawn document and confirm that its extracted record can be traced, corrected or removed according to policy. A dependable document-understanding workflow provides structured information with evidence, known uncertainty and safe downstream boundaries. It does not merely produce JSON that looks convincing.