A hospital receives several thousand reimbursement forms as PDFs and needs a structured record of claim identifiers, dates, providers and amounts. Each form is different enough that rigid text matching fails, yet the business requirements are precise: fields must be traceable to a source document, invalid values must be rejected, and a missing field cannot quietly become an invented answer. This is an extraction architecture problem, not a contest to write the most elaborate prompt.
Claude Certified Architect – Foundations—official exam code CCAR-F and site shorthand CCA-F—covers prompt design, structured output, validation and batch processing. These topics connect, but none is interchangeable with the others. A schema can produce valid JSON containing the wrong amount. A batch can make large jobs economical while hiding individual failures unless the application keeps a reliable manifest. A second model pass can catch inconsistent facts but still needs independent checks.
Choose batch processing when the workload can wait
For a customer on a live support call, the correct response is usually constrained by interactive latency. For a nightly archive of forms, individual documents can often be processed independently and the system can wait for an asynchronous completion window. The Message Batches API supports groups of independent message requests with per-request identifiers and a separate results-retrieval step. Its published processing window can extend to 24 hours, so an application must not promise immediate completion merely because batch submission succeeded.
A batch is not a long-running agent conversation. It does not automatically continue a multi-turn sequence of client-side tool calls across many messages. Each request should contain the information and parameters needed to perform a bounded task. If extraction requires repeatedly searching an external case-management system, the application should arrange that retrieval separately or choose a different workflow for those documents.
A useful rule is to batch stable, self-contained work and keep interactive or state-changing decisions behind the appropriate service. Bulk classification, field extraction and first-pass comparison can be strong candidates. Issuing refunds, modifying medical records or escalating an ambiguous claim should not be smuggled into a bulk inference result without authorization and additional controls.
Architects should also measure the actual shape of the workload. A small batch of highly heterogeneous documents may not benefit as much as a predictable nightly volume. If the business must return a result in seconds, adding an asynchronous queue solely to access a batch API makes the service worse, even if the unit inference price is lower.
Design a manifest before sending the first request
The most common operational mistake is to treat a successful submission receipt as evidence that every document was processed. Instead, assign stable work IDs and persist an input manifest that records source document, revision or hash, expected schema version, destination record and the request’s custom_id or equivalent correlation identifier.
Batch results can arrive out of the original request order. The application should join them to the manifest by ID, not by position in a returned file. A partial completion record must explain which inputs succeeded, failed, expired or need follow-up. If the process restarts, it should load the manifest and avoid resubmitting completed work just because the original client lost its in-memory list.
Imagine documents A, B and C are submitted together. A succeeds, B reports invalid input, and C expires without a result. A weak coordinator retries the whole batch and risks duplicate downstream records. A stronger coordinator records A as accepted pending semantic validation, sends B for input repair, and resubmits only C under a new attempt ID while preserving the logical work ID. The difference becomes critical when a later stage writes to a database.
A manifest also supports legal and operational audit. If a reviewer disputes a value, the team should be able to identify the source revision and model-request version that produced it. The original document can remain under the proper access policy while the manifest stores only the metadata necessary for reconciliation.
Make schemas express uncertainty, not hide it
A form that sometimes omits a provider address should not require the model to invent one. A useful JSON schema distinguishes required output keys from required real-world values. The output object may always contain an address field while allowing the value to be null, together with an evidence reference or extraction status. The schema should use constrained types and enumerations where they genuinely help, without making the data model so brittle that ordinary exceptions are impossible to represent.
For example, an invoice extraction record might include invoice_id, currency, total_amount, evidence_span and needs_review. A currency code can be an allowed enumeration if the business supports a known set. A missing invoice ID should not become 0000 simply to satisfy a string constraint. If the document contains two totals, the model can report the ambiguity and the application can flag it for review.
There is a difference between output-format constraints and the content of the extracted values. A structured-output mechanism can guide Claude toward a valid JSON representation. A tool-use interface can define arguments to be passed to a downstream function. Neither validates that the reimbursement amount is supported by the source document. Syntax and factual correctness therefore require separate validation.
Schema versions should be explicit. If a team adds a new optional category next month, old results should not be silently interpreted as though they were produced under the new rules. The downstream pipeline can handle supported versions or migrate them through a tested transformation.
Extract evidence before making business decisions
An extraction model should return source-grounded observations before the application decides whether to approve a claim. Separating those steps prevents the agent from blending what the document says with a policy judgment that belongs to another service or reviewer.
Consider a handwritten reimbursement total that looks like either 1,080 or 1,030. A useful extraction does not merely select one number confidently; it identifies the uncertain location or evidence span and marks the value for verification. The application can compare the alleged total with the line-item sum and other structured fields. Where source evidence cannot be resolved, the correct outcome may be a targeted human review rather than a guessed value.
Long documents create a context problem. Splitting them into chunks can help with size, but it can separate a table from its header or a value from the section that qualifies it. Preserve document structure, page references, table relationships and revision metadata. Retrieval should supply the evidence that actually supports the requested fields, not simply the top-ranked text chunk.
The system may need to combine facts from multiple pages, such as a policy start date on page one and a signature on page six. In that case, the extraction output should identify both sources. If the pages conflict, the result should explain the conflict without collapsing it into a single unsupported claim. A clean JSON object is not useful if it makes ambiguity invisible.
Validate in layers and retry only the layer that failed
A robust processing pipeline has several checks. Structural validation asks whether the output parses and matches the expected schema. Semantic validation asks whether numeric ranges, date ordering, totals and business rules make sense. Source verification asks whether key values appear in the document or can be corroborated by the permitted evidence. Authorization determines whether the consumer may view or change the resulting record.
Different failures call for different repairs. Malformed JSON may justify a bounded regeneration attempt or a different structured-output setting. A mismatched date range may require the model to revisit a cited span. A missing source file cannot be solved by rephrasing the prompt; it requires retrieving the input or declaring the request blocked. An unauthorized destination record must remain inaccessible even if the extraction itself is excellent.
A retry should preserve the original logical work item, increment its attempt number and record why the retry happened. Do not repeat an entire successful batch because one document failed a semantic check. For state-changing downstream operations, use idempotency or reconciliation where supported, so a network timeout cannot silently create a duplicate record.
A human-review queue needs specific reasons. “Low confidence” without evidence is not actionable. A better review item states “Two totals detected on pages three and four; line-item sum disagrees by fifty units; source page references attached.” The reviewer can resolve the ambiguity without repeating the whole extraction process.
A reconciliation record for a failed batch
Imagine 700 intake documents submitted in a single logical job. At the end, 693 parse successfully, four have unreadable attachments, two pass syntax checks but contain unsupported figures, and one job is still running. The pipeline must not turn “693 valid outputs” into “700 processed.” A durable manifest should record an immutable document key, submission attempt, provider request identifier if available, completion state, input-version hash, schema-validation outcome and acceptance decision. The human reviewer needs to know which results were accepted and why the other seven were not.
A provider batch marked complete only means its request processing finished; it does not prove every individual item is business-valid. After downloading results, reconcile them against the manifest by stable request identifiers. An item may have a provider error, a model refusal, a truncated object or an apparently correct answer with no supporting span in the document. Handle those cases differently. A transient provider error may merit a bounded retry; a missing signature requires source review; a duplicated claim reference should be rejected as an input-integrity problem. Retry only rejected items whose cause can actually be repaired, and preserve the original versions for comparison.
Set two separate metrics: technical completion (received result, parsed representation) and accepted business evidence (validated fields with documented provenance). In a claims example, acceptance may require that a date appears in the input, the monetary amount fits the specified currency, and the claim identifier matches the intake record. A batch can reach 100% technical completion while still requiring human adjudication. Explicit states make that visible to operations teams and create a defensible audit trail.
Use independent review where errors are consequential
A second model pass can help when the first output needs critical examination. For example, one pass extracts structured fields, while another compares those fields against source snippets and flags missing evidence. The second pass should receive a fresh or deliberately restricted context; simply asking the same reasoning trace to declare itself correct is a weak independence test.
Independent review is still probabilistic. It can miss errors shared by both passes or overlook a false source citation. Use deterministic arithmetic and schema checks wherever they apply, and sample outputs against source documents to measure actual quality. Compare performance by document type, layout and language rather than reporting one blended accuracy number.
For a large corpus, stratified human sampling can reveal failures that uniform random sampling misses. Oversample documents with complex tables, poor scans, conflicting dates or unusual formats. If the system performs well on straightforward forms but badly on handwritten amendments, the release plan should reflect that difference.
Evidence must survive across the pipeline. A review that says “field invalid” is less useful than one that provides the field, source location, validation rule and remediation state. This is why evaluation of Claude agents should include process outcomes and provenance, not only a favorable final-answer grade.
Measure the whole operation, not the per-token price
The cost of batch extraction includes input preparation, tokens, storage, retries, validation, human review and failed downstream work. A cheaper inference request that produces more invalid results may raise the cost per accepted record. Likewise, a high-capability model can be economically sensible if it reduces costly handoffs for ambiguous documents. Claims about fixed batch discounts or pricing should be checked against current issuer documentation before a team adopts them as planning assumptions.
Latency should be measured against the workflow’s real service-level agreement. A batch that completes within an overnight window may be ideal for archive indexing and unsuitable for a customer waiting for a decision. Queue age, completion distribution, result retrieval and review backlog matter as much as the model’s generation speed.
Security requirements may also determine architecture. Some documents should not be retained in logs, returned to an unrelated agent or sent to an unapproved location. Keep input references and evidence under the appropriate access controls, and use the minimum necessary content for each processing step.
Test the failure cases before calling the pipeline complete
A strong CCA-F design exercise uses ten synthetic documents: a normal invoice, a missing field, a conflicting total, an out-of-order page, a duplicated document, an invalid currency, a truncated scan, a multi-page table, a permission-denied record and a simulated batch expiration. Define expected outcomes before generating any response.
Verify that every request maps to one logical input, that results reconcile even when returned out of order, that invalid values do not enter the accepted dataset, and that only unfinished work is retried. Confirm that a human can follow source evidence without access to secrets or unrelated customer records. Finally, interrupt the pipeline after a downstream write and check whether it can resume without duplicating the effect.
This approach teaches the architectural distinction that matters most: prompting produces candidate observations, schemas control their shape, validation tests their meaning, batches coordinate independent work, and governed services decide what may happen next. A complete system makes those boundaries inspectable and recoverable rather than expecting Claude to guarantee all of them through fluent output.