INSIGHTS
AI & Data

AI-103: Building AI Apps and Agents on Azure

Practical overview of Microsoft AI-103 objectives, Foundry, safe agent development, retrieval, multimodal processing, and production evaluation.

In this article
  1. Understand the five domains as connected engineering work
  2. Separate Microsoft Foundry resources, models, and application code
  3. Choose a model for the task instead of the biggest model
  4. Build retrieval as a source-and-evidence system
  5. Design agent tools as privileged application interfaces
  6. Treat vision, speech and extraction as separate contracts
  7. Evaluate a system, not merely the eloquence of its answers
  8. Practice with one evolving project instead of disconnected labs
  9. Define exam readiness as repeatable decisions

The Microsoft AI-103 exam tests a practical question: can a developer turn an AI capability into a dependable Azure application, not merely demonstrate a model in a playground? A team may have an impressive prompt that answers product questions, yet still have no control over whose documents are retrieved, what a tool may change, how releases are tested, or what happens when a model endpoint slows down. Those are the concerns of an application engineer. They are also why Developing AI Apps and Agents on Azure spans Microsoft Foundry, retrieval, agent workflows, prompt and structured-output design, multimodal processing, operations, and responsible AI.

AI-103 earns Microsoft Certified: Azure AI Apps and Agents Developer Associate. Microsoft retired AI-102 in June 2026, so a person studying for an active exam should not silently substitute its old objective list. The skills measured for Microsoft AI-103 describe a developer familiar with Python, Azure services, generative AI, and the work of building and maintaining applications. The current Microsoft exam study guide is the authority for scope and weightings; Microsoft periodically revises exam content.

Understand the five domains as connected engineering work

Microsoft weights planning and managing AI solutions at 25–30%, generative AI and agents at 30–35%, computer vision at 10–15%, text analysis at 10–15%, and information extraction at 10–15% in the April 2026 blueprint. These bands are not an exact prediction of the questions any candidate will receive. They indicate the breadth of implementation decisions a prepared developer should be able to explain and test.

A useful cross-domain scenario is a customer-support agent that reads invoices and photographed equipment labels, answers troubleshooting questions from an internal manual, transcribes a caller’s request, and creates a service ticket only with the caller’s approval. That one workflow needs Foundry model deployment, secure retrieval, an agent tool contract, OCR or visual understanding, speech processing, structured extraction, and monitoring. If the model answers correctly but submits a ticket for the wrong customer, the system has still failed. Study the complete workflow, not isolated service names.

Separate Microsoft Foundry resources, models, and application code

Start by drawing the ownership boundary. A Microsoft Foundry project provides a place to configure models, agents, connected tools and evaluators. Azure resources, identities, networking and monitoring settings supply its operating environment. Your application still controls the user session, business authorization, input handling, product experience and downstream writes. A deployed model endpoint is not automatically a secure application or a usable agent.

For practice, create a small Python client that calls a deployed model through a supported Foundry endpoint, then trace how the request is authenticated and which resource is charged. Move configuration such as project endpoint, deployment name, timeout and logging level into environment-specific settings. Never paste subscription keys into source files. Azure RBAC matters because deployment permission, model invocation, data-store access and the authority to run a tool are distinct permissions.

Choose a model for the task instead of the biggest model

Model choice is a decision among output quality, modality, latency, context needs, availability, cost and regional support. A small model may classify a bounded request cheaply, while a larger multimodal model handles ambiguous diagrams or multi-step reasoning. A retrieval pipeline may solve a knowledge problem better than changing models. A deterministic parser may be preferable to any model for an invoice number printed in a consistent field.

Make a simple comparison dataset: ten straightforward questions, ten ambiguous requests, and ten cases where the correct action is to say the evidence is insufficient. Record quality and latency separately. Include input and output token estimates, because a model that is fast on short prompts can become expensive when the application attaches full documents to every call. Check current model availability and region restrictions rather than hard-coding an example deployment from an old tutorial.

Build retrieval as a source-and-evidence system

Retrieval-augmented generation gives the application a way to consult enterprise information. It typically starts with ingestion, parsing, chunking, an index and a retrieval query, not with an instruction that says “use company documents.” A useful chunk has content, source identity, update time and access metadata. Retrieval must respect the user’s authorization boundary before results reach the model. A highly relevant but unauthorized paragraph is not a valid answer source.

Azure AI Search combines keyword and vector retrieval; its hybrid search can merge both result sets, with optional semantic reranking. Test a document identifier, a paraphrased natural-language question and an outdated policy version. If the exact identifier disappears from retrieval, tune indexing or the lexical path. If old policy wins, improve freshness and filtering. If relevant chunks arrive but the final answer invents a policy, evaluate response groundedness independently of retrieval quality.

Design agent tools as privileged application interfaces

An agent can call a search API, retrieve a customer record or request a change to a system. The model choosing a tool does not grant authorization to execute it. Design a tool interface with a specific name, required typed arguments, a narrow set of allowed operations and a clear result schema. The application or tool service validates input, checks the user’s authority, and decides whether a sensitive action needs confirmation.

For example, a ticket tool might accept an authorized customer ID, an issue category and a proposed note; it must not accept arbitrary SQL, a destination URL or an instruction to impersonate an administrator. A retryable ticket creation should use an idempotency key so a model or network retry cannot create duplicate tickets. Microsoft Foundry Agent Service supports managed agents and tool integrations, but application ownership of business rules remains essential.

Treat vision, speech and extraction as separate contracts

Vision analysis asks what an image or video contains. Image generation produces new visual media. OCR extracts printed text. Document understanding recovers layout, fields and relationships. Speech recognition turns audio into text; synthesis turns text into speech. Translation changes language, while entity recognition finds named things in text. A multimodal model may span several of these tasks, but the validation contract differs for each.

Imagine a scanned purchase order. OCR can recover the supplier name, but a downstream payment workflow needs field typing, confidence thresholds, comparison with purchase-order records and a review path when values disagree. Treating a convincing JSON response as verified data is a common engineering error. Content Understanding supports structured extraction, yet the business application must still validate output and preserve provenance.

Evaluate a system, not merely the eloquence of its answers

An evaluation suite should contain representative user intents, authorized and unauthorized documents, benign and hostile tool content, expected output formats, and actions that must not occur. Different failures call for different measures: retrieval relevance, groundedness, exact-field correctness, tool-choice correctness, safety compliance and task completion. A single “overall quality” score hides the cause of regressions.

Keep at least one end-to-end trace for each failure type. A response can be factually plausible but based on the wrong document version; a correct answer may still contain another user’s private data; a tool result can be accurate while an unintended action follows. Connect evaluators to regression gates so that changing a model deployment or prompt requires evidence, not just a successful playground chat.

Practice with one evolving project instead of disconnected labs

Build a small service-desk prototype in layers. First deploy a model and write a Python request with retries and structured error reporting. Next index a deliberately small manual and attach source identifiers to answers. Add a read-only lookup tool before introducing any write action. Test a non-existent customer ID and an attempt to request a tool outside the user’s role. Finally add one image or audio input and record separate extraction and generation quality measures.

Document the decisions you make: why this model, why hybrid retrieval, which identity accesses storage, what happens under a rate limit, and how an operator can reconstruct one failed session. Repeat the deployment through a second environment with the same configuration process. This yields evidence of engineering competence more valuable than memorizing portal screenshots. The detailed AI-102-to-AI-103 transition is also worth understanding: the newer exam does not simply rename the old service list.

Define exam readiness as repeatable decisions

For each domain, practice explaining the constraint first and the Azure service second. If a question says a user needs an exact current policy paragraph, what retrieval and authorization checks are necessary? If it says an agent should send a payment instruction, what approval boundary is required? If it says image content must be accessible, how will the model’s caption be checked against visual evidence and edited for accuracy? If a model endpoint returns rate-limit responses, how will the client back off without duplicating side effects?

Use the official skills guide to maintain a coverage checklist and turn it into a hands-on AI-103 lab plan that exercises the same tasks in a controlled Azure environment; update both when Microsoft changes the exam. Avoid treating retired AI-102 dumps or screenshots as current questions, and never infer that a certification title by itself proves a feature is generally available in every region. Passing readiness comes from being able to make and validate decisions across the complete application, from data ingestion to safe production operation.

Filed under AI & Data