A production AI application has at least two kinds of health. Operational health asks whether requests complete within service targets and whether infrastructure is available. Behavioral health asks whether answers remain useful, grounded, safe and consistent with the task. A service can meet its uptime objective while gradually returning worse answers because its index is stale, its prompt changed, or users began asking questions its evaluation set never represented.
The Microsoft AI-103 scope brings those concerns together: evaluation before release, model and agent tracing, safety signals, retrieval quality, drift, cost and analysis of failures. The useful engineering discipline is to define quality measures before selecting a dashboard. Otherwise monitoring shows attractive graphs with no agreed decision about what should cause rollback.
Define success for the exact application task
An internal document assistant should cite approved sources and avoid inventing policy. A classification endpoint should return valid categories and distinguish unsupported inputs. A transactional agent should select permitted tools and wait for approval before a consequential change. Each requires different measurements. A universal “response quality score” cannot represent all three reliably.
Build a task inventory with expected outputs, unacceptable behavior and representative variants. For document questions, include source records and permissible answer ranges. For an agent, record expected tool selections and forbidden actions. For extraction, define typed field correctness and when to send a record to human review. Use requirements a subject-matter expert can defend, not just a model judging its own writing style.
Build evaluation datasets with real failure modes
A good dataset includes easy inputs and the cases most likely to harm users: missing evidence, conflicting documents, partially unreadable images, multilingual material, long conversations, ambiguous customer IDs and injected instructions. Keep synthetic tests distinct from real anonymized production samples, because they may fail in different ways. Never copy confidential customer content into an evaluation store without the required permissions and data-handling controls.
Version the dataset. When a source policy changes, label which tests depend on the old version and update expectations explicitly. This prevents a perfectly valid improvement from looking like a regression only because the ground-truth file is obsolete. Conversely, retaining a few intentionally outdated-source tests verifies that the application refuses to treat historical material as current authority.
Separate retrieval, grounding and task completion
A RAG system can retrieve the correct source and still misstate it. It can generate a convincing answer from an unauthorized source. It can produce a grounded answer that does not address the user’s actual task. These are different defects. Use retrieval relevance or labeled document metrics for search, groundedness measures for claims supported by supplied context, and task-level measures for whether the response solves the requested problem.
Microsoft’s Foundry RAG evaluator documentation distinguishes system and process evaluation. Use those distinctions when troubleshooting rather than promoting a model solely because one aggregate score improved. A release that increases fluency while lowering source fidelity is not necessarily an improvement.
Trace model calls, tool calls and source decisions
An end-to-end trace should show where a user request entered, which deployment answered, which retrieval queries were made, which documents were returned, what tool calls were proposed, and whether those calls succeeded. Use correlation identifiers and timings across steps. Logs that contain only the final response cannot explain why a wrong order was created or why one user experiences far greater latency than others.
Microsoft’s Foundry tracing guidance uses Application Insights and OpenTelemetry conventions for agent operations. Choose a trace retention period, access model and redaction policy before enabling detailed payload capture. Prompts and tool results may contain account numbers, personal details or secrets. An audit trail should support investigation without becoming an uncontrolled secondary data store.
Read latency as a distributed budget
A request may wait for user authentication, vectorization, search, semantic reranking, model inference, tool execution and output validation. Measure those stages independently. A five-second response with a four-second search step should not trigger a blind change to model generation parameters. Similarly, an agent that calls the same tool five times may have a workflow design defect even when each individual call is fast.
Track median and tail latency, tool fan-out, retries, token volume, rate-limit responses and queueing. Map each to a user-facing service target. A live voice application may need strict turn-level responsiveness; a scheduled document-analysis job may tolerate longer processing but require cost predictability and durable retry state. The correct thresholds follow application intent.
Use safety signals alongside quality measures
An agent can be helpful and unsafe: it may answer accurately while revealing data outside the user’s scope or executing an unapproved operation. Keep authorization checks and harmful-content detection separate. Use adversarial evaluation cases where document text attempts to override instructions, where a tool returns a malicious URL, and where the user tries to access another tenant’s records.
Safety evaluators may detect types of risk, but they do not prove the correctness of every permission decision. Log deterministic allow/deny results from the tool layer. Track false positives as well as false negatives so safety protections do not block ordinary authorized tasks. A safety regression should be able to stop a deployment independently of improvements in answer fluency.
Detect distribution change without vague drift alarms
Applications change because their users, source data and connected systems change. A new product line may increase queries with unfamiliar terms; a language rollout may alter translation requests; a policy update may invalidate previously correct answers. Compare production request categories, document freshness, retrieval distribution, safety events and evaluation samples over time. Separate actual concept change from telemetry pipeline errors.
When a quality drop appears, identify what changed before retraining or switching models. Did the index ingest the new documents? Did an API connector change schema? Did the prompt release alter a critical rule? Did a user segment begin sending attachments? A useful drift alert leads to a testable hypothesis and a repair step, not simply an alarm that says “AI quality low.”
Connect evaluation to deployment gates
Before promotion, run the current test suite against the proposed model, prompt, retrieval index and tool definitions. Record the exact versions. Require critical authorization and safety tests to pass, compare task-quality metrics against a baseline, and investigate material regressions. After release, canary a limited audience where practical and monitor errors and behavioral signals before broad rollout.
Rollback can require restoring multiple components. Reverting application code while leaving an incompatible agent tool schema in production may keep the problem alive. Maintain a compatibility matrix or release manifest linking model deployment, prompt version, index version and tool contract. Microsoft’s Foundry evaluation workflows support pre- and post-deployment testing, but a team still has to decide which outcomes qualify as acceptance.
Consider a release that changes both the embedding model and the prompt used by a policy-answering assistant. The team’s aggregate “quality” metric rises, but the assistant has begun returning accurate answers from outdated policy versions. A meaningful evaluation set must include source recency, required citation identifiers, user permissions and an explicit correct-no-answer category. Otherwise the release looks better precisely because the metric cannot observe its most damaging mistakes.
Build a small adjudicated suite with labeled expected documents. For retrieval, measure whether the authorized current document appears within the top results; for groundedness, check that generated claims follow from those results; for the full task, check exact allowed actions and the presence of uncertainty when evidence is insufficient. Keep safety tests separate: an agent might be factually correct yet reveal private customer data through its tool arguments. Different failures require different fixes, so they should not be collapsed into one averaged number.
A release gate needs an interpretable threshold and a review rule. For example, on a suite of 60 authorized factual questions and 15 no-evidence cases, investigate any new instance in which unauthorized material is retrieved or a write occurs without approval, regardless of changes in average fluency. A small decline in stylistic preference may warrant review; a single newly successful unauthorized tool execution is a release blocker. Record evaluation-set versions and changes to the model, index and prompts together so regressions can be reproduced.
If a judge model scores groundedness, calibrate it against human-reviewed cases. Human reviewers should inspect disagreement examples and sample both passing and failing cases, because a model judge can share the generator’s preferences or overlook subtle factual omissions. Evaluation is an evidence practice, not permission to replace editorial or operational review with another automated model.
Run an incident rehearsal with a broken source
As an AI-103 exercise, deploy a small RAG assistant with known-answer cases. Replace one authoritative document with an outdated revision, disable one permitted tool, and throttle the model endpoint. Observe whether you can distinguish these three failures from traces and metrics without guessing. Each incident should produce a specific repair: reindex the correct version, restore the tool permission or apply bounded backoff to retryable rate limits.
Close the rehearsal by writing a short timeline: when the quality changed, what signal detected it, which users were affected, which action repaired it, and what test would catch it next time. That discipline is more valuable than a dashboard screenshot. Monitoring is successful when operators can explain and control the system’s behavior under change.