INSIGHTS
AI & Data

Microsoft AI-300: Model Evaluation & Monitoring in Microsoft Foundry

In this article
  1. Start with a task-specific definition of quality
  2. Build evaluation datasets that resemble production
  3. Combine deterministic checks with AI-assisted evaluation
  4. Measure retrieval and generation separately
  5. Evaluate agents at the trajectory level
  6. Use tracing to explain why a result happened
  7. Monitor production with a baseline and thresholds
  8. Connect evaluation failures to the release process
  9. Keep security and safety signals distinct but connected
  10. Use evidence to decide whether a release is healthy

Putting a generative AI application into production creates two separate questions. The first is operational: is the service available, responsive, and within its resource limits? The second is behavioral: are the answers useful, grounded, safe, and consistent with what the application is supposed to do? Traditional application monitoring answers the first question well. Production AI engineering has to answer both.

Microsoft Foundry provides evaluation, tracing, and monitoring capabilities that fit this broader operating model, and those capabilities are reflected in the current AI-300 scope. The exam does not treat deployment as the finish line. It expects practitioners to evaluate generative AI solutions, monitor them after release, use observability signals, and improve the system based on evidence.

The distinction matters because generative quality is not captured by a single accuracy percentage. A response can be fluent but unsupported by the source material. It can be relevant but unsafe. An agent can select the correct tool but pass the wrong arguments. A retrieval system can find documents quickly but retrieve evidence that does not answer the user’s question. Effective evaluation therefore starts by defining the kind of failure the application actually needs to prevent.

Start with a task-specific definition of quality

Before choosing an evaluator, define what a good response means for the workload. A customer-support assistant may need correct policy grounding, concise answers, safe escalation, and refusal to invent account details. A developer assistant may need syntactically valid code, correct use of a library, and clear explanation. A research agent may need citation fidelity, completeness, and transparent uncertainty. The metrics should follow the task rather than the other way around.

Common generative evaluation dimensions include relevance, groundedness, coherence, fluency, completeness, and safety. These labels are useful, but they remain abstract until they are tied to an application contract. “Grounded” might mean every factual claim is supportable by retrieved enterprise documents. “Relevant” might mean the response directly addresses the stated request without unnecessary content. “Safe” might include harmful-content limits as well as business-specific restrictions on what an assistant may reveal or perform.

Do not compress all of those dimensions into one score too early. A single composite can hide a serious weakness. An average score may look acceptable even when a safety metric has deteriorated or one high-impact scenario consistently fails. Keep important dimensions visible so release decisions reflect the actual risk profile.

Build evaluation datasets that resemble production

An evaluation set is useful only if it exercises the behaviors that matter. Idealized demo prompts tend to overstate quality because they represent the easiest and most familiar cases. A production-oriented set should include routine requests, ambiguous wording, multi-turn context, incomplete information, difficult edge cases, and requests the system should refuse or escalate.

For retrieval-augmented generation, include questions whose answers are present in the corpus, questions whose answers are absent, and questions where multiple documents contain potentially conflicting information. This reveals whether the system can distinguish “I found evidence” from “I can generate a plausible answer.” For agents, include tasks that require the right tool, the wrong tool to be rejected, invalid arguments to be handled, and high-impact actions to require additional control.

Evaluation data should be versioned. If a release is approved against one test set and the test set later changes, the team needs to know which evidence supported the original decision. Versioning also makes it possible to add real failure cases from production without silently changing historical comparisons.

Sensitive examples deserve the same governance as production data. Do not build a convenient benchmark by copying customer conversations into an unrestricted repository. Mask or synthesize data where practical, apply appropriate access controls, and define retention. The broader secure data lifecycle applies to test prompts, evaluator outputs, and traces as much as it applies to training data.

Combine deterministic checks with AI-assisted evaluation

Some properties should be tested deterministically. A response can be checked for required JSON structure, a mandatory citation field, prohibited HTML, schema conformance, tool-call parameters, or whether a retrieved document identifier belongs to the allowed set. Deterministic checks are repeatable and easy to gate in continuous integration.

Other properties are semantic. Relevance, groundedness, coherence, and some safety dimensions may require model-based evaluators or human judgment. AI-assisted evaluators can scale across large test sets and make regression testing practical, but they have their own limitations. Their scores can vary, evaluator prompts can introduce bias, and an evaluator may favor a particular writing style rather than the real task objective.

For high-impact scenarios, use layered evidence. Deterministic checks can enforce hard constraints, model-based evaluators can cover broad semantic behavior, and human reviewers can inspect representative or disputed cases. The evaluation system itself should be validated periodically rather than assumed to be perfect.

This layered approach also reduces the temptation to optimize the application for a single evaluator. A team that chases one score can accidentally improve benchmark performance while making real responses less useful. Multiple signals keep the focus on the user task.

Measure retrieval and generation separately

A retrieval-augmented system can fail before generation begins. If the search layer retrieves irrelevant, stale, or incomplete documents, even a strong language model will struggle to produce a grounded answer. Evaluation should therefore separate retrieval quality from answer quality where possible.

Inspect whether relevant evidence was retrieved, whether important evidence was missed, whether ranking placed the best material near the top, and whether filtering rules excluded information incorrectly. Then evaluate whether the generated response used that evidence accurately. A groundedness problem caused by poor retrieval requires a different fix from a groundedness problem caused by the model ignoring good retrieved context.

Chunking, metadata, query rewriting, hybrid search, filters, index freshness, and top-k settings can all change retrieval behavior. Record those configurations with the application version so a quality regression can be traced to the part of the system that actually changed.

Evaluate agents at the trajectory level

Agentic applications add another unit of analysis: the sequence of steps taken to complete a task. Looking only at the final answer can hide dangerous or expensive behavior. An agent may eventually produce the correct result after calling unnecessary tools, exposing too much data, repeating an action, or taking a route that would be unacceptable in production.

Trajectory evaluation asks whether the agent selected appropriate tools, used them in a sensible order, passed valid parameters, respected authorization boundaries, and stopped when the task was complete. It can also measure task success, tool-call errors, iteration count, latency, and cost.

The governance implications connect directly with AB-100, which focuses on designing and governing agentic business solutions. For an agent that can change records, send messages, approve transactions, or trigger workflows, evaluation should include authorization failures, prompt manipulation attempts, and cases where human approval is required. A high-quality final sentence does not compensate for an unsafe action trajectory.

Use tracing to explain why a result happened

Scores tell you that something changed; traces help explain why. Distributed or AI-specific tracing can capture the sequence of model calls, retrieval steps, tool invocations, prompt construction, and latency across a request. When an evaluation fails, the trace can reveal whether the problem began with retrieval, a prompt branch, a model call, a tool timeout, or a downstream system.

Tracing is particularly useful for intermittent failures. A user may report that an agent “sometimes” gives the wrong answer. Without a trace, the team may be unable to reproduce the exact retrieval result, tool path, or model configuration involved. Correlating a request with its deployed version and trace gives engineering teams a concrete artifact to investigate.

Observability data can itself be sensitive. Prompts may include personal information, proprietary data, credentials mistakenly pasted by a user, or confidential retrieved text. Traces should be collected according to a deliberate policy: minimize unnecessary payloads, restrict access, scrub secrets where possible, and retain detailed content only as long as it serves an operational purpose.

Monitor production with a baseline and thresholds

Offline evaluation establishes whether a candidate appears ready. Production monitoring asks whether it continues to behave within expected bounds. Start by recording a baseline for the released version: latency, error rate, token usage, evaluation metrics, tool success, retrieval behavior, safety events, and relevant business outcomes.

Thresholds should be tied to action. A small change in average relevance may justify investigation but not rollback. A sharp increase in prohibited-content responses may require immediate escalation. A rise in tool authorization failures may indicate a deployment problem or an attack. Different signals deserve different severity levels and owners.

Sampling strategy matters because continuous semantic evaluation of every interaction can be expensive and may create unnecessary data-retention risk. Many systems can evaluate a representative sample, with heavier sampling after a new release or when another metric crosses a threshold. High-risk applications may require more comprehensive controls.

Monitor by segment where it improves diagnosis. One global score can hide a regression limited to a particular language, product line, data source, tool, or user workflow. Segmenting every possible dimension creates noise, so choose segments that map to real differences in data, risk, or expected behavior.

Connect evaluation failures to the release process

Monitoring becomes useful when it changes what the engineering team does. A production regression should create a clear path to triage: identify the affected release, reproduce representative cases, determine whether the change came from code, prompt, model, retrieval data, or external dependency, and decide whether to rollback or correct forward.

Prompt and model versions should be visible in that process. If a team changes a system instruction without changing the application build, the monitoring record must still identify the new behavioral version. Likewise, a model-provider update should be treated as a meaningful dependency change when it can alter output behavior.

Regression cases should feed back into the evaluation set. A failure discovered in production is valuable evidence because it represents a scenario the previous test process missed. Add it to the controlled benchmark after appropriate sanitization so the same class of defect is more likely to be caught before a future release.

This creates a continuous improvement loop rather than a one-time certification exercise: evaluate, release, observe, learn, expand the benchmark, and evaluate again.

Keep security and safety signals distinct but connected

Quality and security overlap, but they are not interchangeable. A response can be low quality without being dangerous, and an attack can succeed even when the response appears coherent. Monitoring should therefore include security-specific signals such as prompt-injection attempts, anomalous tool calls, unusual access patterns, repeated policy-boundary probing, and attempts to exfiltrate system instructions or protected data.

The existing overview of AI security risks helps frame why these signals need their own response path. A safety evaluator that detects harmful content may trigger a product-quality review, while evidence of unauthorized tool use may need immediate security escalation. Both can appear in the same trace, but the ownership and containment steps differ.

Responsible-AI monitoring also needs to look beyond explicit attacks. Repeatedly poor behavior for a protected or high-impact segment, ungrounded claims in a regulated workflow, or a model that becomes overly confident after an update can create risk without any malicious input. The monitoring design should reflect the actual consequences of the application.

Cost and latency should be analyzed beside quality. A candidate can improve groundedness while adding several retrieval or model calls, making the user experience slower and the service more expensive. Evaluation reports should therefore include operational trade-offs for representative tasks. A release gate can allow a small cost increase for a meaningful quality gain, but the trade-off should be visible rather than discovered after billing or latency alerts appear.

Evaluator calibration should be revisited as the application evolves. When domain experts repeatedly disagree with an automated evaluator, investigate whether the rubric, examples, or threshold still represents the intended behavior. The goal is not to make human reviewers agree with the metric; it is to ensure that the metric remains a useful proxy for real quality.

Evaluation ownership should be explicit as well. Product teams can define task success, security teams can define abuse cases, domain experts can judge correctness, and platform teams can maintain the execution framework. A benchmark without an owner eventually becomes stale; named ownership keeps cases, thresholds, and escalation paths aligned with the deployed application.

Use evidence to decide whether a release is healthy

A mature evaluation program does not aim to prove that an AI system is perfect. It aims to make its known behavior measurable enough that teams can decide whether a release is acceptable, detect meaningful changes, and investigate failures efficiently. That requires task-specific metrics, representative datasets, deterministic and semantic tests, traceability, production baselines, and a response process.

For teams in the Microsoft ecosystem, Foundry provides important pieces of that workflow, but the operating discipline is broader than a product feature. Evaluation assets need version control. Releases need gates. Traces need governance. Production signals need owners. Incidents need to become future regression tests.

When those practices are connected, model evaluation and monitoring stop being separate activities performed by different teams. They become one evidence chain from pre-release testing to real-world operation. That chain is what makes it possible to improve generative AI systems without treating every prompt change, model update, or production complaint as a new experiment from scratch.

Filed under AI & Data