An AI agent can produce impressive demonstrations and still be unsafe to release. The difference is that production exposes the agent to messy requests, incomplete data, adversarial content, tool failures, policy boundaries, and combinations of conditions that a developer may never type into a playground. Pre-production evaluation is the discipline that turns those unknowns into explicit test evidence.
The current AB-100 scope requires architects to recommend agent testing processes and metrics, create validation criteria, design end-to-end scenarios, monitor performance, and address security and governance. Microsoft Foundry also supports agent-focused evaluation and observability. Together, these capabilities point to a practical rule: evaluate the complete agent behavior, not only the text of its final answer.
That means measuring whether the agent completed the task, used the right sources, selected appropriate tools, respected constraints, handled uncertainty, and stopped or escalated when necessary. A production gate should answer whether the system is ready for its intended risk level, not whether a few hand-picked conversations looked good.
Define success in business terms before selecting metrics
Evaluation starts with the job the agent is supposed to perform. For a service agent, success might require correct intent classification, use of the current policy, an accurate proposed resolution, and no unauthorized account changes. For a research agent, success might require source coverage, factual consistency, citation quality, and a clear statement when evidence is insufficient.
Convert those expectations into observable criteria. Some can be deterministic: required fields present, tool call within an allowlist, no write action without approval, response includes a case identifier. Others need semantic scoring: answer relevance, groundedness, instruction adherence, completeness, or tone. The combination gives a more faithful picture than one aggregate model score.
Define failure severity as well. A slightly awkward sentence and an unauthorized financial action are both “errors,” but they should not have the same release consequence. Classify failures by business impact so a small number of critical violations can block deployment even if average quality is high.
Create a failure taxonomy early: factual error, unsupported claim, retrieval miss, policy violation, unsafe tool call, privacy exposure, incorrect escalation, latency breach, and workflow failure are examples. The categories make results actionable because the team can assign an owner and remediation path instead of debating a single blended quality score.
Build a representative evaluation dataset
A test set should resemble the workload the agent will actually see. Sample real patterns where privacy rules allow, then de-identify them and add intentionally designed cases for rare but important conditions. Include short and long requests, incomplete information, conflicting instructions, different user roles, historical edge cases, and variations in terminology.
Do not allow the dataset to become a static collection of easy examples that the team implicitly optimizes against. Add new cases when production incidents, user feedback, policy changes, or tool failures reveal a gap. Maintain holdout cases so a prompt or model change cannot simply overfit the visible evaluation set.
Balance the dataset by business importance rather than raw frequency alone. Rare events such as account closure, fraud escalation, or legal hold may represent a small share of traffic but deserve dedicated cases because the consequence of failure is high. A frequency-weighted average can otherwise make a system look excellent while ignoring precisely the scenarios that require the strongest controls.
Grounded agents need source variation in the dataset too. Test relevant documents, irrelevant retrievals, stale versions, conflicting sources, missing metadata, and permission-restricted material. A system that works only when retrieval returns the perfect passage has not been evaluated realistically.
Include counterexamples in rubric design. If a correct response must cite policy, create examples where the prose is factually plausible but unsupported by the retrieved source. If the agent must escalate when information is incomplete, create cases where guessing would be easy. Counterexamples prevent evaluators from rewarding fluent output that violates the operational contract.
Where qualified reviewers disagree, document the disagreement instead of forcing a false binary label. Some business tasks contain genuine judgment. In those cases, evaluation can measure whether the agent identifies the relevant considerations and remains within an acceptable range rather than insisting on one exact wording or conclusion.
Treat the evaluation dataset as a versioned product. Record where each case came from, the scenario it represents, why it is important, and what sensitive fields were removed or synthesized. Stratify the set so common requests do not overwhelm rare but consequential cases such as permission boundaries, ambiguous instructions, missing documents, conflicting source data, or tool failures. When production incidents reveal a new failure mode, add a sanitized regression case and keep it in future releases rather than testing the incident once and forgetting it.
Representative does not mean “large at any cost.” A smaller set with explicit coverage labels can be more useful than thousands of uncurated conversations. Track coverage across task types, user roles, languages, data conditions, tools, and risk classes. That turns evaluation design into a measurable question: the team can see which business scenarios have evidence and which still depend on intuition.
Use rubrics that describe acceptable behavior
A strong rubric gives the evaluator criteria rather than a vague instruction to rate the answer. For example, a policy-support response may need to identify the governing rule, apply it to the facts, cite the approved source, avoid inventing exceptions, and escalate if required information is absent. Each dimension can be scored independently so teams know what actually regressed.
Rubrics are useful for model-assisted evaluation, but they still need calibration. Compare evaluator judgments with qualified human reviewers on a sample. If the automated evaluator consistently rewards verbosity or misses a domain-specific defect, refine the rubric or use a different measurement method.
Pass thresholds should reflect risk. A low-impact internal drafting aid may tolerate occasional stylistic variance. An agent that recommends access changes or financial decisions needs stricter criteria around policy adherence, source accuracy, and escalation. Avoid treating an arbitrary percentage as universal quality.
Evaluate tool choice and agent trajectory
For an agent, the path matters. A correct final answer can hide unnecessary or dangerous intermediate actions. Record which tools the agent chose, in what order, with what parameters, and what result each tool returned. Then evaluate whether the trajectory was appropriate for the task.
Test overuse and underuse. An agent may call an expensive search tool for every trivial question, or it may answer from model memory when policy requires an authoritative lookup. It may call a write tool before gathering the evidence needed to justify the action. These are orchestration defects even if the final prose sounds plausible.
Tool parameters deserve deterministic checks wherever possible. Resource identifiers, amount limits, environment names, email recipients, and change scopes can often be validated against known constraints. The application should reject invalid values, but pre-production evaluation should reveal whether the agent frequently tries to produce them.
Test grounding and citation behavior under imperfect retrieval
Grounded systems can fail in at least three different layers: retrieval may find the wrong material, the context may contain conflicting evidence, or the model may ignore accurate evidence and answer from prior knowledge. Evaluation should isolate those layers so the remedy targets the actual cause.
Measure retrieval coverage for questions with known supporting documents. Check whether the agent uses the most current source and whether citations actually support the claims attached to them. Introduce distractor documents with similar vocabulary to see whether ranking and reasoning stay focused on authoritative evidence.
When no approved source answers the question, the expected behavior may be to ask for clarification or state that the evidence is insufficient. Rewarding an answer in every case trains the system toward confident invention. In enterprise settings, abstention can be a correct outcome.
Run safety, privacy, and adversarial evaluations
Normal functional tests do not reveal how an agent reacts to malicious instructions, sensitive data, or attempts to expand its authority. Add prompt-injection cases, requests for restricted records, attempts to reveal system instructions, tool-argument manipulation, and content designed to bypass workflow rules. The broader set of AI security risks should inform the threat scenarios.
Privacy tests should verify that the agent does not expose information from another user, previous session, or over-broad retrieval result. The practices described for protecting PII in AI usage are particularly relevant when evaluation datasets themselves contain realistic enterprise material. Use sanitized or synthetic data where possible and control access to evaluation traces.
Safety evaluation also includes actions. Attempt to persuade the agent to skip approval, change the target of a transaction, call an unapproved endpoint, or treat retrieved text as higher authority than system policy. A good result is not merely a refusal message; the protected operation must remain impossible through the tool boundary.
Test identity and authorization boundaries as part of functional evaluation. The same prompt from two users with different roles should not necessarily produce the same retrieved evidence or available actions. Create paired cases that verify the system respects data and tool permissions while still giving each user an appropriate response.
Also test temporal behavior. Policies, prices, inventory, and operational states change. An evaluation suite should include versioned or time-sensitive cases so the agent proves it uses current evidence rather than a memorized or cached answer. This is especially important after index refresh, source migration, or model upgrades.
Measure latency, cost, and reliability with quality
Agents can pass quality tests and still be operationally unsuitable. Measure end-to-end response time, model latency, tool latency, token use, retries, timeouts, and concurrency behavior. Long multi-agent chains may improve a benchmark while making a customer-facing workflow too slow or too expensive.
Run the same scenarios repeatedly because agent behavior can vary. A workflow that succeeds nine times and fails once on the same input has a different operational profile from deterministic software. Estimate variance for important metrics and test enough repetitions to expose unstable routes or tool selection.
Failure injection is useful. Make a retrieval source unavailable, return a tool error, delay an approval, or simulate throttling. The agent should surface a controlled failure or recovery path rather than invent a result. Reliability includes how the system behaves when dependencies are not healthy.
Keep failure examples, not only successes. When a candidate release fails a case, preserve the trace and categorize the defect: retrieval, routing, prompt, model behavior, tool validation, authorization, or environment. Over time this produces an engineering history of what kinds of changes tend to create regressions and which controls catch them.
Release comparison should use statistical caution. Small evaluation sets can make a one- or two-case difference look like a dramatic percentage change. For critical metrics, use enough samples to understand variance, repeat nondeterministic scenarios, and investigate material differences rather than declaring victory from a tiny score improvement.
Turn the evaluation suite into a release gate
Evaluation has the most value when it runs every time the system changes. Prompt edits, model upgrades, new tools, grounding changes, orchestration updates, and safety-policy changes can all alter behavior. Store the evaluation configuration and test cases with the release process, then compare the candidate version with the current production baseline.
The lifecycle practices associated with AI-300 fit naturally here: version the deployable assets, automate evaluation, require thresholds, promote gradually, and preserve rollback. Source-control practices such as reviewable Git history help make prompt and configuration changes visible to the same governance process as code.
Do not use one overall score as the only gate. Require critical dimensions to pass independently. A candidate should not compensate for a security regression by becoming more eloquent, nor should a cost improvement compensate for worse policy adherence.
Red-team and abuse testing deserves a separate track from ordinary regression tests. Functional datasets ask whether the system can complete its intended work. Adversarial datasets ask whether it can be induced to exceed that work. Keeping both visible prevents teams from trading away security because average task quality remains high.
Document residual risk at release time. No evaluation proves that an agent will never fail. The release decision should state what was tested, which thresholds passed, what known limitations remain, what permissions the agent has, and what monitoring or human controls reduce the remaining exposure. That makes readiness a governed decision rather than an implied claim of perfection.
Release gates work best when they distinguish blocking regressions from review signals. A security-policy failure or unauthorized tool action may be an automatic stop, while a small latency increase may require a trade-off review. Define those rules before seeing the candidate release so teams do not move thresholds to make a favored version pass. Preserve the baseline, candidate scores, failed cases, and reviewer disposition with the release record; that creates an auditable explanation for why a version was promoted.
Use production feedback without turning users into test subjects
Pre-production evaluation cannot predict every real interaction. After release, sample traces, monitor quality metrics, capture explicit user feedback, and add newly discovered failure modes to the test set. This creates a feedback loop from production evidence back into controlled validation.
Rollouts should limit exposure while evidence accumulates. A new version might begin with an internal cohort, a small percentage of traffic, or read-only tasks before receiving broader action permissions. Define rollback triggers before release so the team does not debate acceptable degradation during an incident.
Within the Microsoft agent ecosystem, the goal of evaluation is not to prove that an AI system is generally “smart.” It is to show that a specific version reliably performs a specific business job within its allowed authority and risk tolerance. Representative datasets, explicit rubrics, trajectory checks, adversarial testing, operational measurements, and release gates give architects the evidence needed to make that decision before users depend on the agent.