Agent testing is different from testing a deterministic workflow because the same model-driven system can produce variation across runs. A production-quality Agentforce agent therefore needs a test set broad enough to evaluate routing, action selection, grounded knowledge retrieval, response quality, policy adherence, and multi-turn behavior. Salesforce’s Testing Center is designed for exactly this lifecycle, supporting generated and custom test cases, built-in scorers, and large-scale batch evaluation.
The current Salesforce Certified Agentforce Specialist guide includes Testing Center, evaluations, sandbox-to-production deployment, governance, and observability. Testing should therefore be treated as release engineering rather than as a one-time demo before launch.
Define expected behavior before scoring
Write down what success means for each use case. A service agent may need accurate answers, correct subagent routing, safe action execution, acceptable latency, and escalation when information is insufficient. Different requirements need different tests.
Do not collapse everything into one “quality” number. A high average response score can hide a severe action-selection failure.
Test subagent routing explicitly
Salesforce’s current terminology uses subagents where older documentation used topics. Testing Center includes subagent assertions so teams can check whether the agent routed an utterance to the intended capability.
Use near-boundary utterances, ambiguous wording, and overlapping intents. A routing design that works only for obvious examples will fail in real conversations.
Test action selection and action inputs
Action assertions help verify that the agent chose the intended action. Go further by validating input values, authorization context, and downstream effects where possible. A correct action called with the wrong record ID is still a failure.
Write actions so they are independently testable. Narrow inputs and outputs make it easier to identify whether a failure belongs to reasoning, parameter extraction, action code, or the downstream system.
Evaluate response accuracy and instruction adherence
Testing Center provides response-accuracy evaluation and additional quality scorers such as completeness, coherence, conciseness, latency, and instruction adherence. Select scorers based on the application’s risks rather than enabling every metric without purpose.
Instruction adherence matters when the agent has strict policies, response formats, or escalation requirements. A useful-sounding answer that violates a mandatory instruction should fail the test.
Use custom scorers for business outcomes
Salesforce supports custom scorers for organization-specific criteria such as tone, brand voice, policy compliance, resolution quality, or another business outcome. These are valuable when generic LLM quality metrics do not capture the actual acceptance standard.
Write scorer criteria clearly enough that a human reviewer can understand why a result passed or failed. Opaque evaluation logic is difficult to maintain and easy to game unintentionally.
Build representative scenario sets
Testing Center can generate test cases from subagents, actions, and connected data libraries, but custom scenarios are still necessary for edge cases. Include normal requests, ambiguous requests, missing-data cases, authorization boundaries, prompt injection, stale knowledge, and malicious inputs.
The principles in AI risk awareness are relevant because reliable testing needs to cover failures that ordinary happy-path conversations rarely expose.
Test grounding separately from generation
If a response is wrong, determine whether the correct knowledge was retrieved before blaming the model. Test cases should verify expected sources or facts where practical, especially for data libraries and RAG-style grounding.
Keep retrieval failures visible. A generation scorer may judge a plausible answer positively even when it relied on the wrong source.
Test in sandbox when actions can change data
Salesforce explicitly warns that Testing Center can modify CRM data and recommends sandbox use. Build representative test records and isolate destructive actions so batch testing cannot accidentally alter production customers.
Sandbox data should still reflect important permission, sharing, and relationship patterns. A completely unrealistic dataset can hide security and routing problems that appear only in production-like contexts.
Version tests with the agent release
Prompt changes, action changes, data-library changes, model changes, and subagent changes can all alter behavior. Store the test set and expected assertions with the release process so a candidate agent is evaluated against a stable baseline.
The broader discipline of version-controlled development supports reproducibility even if Testing Center manages the execution itself.
Production feedback should continuously expand regression coverage. After launch, add confirmed failures to the test suite. Repeated escalation, incorrect actions, retrieval misses, or policy violations should become regression cases so the same behavior does not return after a later update.
The Salesforce Agentforce lifecycle becomes safer when evaluation is continuous: define expected behavior, test at scale, inspect failures, refine subagents and actions, deploy deliberately, and turn production incidents into future tests.
Test datasets should be versioned and labeled by purpose. Maintain a small release-gating suite for must-pass behavior, a broader quality suite for representative coverage, and targeted adversarial suites for security and policy. Mixing all cases into one undifferentiated score makes release decisions harder.
Ground truth can be an expected answer, expected action, expected subagent, required fact, prohibited behavior, or business outcome. Choose the assertion type that matches the behavior being tested. Not every conversational response has one exact acceptable sentence.
Multi-turn testing is important because later agent behavior depends on conversation history. Include scenarios where the user changes intent, corrects earlier information, refers to a previous record, or tries to use prior context to bypass a policy. Single-turn tests can miss these state-dependent failures.
Non-determinism means one passing run is weak evidence for critical behavior. Repeat high-risk cases or test enough variants to understand consistency. A rare but severe wrong action should not be hidden by many easy successful conversations.
Subagent assertions should include negative cases. Test utterances that must not route to a sensitive subagent even when they contain related keywords. This helps detect overly broad subagent descriptions that capture requests outside their intended authority.
Action testing should validate side effects after execution. If an agent updates a Salesforce record, inspect the record state and audit fields rather than trusting the model’s statement that the update succeeded. The final response and actual system effect are separate test targets.
Testing Center can generate scenarios from connected data libraries, which is useful for grounding coverage. Add curated edge cases where the correct answer depends on a specific passage, an exception, or a version distinction so retrieval quality is exercised beyond generic questions.
Latency scoring should be interpreted alongside action and retrieval complexity. A slower response may be justified when it performs a necessary multi-step workflow, while a simple knowledge answer should not spend several seconds making unnecessary calls. Segment latency expectations by use case.
Completeness and conciseness can conflict. A service agent may need to include mandatory policy language even when the most concise answer would omit it. Custom scorers can encode those business constraints more accurately than optimizing every built-in quality metric independently.
Human review should calibrate automated scorers. Periodically compare evaluator judgments with domain experts, especially for policy compliance and nuanced business outcomes. If the scoring model disagrees consistently with expert standards, adjust the rubric rather than optimizing the agent toward the wrong target.
Tests should capture permissions. Run the same scenario as different user types where access changes expected behavior. A support supervisor and a customer may ask identical questions but should receive different data or action options.
Deployment gates should distinguish blocking and non-blocking failures. Wrong record updates, policy violations, and unauthorized data exposure should block release; minor style regressions may create follow-up work without stopping deployment. Explicit severity prevents release decisions from becoming subjective.
Production observation should feed back into evaluation. Agent analytics, escalations, user feedback, failed actions, and incident traces reveal scenarios that development teams did not predict. Convert validated failures into permanent regression cases with ownership.
Benchmark maintenance is itself a lifecycle. Remove obsolete policy cases, update examples when Data 360 or Agentforce terminology changes, and record why expected behavior changed. Otherwise a test suite can prevent correct modernization by enforcing historical assumptions.
Good evaluation changes the conversation from “the agent seems better” to “the candidate routes correctly, chooses safe actions, retrieves authoritative evidence, meets quality thresholds, and does not regress critical scenarios.” Testing Center provides the execution framework, but disciplined scenario design makes the results meaningful.
Test coverage should be mapped to the Agentforce architecture. Every subagent should have positive and negative routing cases; every write action should have authorization and idempotency tests; every grounding source should have freshness and permission cases; and every mandatory policy should have instruction-adherence checks.
Scenario generation can accelerate coverage, but generated tests should be reviewed for business realism. An LLM may produce syntactically plausible questions that no real customer asks while missing important contractual or operational edge cases known to domain experts.
Test data should be safe to mutate. Create dedicated records for actions that update CRM and reset or recreate them between runs. Shared test records can make results depend on execution order, which turns the evaluation suite itself into a source of nondeterminism.
Expected action sequences can matter for multi-step workflows. An agent that checks entitlement after issuing a refund is unsafe even if the final state eventually looks correct. Trace assertions should verify the order of operations when sequencing is a policy requirement.
Evaluation should distinguish recoverable conversation mistakes from irreversible system mistakes. A slightly verbose answer can be corrected in the next turn; an unauthorized record update cannot. Weight severity accordingly when deciding release gates.
Performance testing should include batch scale. Testing Center can run many scenarios, which is useful for finding patterns that manual testing misses. Monitor evaluation duration and cost, but do not reduce coverage solely to make the dashboard finish faster.
Locale and language testing matter when agents serve multilingual users. Translation can change intent classification, retrieval terms, and policy wording. Use native-language scenarios for important markets rather than assuming English test results transfer automatically.
Model upgrades require re-baselining some scorer behavior. If the evaluator or agent model changes, small differences in quality scores may reflect the scoring system as well as the agent. Preserve benchmark versions and inspect example-level changes.
Agent testing should also include cancellation, interruption, and correction. Users may say “stop,” change the requested account, or correct a date after an action plan has begun. The agent should respond safely to changed intent and not execute stale steps.
Production analytics can prioritize which scenarios deserve more test coverage. High-volume subagents, frequently failing actions, common escalations, and high-latency paths should receive more attention than rarely used capabilities.
A testing program matures when every significant production change has a corresponding evaluation plan and every meaningful production failure becomes a regression test. The suite then becomes institutional memory for agent behavior rather than a one-time launch checklist.
Security and quality tests should share some scenarios. A prompt-injection attempt can be evaluated for both refusal behavior and whether the agent avoided calling a sensitive action. Testing only the final text can miss a dangerous tool call that happened before the refusal.
Conversation-level scorers are important for workflows where correctness depends on several turns. The agent may eventually reach the right outcome but ask unnecessary questions, lose context, or reveal information during the path. Evaluate the session, not only the last message.
Release reports should summarize which tests changed from pass to fail and which improved. Example-level diffs are more actionable than one average score because they show exactly which use cases were affected by the candidate release.
Evaluation ownership should be shared with domain experts. Technical teams can maintain Testing Center and scorers, while business owners define correct policy outcomes and acceptable escalation. This prevents benchmarks from drifting away from real operational expectations.
Test maintenance should include removing obsolete action names and subagent labels after architecture changes. Stale assertions can create false failures that encourage teams to disable useful controls rather than update the benchmark.
Use holdout scenarios that builders do not tune against directly. If every failing example is repeatedly used during prompt editing, the agent can improve on that fixed set without becoming more robust. A separate holdout set gives a better signal of generalization.
Testing should include tool unavailability and partial failures. The agent must behave safely when a Flow errors, an API times out, or a data library returns no result. Resilience is part of quality, not merely an infrastructure concern.
For actions with side effects, compare the expected and actual record state after each test and reset the sandbox fixture. This creates a deterministic test harness around a nondeterministic reasoning layer.
Quality thresholds should be reviewed as the use case matures. A pilot may tolerate more manual escalation, while a production customer channel may require stricter accuracy and latency. Evaluation criteria should evolve with business risk.
Release evidence should include the agent version, test-suite version, model, data-library state, scorers, pass/fail summary, and known accepted exceptions. This creates an auditable reason for promotion and a baseline for the next regression comparison.
A mature test program should also track coverage by subagent, action, policy, language, and risk tier so teams can see where confidence is strong and where important behavior is still represented by only a few examples.
That coverage map turns testing from a collection of examples into a measurable quality system that can evolve with the agent.
Keep the suite aligned with the risks the production agent actually carries.
Keep evaluation tied to production risk.
Regression suites should retain confirmed production failures as future test cases, ensuring that a prompt, topic, action, or model update does not quietly reintroduce behavior the team already corrected.