A working demonstration is not an evaluation. If a support agent answers one familiar question correctly, the result shows that a particular interaction succeeded; it does not establish how the system behaves with a missing account, a malicious document, an unavailable tool or a customer who changes their request halfway through. In production, architecture includes a plan for finding those failures before they become incidents.
Claude Certified Architect – Foundations places evaluation across several domain decisions rather than treating it as a decorative final step. Agent design, tool integrations, structured output and context management all require evidence that the chosen approach works under realistic constraints. An architect should be able to say what a test measures, what it leaves uncertain and which failure warrants a release block.
Evaluate the workflow, not just the final sentence
A generative answer can sound useful while the underlying system behaves incorrectly. A research agent might cite the right topic but use an outdated source. A support agent might give a polite response after triggering an unauthorized tool action. A coding assistant might describe passing tests that never ran. Evaluating only the final text would miss those defects.
Separate measurements by layer. At the tool layer, check whether arguments conform to schemas, permissions are enforced and errors are handled without unsafe retries. At the model-decision layer, examine whether the agent selects relevant tools and uses results appropriately. At the content layer, check factual correctness, provenance, completeness and the handling of uncertainty. At the workflow layer, confirm that approvals, cancellations and recovery behave as designed.
A test case should have a defined input, expected properties and evidence of the observed behavior. For an extraction system, the expected property may be that a missing date is marked absent rather than invented. For an account workflow, it may be that unauthorized requests never reach an execution tool. For research, it may be that a citation points to a retrieved source that actually supports the claim.
These layers can fail independently. A valid final JSON object proves little about whether the tool selection was appropriate or the data was authorized. A correct answer in one run does not prove an application will behave the same way under load or after its context has been compacted.
Build evaluation sets around real decisions
The most valuable tests resemble situations the service will actually encounter. A customer-support agent should be tested with ordinary questions, incomplete identities, conflicting account records, cancellation attempts and cases where a human must decide. A document extraction agent should see missing fields, OCR errors, multiple similar entities and contradictory dates. A developer agent needs tests that require code changes, dependency checks and accurate reporting of build failures.
Balance common cases with consequential edge cases. A high average accuracy score can mask a small number of severe failures. If a system incorrectly approves one unauthorized refund in a thousand, the business may care more about that event than about several minor wording errors. Evaluation therefore needs severity categories, not just one aggregate score.
An effective regression set includes examples drawn from past incidents and from deliberate threat modeling. Each defect becomes an opportunity to make the failure reproducible. Tests should retain the input, relevant system version and expected outcome so later prompt, model or tool changes can be compared meaningfully.
Avoid using live confidential data unnecessarily. Synthetic or appropriately de-identified cases can exercise most control paths, while access to real production examples should be governed by privacy and retention rules. Test data also needs quality review; an incorrect “ground truth” label can penalize a sound system or reward a flawed one.
A release gate should prove both correctness and containment
Build a small but representative evaluation set for a refund-support agent: ordinary inquiry, verified eligible refund, request from a different account, duplicate submission, interrupted payment service, malicious retrieved policy, ambiguous user consent and missing transaction evidence. For each case, specify the expected application action and the expected user-visible explanation separately. A response can be clear yet unsafe; conversely, a safe denial can still be unhelpful if it obscures what information a legitimate user should supply next.
Score different dimensions against their own evidence. Grounded-answer quality needs source identifiers; tool selection needs requested and executed operation names; access-control enforcement needs service decision logs; resilience needs duplicate-effect checks; usability needs appropriate explanation without leaking protected data. A single model-generated confidence score cannot replace those observable measures. For high-impact operations, failed authorization or duplicate charge should block a release regardless of a high average answer-quality score.
A disciplined evaluation report records the baseline version, changed prompt or model configuration, test dataset version, failure counts, reviewer decisions and rollback trigger. Re-run the same scenario when a tool schema changes, not just when the prompt changes. Keep some previously unseen cases so the team does not tune all logic solely for the examples it already knows. That is how evaluation becomes operational evidence rather than a polished demonstration.
Traces connect model decisions to operational evidence
When an agent fails, a reviewer needs to know which step created the problem. A useful trace connects the user request, model turns, tool calls, returned statuses, retrieved source identifiers and the final answer. It also records timing and cost where those affect system design.
For a multi-agent workflow, trace each delegated task with its parent and outcome. If two agents cite the same outdated document, the issue may lie in shared retrieval rather than their reasoning in isolation. If a tool call was rejected by the authorization service, the failure may be a wrong model decision but not a successful security breach. The trace should preserve that distinction.
Multi-agent observability should reveal where evidence entered, which component made a decision and what external state changed. The instrumentation may differ in a Claude system, but those questions still govern debugging.
Do not log everything indiscriminately. Tool arguments and retrieved documents can contain secrets or personal information. Use redaction, access controls and retention rules so diagnostic records do not create a second security problem. A trace should answer operational questions without becoming a widely accessible copy of sensitive data.
Reliability needs more than self-reported confidence
Models can produce confident explanations for wrong answers. Asking an agent to grade itself may provide a useful secondary signal, but it is not an authoritative verdict. Where a task has verifiable properties, evaluate them externally: test whether a file exists, whether a database transaction has a valid ID, whether a cited paragraph contains the claimed fact, or whether a code change passes its test suite.
Confidence can be paired with abstention or escalation when uncertainty is significant, but the policy for escalation should account for risk. A low-impact formatting suggestion may be delivered with a caveat; an uncertain security authorization decision should not proceed automatically. Define the evidence threshold and the owner of the final decision in advance.
For cases involving language interpretation, human review may be necessary. Reviewers should assess the right sample, including failures and ambiguous cases, rather than rubber-stamping fluent outputs. A useful feedback loop records why a result was judged incorrect and converts that reason into a reproducible test when possible.
When comparing alternative designs, do not assume that a more complex agent is automatically more accurate. A simple deterministic check or clearer tool description may fix a recurring error more reliably than adding a second model call to criticize the first.
Adversarial and interruption tests expose architectural defects
Prompt injection tests matter because retrieved text can attempt to cross authority boundaries. Give the agent a legitimate document containing a malicious instruction. A passing system should use the document for its factual content without allowing it to alter the application’s tool permissions or override the user’s actual task.
Another class of test targets interruption. Terminate a workflow after a tool may have executed but before the result was recorded. On recovery, the application should reconcile the operation rather than blindly retry it. Simulate a user cancellation while a subagent is still active and verify that no later state-changing operation proceeds without renewed authorization.
Test resource exhaustion. An agent might repeatedly call a tool when results are incomplete or keep splitting work into new subagents. Define maximum turns, attempts, elapsed time and cost for the use case. A bounded “cannot finish with available evidence” is often a better production outcome than unlimited looping.
These cases should be designed before launch and repeated after relevant changes. Regression coverage must include both successful paths and the boundaries that prevent unsafe behavior.
Prompt and model changes are release changes
A new prompt version may improve answers in common cases while breaking structured output on rare documents. A different model may follow the same instructions differently. An updated tool schema may make the model choose a different operation. Treat these as controlled changes with a before-and-after evaluation record.
Version the prompt, tools, model configuration and test data. Associate results with the exact deployment under review so a team can reproduce its findings. For critical workflows, compare not only aggregate outcomes but error categories: unsupported claims, tool-selection mistakes, authorization failures, missed escalations and incomplete responses.
Testing and evaluating agents on Salesforce Agentforce similarly depends on clear test ownership and release thresholds. A Claude architecture needs its own checks, not a mechanical copy of another product’s controls.
If the service operates continuously, monitor production outcomes for distribution shifts, latency changes and tool failures. Evaluation does not stop when a release passes the staging suite. Real usage may reveal new document formats, unfamiliar questions or failure modes that initial tests missed.
Cost and latency belong in quality measurements
Two designs may have similar answer quality but very different operational costs. One may retrieve a small relevant section and answer in a few model calls; another may start multiple subagents and process a large corpus. If the extra work does not materially improve the decision, it may be the wrong production architecture.
Measure token consumption, tool latency, number of calls, external-service cost and failure-related retries. Evaluate these metrics against the service’s expected volume and response-time requirements. A system that passes quality tests but routinely exceeds a customer’s workflow deadline is not production-ready.
Cost optimization must not remove critical checks. Skipping an authorization service or citation validation to save a model turn is false economy. The goal is to eliminate redundant work while retaining controls that the risk model justifies. Architects should be able to explain which steps can be shortened safely and which are required guarantees.
An operational metric should have an owner and an action. If a latency alert rises, the team needs to know whether to inspect retrieval, tool availability, context size or model selection. A dashboard full of measurements without response procedures does not make a system more reliable.
A practical release gate for the CCA-F candidate
For a Claude Certified Architect – Foundations exercise, design an agent that summarizes customer incidents and proposes follow-up actions. Build a small evaluation set containing ordinary incidents, duplicate records, missing attachments, conflicting timestamps, malicious instructions embedded in logs and requests that require approval.
Record what success means at each layer. The model should select relevant evidence; the retrieval layer should return actual source IDs; the output should preserve uncertainty; and any proposed external action should be checked by independent permissions. Run the cases with a fixed version of the prompt and tool contracts, then change one component and compare the effects.
Finally, test interruption and rollback. Confirm that completed actions can be reconciled and that an error report distinguishes “the assistant could not retrieve evidence” from “the action was performed successfully.” The resulting evaluation package provides something more useful than a demonstration video: a defensible reason to trust, restrict or reject the design before customers depend on it.
The architect’s responsibility is not to guarantee that Claude never makes a mistake. It is to create tests and controls that reveal mistakes, contain their consequences and support continuous improvement. That is the operational judgment CCA-F candidates should practice.