INSIGHTS
AI & Data

Agent Loops and Delegation in the CCA-F Exam

See how bounded Claude agent loops, delegation, error handling and approval boundaries fit the Architect Foundations exam.

In this article
  1. The agent loop is a state machine, not an endless conversation
  2. Choose deterministic steps where determinism is possible
  3. Delegation should change the information flow
  4. Failure handling belongs outside optimistic prompt wording
  5. Inspect the actual tool-use protocol before deciding to retry
  6. Human approval is a designed branch
  7. Parallel agents introduce race conditions and costs
  8. Test the boundaries with failure scenarios

The most expensive agent architecture mistake often appears before the first tool call: a team gives a language model control over a workflow that would have been safer as ordinary application code. An agent can be valuable when it must inspect uncertain evidence, choose among unfamiliar sources or adapt its next step. That does not mean it should decide whether an account holder is authorized to transfer funds, whether a refund has cleared, or whether a production deployment is permitted.

Anthropic’s Claude Certified Architect – Foundations exam, formally coded CCAR-F and commonly called CCA-F, tests the distinction through production scenarios. The practical question is not how to get an agent to do more work. It is how to let it make useful decisions without turning unpredictable output into uncontrolled action.

The agent loop is a state machine, not an endless conversation

A tool-using Claude application typically sends a set of instructions and context to the model, receives either a final response or a request to use a tool, runs an allowed tool through the application, and returns the tool result for a further model decision. The application owns the loop. It decides which tools exist, how a tool call is authenticated, what happens on failure, and when the workflow must stop.

A successful turn is not identical to the absence of a tool call. The Claude API includes response stop reasons that distinguish a natural completion, a tool request and other conditions such as output limits. Code that merely assumes “no exception means finished” can silently accept partial work. A robust controller interprets the response structure and records whether it has a user-ready result, a tool request, or a state that needs continuation or review.

The controller should also own budgets. Limits might include total elapsed time, number of model turns, number of tool attempts, money spent, and amount of data retrieved. These do not have to be identical across tasks. A document-classification workflow with a short deadline should not use the same loop budget as a research task that may consult twenty sources.

Suppose a model repeatedly requests a search that returns empty results. The correct behavior may be to ask for a missing identifier, switch to an alternate search strategy or return an explicit “not found.” Continuing the same search indefinitely is not persistence; it is a missing termination condition.

Choose deterministic steps where determinism is possible

Imagine an insurance claims service. The model can read a narrative, identify potentially missing information and suggest which policy section needs attention. Coverage thresholds, customer eligibility and disbursement controls should come from authoritative rules and systems. Treating the model’s prose as the definitive source of those decisions produces a design that is hard to test and harder to audit.

The boundary can be explicit: Claude proposes review_claim(claim_id, concerns); an application validates the claim ID, checks the caller’s permissions, fetches required records, and either schedules a human review or returns a typed rejection. This is better than a tool that accepts arbitrary SQL or can modify every claim field. Even a well-instructed model may construct a dangerous argument when source data is misleading or its context is incomplete.

Deterministic checks are especially important for invariant properties. A support agent may confidently state that it has “never refunded more than the allowed amount.” That statement cannot replace a transaction service whose code enforces maximum amounts. In a production design, natural-language judgment informs bounded choices while the application enforces business rules.

Agentic application architecture on Amazon Bedrock faces a comparable separation of model-chosen steps from platform-enforced controls, even though the underlying services differ.

Delegation should change the information flow

A subagent is useful when it receives a sharply defined assignment, a suitable tool set and a result contract that lets the parent continue. For example, a legal research assistant may delegate contract-clause extraction to one specialist and jurisdictional background research to another. Each worker can operate on a narrower context and return evidence with citations. The parent still needs to reconcile inconsistent terminology and explain uncertainty.

Delegation does not automatically improve quality. If the parent forwards the entire original request to several indistinguishable subagents, the system may simply repeat work and produce mutually reinforcing errors. Coordination overhead also grows: results have to be awaited, validated, combined and sometimes retried. Before adding another agent, ask whether the task truly benefits from separate context, independent tools, parallelism or specialized instructions.

Define a handoff contract. A research worker might return a finding, source identifiers, confidence-limiting observations and unresolved questions. The parent should not silently treat a missing citation as a successful source check. If a worker reports a transport error, the parent can retry under a budget, choose another source or disclose the limitation. A vague “subagent completed successfully” is not enough evidence for a consequential answer.

Multi-agent observability requires each handoff to have an identifiable owner, a traceable outcome and enough context to distinguish a tool failure from a model or orchestration mistake.

Failure handling belongs outside optimistic prompt wording

Three different failures can look similar to a user: a tool may time out; a tool may return a valid but empty response; or it may return an apparently valid result containing stale or false data. An architecture must respond to each differently. Timeouts can justify bounded retries with backoff. Empty results may require user clarification. Stale information needs provenance, freshness checks or an alternate authoritative source.

Retrying an irreversible action is particularly dangerous. A payment tool that times out might still have completed the payment. The agent must not issue another charge merely because it did not receive a success message. Instead, the transaction service should use idempotency keys or durable operation identifiers and allow the application to reconcile the original request before retrying.

Tool errors also need structured responses. A single string such as “failed” leaves the agent to guess whether the problem is authentication, authorization, invalid input, temporary unavailability or a permanent conflict. A typed error with a stable code can guide controlled handling. Sensitive diagnostic details should stay in operational logs rather than being exposed unnecessarily in the model’s context.

A production controller should maintain an explicit state: pending approval, awaiting tool, completed, failed, timed out, or cancelled. This helps recover from interruptions without reconstructing the truth from conversational text. A model’s narrative about what it thinks happened is not a transaction log.

Inspect the actual tool-use protocol before deciding to retry

For the Claude Messages API, a response with stop_reason equal to tool_use is not a final business outcome. The model has emitted a tool_use block with an identifier, name and arguments; the application must validate the request, execute the authorized client-side operation and return a tool_result matched to that identifier. The agent may then continue. An ordinary natural-language answer can arrive with end_turn; a response truncated at the configured output limit can arrive with max_tokens. These states must not be collapsed into one generic “Claude completed” flag.

Suppose the model requests a refund proposal. The tool service commits a proposal but the network response times out. Repeating the tool call immediately could create a duplicate proposal or, in a poorly designed system, a duplicate payment. The controller should first query the authoritative ledger using its idempotency key or operation identifier. A missing tool_result is a protocol problem; a missing ledger record is a different business-state problem. The retry policy should depend on which evidence is actually absent.

Anthropic also documents pause_turn for a paused server-side tool sequence. Its continuation is not the same as returning a client-tool result. Architects need to distinguish tools executed by the application’s own code from server tools run by the provider, and preserve sufficient history for the correct continuation. This level of protocol awareness prevents a subtle but costly design error: telling operators an agent succeeded because the model stopped, when the action was merely requested or suspended.

Human approval is a designed branch

A human-in-the-loop process should identify the precise event that requires review. It might be an attempted transfer above a threshold, publication of content to a customer-facing site, access to sensitive information, or an update that cannot be reversed. The application pauses the action, presents the facts needed for a decision and records the reviewer’s choice in an authoritative system.

An approval checkpoint is not the same as asking the model whether it feels confident. Self-rated confidence can be useful as a weak triage signal but is not a substitute for risk rules. For a high-impact operation, the permission boundary must be enforced by software the agent cannot override. The reviewer should see the proposed operation, target, relevant evidence and potential consequences, not a decorative “approve” button detached from the real tool arguments.

A well-placed approval gate also improves usability. Requiring confirmation for every harmless query creates fatigue and delays. Risk-based approval reserves human attention for decisions involving money, access rights, irreversible state or significant uncertainty. The balance is a design choice that should be reviewed against the organization’s actual risk tolerance.

Human oversight of AI-agent workflows works only when the approval has an enforced effect on subsequent actions; a Claude application must implement that boundary through its own authorization services.

Parallel agents introduce race conditions and costs

A team may divide a research problem among several agents to reduce elapsed time. Parallel work succeeds only when output is independently bounded and combined safely. If two workers can update the same record, concurrency can lead to lost changes or conflicting state. If each worker can discover more tasks without a budget, cost can increase faster than the parent can supervise it.

Prefer read-only delegation where possible. When parallel workers must write, specify ownership of each resource, concurrency controls, idempotent operations and conflict handling. A parent agent should reconcile findings from separate workers before committing a final recommendation; agreement between agents trained similarly is not proof of factual independence.

Think about context costs as well. Passing all preceding messages to every worker increases token usage and may expose information irrelevant to the worker’s assignment. A smaller task brief with a limited evidence bundle can improve focus and reduce unnecessary data sharing. The parent then combines explicit outputs instead of accumulating unstructured chat history.

An asynchronous workflow also needs a durable place to store results. A network process or browser session cannot be the only record of a task that may continue after a worker restarts. Without persistent identifiers, a partial retry may duplicate the work or lose the fact that a step was already approved.

Test the boundaries with failure scenarios

The best rehearsal for the CCA-F exam is an architecture review in which something goes wrong. Give a support agent a tool response that contradicts the customer’s message. Give a research agent three sources with incompatible dates. Interrupt a delegated task halfway through. Simulate a timeout after a successful state-changing operation. Ask where the application learns the truth and who has authority to proceed.

A useful test record contains the requested operation, tool inputs after validation, tool result, decision taken, and the reason that decision was permitted. For content-producing agents, retain source references so reviewers can distinguish grounded findings from unsupported synthesis. For sensitive systems, logs themselves need restricted access, retention policies and redaction.

The final design should not promise that the model will always pick the right action. It should make wrong actions harder to execute, easier to detect and less damaging to recover from. An architect understands when a free-form agent loop is warranted, when a deterministic workflow is enough and when a human must assume control. That is the practical foundation behind CCA-F’s largest exam domain.

Filed under AI & Data