{"id":3925,"date":"2026-10-11T08:37:54","date_gmt":"2026-10-11T08:37:54","guid":{"rendered":"https:\/\/www.examtopics.info\/blog\/claude-agent-sdk-hooks-sessions-cca-f\/"},"modified":"2026-10-11T17:25:03","modified_gmt":"2026-10-11T17:25:03","slug":"claude-agent-sdk-hooks-sessions-cca-f","status":"publish","type":"post","link":"https:\/\/www.examtopics.info\/blog\/claude-agent-sdk-hooks-sessions-cca-f\/","title":{"rendered":"Claude Agent SDK: Hooks and Sessions for CCA-F"},"content":{"rendered":"<p>A Claude agent can answer a question convincingly while its surrounding application is in a state that should never have been allowed. A support assistant might promise a refund before the payment system confirms it; a research agent might complete a summary after one of three sources timed out; a coding agent might state that a test suite passed while its test process was still running. None of those failures is solved by a better adjective in the system prompt. They are failures of control flow, state and evidence.<\/p>\n<p>The Claude Certified Architect \u2013 Foundations examination, coded CCAR-F in Anthropic&#8217;s published guide and listed as <a href=\"https:\/\/www.examtopics.info\/cca-f\">CCA-F<\/a> on this site, asks architects to recognize these boundaries. The <a href=\"https:\/\/www.examtopics.info\/blog\/cca-f-agent-loops-delegation\/\">agent-loop design<\/a> question is not merely how to make the model call a function. It is how the application knows what the model requested, what really happened, what may happen next and what it is allowed to report to the user.<\/p>\n<h2>A tool call is a control signal, not ordinary prose<\/h2>\n<p>Imagine a customer asking an agent to check a parcel and change its delivery address. A fragile implementation tells Claude, \u201cWhen you are finished, say DONE,\u201d and searches the generated response for that word. It can be fooled by a quoted example, a partial answer or a tool result that happens to contain the same letters. It also conflates a user&#8217;s requested outcome with the application&#8217;s actual state.<\/p>\n<p>The Messages API exposes explicit content blocks and a <code>stop_reason<\/code>. A client-executed tool request appears as a <code>tool_use<\/code> content block, and the response signals that the application should execute tools. The application locates each requested tool, validates the arguments, runs the operation under its own authorization policy, and returns a matching <code>tool_result<\/code> message. The next model turn must receive enough conversation and tool state to interpret that result. The implementation, not a text marker, determines whether to continue.<\/p>\n<p>That distinction becomes important when the assistant requests two read-only tools in the same response. Both results need to be associated with the correct tool-use IDs. Returning only one or swapping results can produce a polished conclusion assembled from the wrong evidence. For state-changing tools, an application may also enforce ordering that the language model did not explicitly request. A shipping address must be validated against the account before a change is committed, even if the model would prefer to issue both calls together.<\/p>\n<p>An <code>end_turn<\/code> response can indicate that the model has finished its current turn; it does not prove the requested business operation succeeded. Conversely, <code>max_tokens<\/code> means generation reached its output limit and should not be treated as a successful final outcome. Server-executed tools can have additional continuation states. Architects should distinguish the stop-reason contract implemented by their chosen API and SDK version from assumptions inherited from an older code example.<\/p>\n<h3>A minimal failure-state contract for an Agent SDK controller<\/h3>\n<p>Treat model output and application operations as separate state machines. A tool call identifies the action Claude wishes to request; a hook can inspect the proposed operation or add an application-defined checkpoint at a documented lifecycle boundary. The protected service still owns authorization and the authoritative operation result. For example, a proposed <code>update_ticket<\/code> call should carry a validated ticket ID, requested change and trace ID. A denial must be reported as a denial rather than removed from the trace and silently retried with different arguments.<\/p>\n<p>A controller built directly on the Messages API should distinguish <code>tool_use<\/code>, <code>end_turn<\/code>, <code>max_tokens<\/code>, a refusal and other documented stop states. After <code>tool_use<\/code>, a <strong>client-executed<\/strong> tool produces a <code>tool_result<\/code> block whose ID corresponds to the model&#8217;s request. The controller must not fabricate a success block if execution failed. A provider-executed <strong>server tool<\/strong> may instead complete inside the provider&#8217;s own loop; a <code>pause_turn<\/code> can mean that sequence must be continued with the prior assistant content. These are different continuation contracts. Anthropic\u2019s published <a href=\"https:\/\/platform.claude.com\/docs\/en\/build-with-claude\/handling-stop-reasons\">Claude stop-reason semantics<\/a> distinguish these protocol states and should be consulted when implementing the controller.<\/p>\n<p>A useful architecture test artificially loses a tool response after a backend side effect. The controller should check the operation ledger, avoid duplicate effects, reattach the real receipt if permitted and preserve enough session state to explain the outcome. A resumable session is not a permission to replay earlier actions. Tests should also assert that a tool failure, absent approval and truncated model response yield different final statuses. These distinctions are the difference between a conversation that sounds complete and an application whose actions can be trusted.<\/p>\n<h2>Separate model deliberation from application state<\/h2>\n<p>A model might ask for <code>update_delivery_address<\/code> after receiving a validated customer identifier. That request is not itself a delivery-address update. The application must capture the verified target account, caller identity, proposed new address, approval requirement and server response. Only a successful operation receipt should allow the conversation to say the change was made.<\/p>\n<p>This creates two timelines. The conversational timeline includes assistant messages, tool requests and explanations. The operational timeline includes transaction IDs, execution timestamps, access checks and the actual committed result. Reconstructing either solely from the other is unsafe. A customer can receive a coherent message even when a payment service returns an unknown result. The reverse is also possible: a transaction succeeds while the model&#8217;s next response is interrupted.<\/p>\n<p>For an order-change task, keep a durable operation record outside the prompt. A useful record might contain <code>request_id<\/code>, <code>account_id<\/code>, <code>expected_version<\/code>, <code>approval_state<\/code>, <code>last_tool_status<\/code> and <code>committed_at<\/code>. A retry can consult that record before attempting another write. Nothing about this requires Claude to keep every raw API response in its context. The model needs the fields required for the next decision, and the application needs the full evidence for audit and recovery.<\/p>\n<p>The boundary becomes even clearer in a multi-user system. A system prompt telling the agent to act only on the current customer&#8217;s order is insufficient. An ordinary authorization layer must establish the customer-to-order relationship and deny other resources. The <a href=\"https:\/\/www.examtopics.info\/blog\/mcp-tool-design-anthropic-ccar-f\/\">MCP tool-contract design<\/a> is useful because it exposes this separation between a convenient tool interface and real execution authority.<\/p>\n<h2>Hooks enforce gates at the moment they matter<\/h2>\n<p>Hooks are attractive because they provide a place for deterministic checks around model-driven work. The architect&#8217;s first question should be whether a policy must be enforced before a tool runs or whether a check is meant to inspect and normalize the result afterward. Those are different jobs.<\/p>\n<p>Suppose a code agent may edit files under <code>src\/<\/code> but must not modify production secrets or deployment credentials. The protected-file rule must be enforced when the proposed operation is intercepted, before a write occurs. A later review that reports \u201ca forbidden file was modified\u201d is useful evidence of a failure, but it is not prevention. Similarly, a high-value payment should be blocked pending approval before an irreversible API call; asking a model to apologize after the call cannot restore a lost payment.<\/p>\n<p>A post-tool hook can help normalize noisy observations. A successful test runner may return hundreds of lines of logs, a process exit code and a concise test summary. A controlled post-tool step can preserve the exit code and failing test identifiers while reducing irrelevant console output before it becomes the next model context. The key is not to falsify failure into success. Normalization improves the quality of evidence without changing the fact that a command failed.<\/p>\n<p>Not every convention needs a hook. It is reasonable to describe preferred naming in <a href=\"https:\/\/www.examtopics.info\/blog\/claude-md-rules-skills-cca-f\/\"><code>CLAUDE.md<\/code><\/a> and verify it through a linter or review. It is not reasonable to rely on preferences alone for forbidden production actions. An effective design allocates responsibilities: prompts explain goals, application code enforces authorization, hooks enforce local workflow gates where supported, and independent tests decide whether the delivered artifact is acceptable.<\/p>\n<h2>Give subagents bounded work and traceable returns<\/h2>\n<p>Delegation can be valuable when a coordinator needs distinct tool access or a separate working context. A research agent might ask a document specialist to extract evidence from an annual report while another specialist compares requirements. The coordinator should not have to read every raw paragraph that either specialist encountered, but it must receive enough to assess coverage, contradictions and uncertainty.<\/p>\n<p>A useful delegation contract includes task scope, input references, allowed tools, time or iteration budget, expected output fields and a clear failure path. A subagent return can report <code>complete<\/code>, <code>partial<\/code>, or <code>blocked<\/code>, along with relevant source identifiers and which questions remain unanswered. This prevents a coordinator from treating an empty result as proof that nothing exists. It also avoids a different failure: confidently asserting complete coverage from a specialist that silently lost access to one source.<\/p>\n<p>Passing only a narrative summary between agents is especially risky for a research workflow. A statement such as \u201cthe vendor says its product is secure\u201d is hard to challenge without a source and version. A more useful handoff identifies the precise vendor document, date, relevant claim, scope and any conflicting evidence. The next agent can then synthesize instead of inventing where a claim came from.<\/p>\n<p>When several agents execute in parallel, result arrival order should not determine truth. A coordinator can merge records by stable source ID, detect duplicates and flag conflicts. It may request a targeted follow-up if the evidence is incomplete. It should not launch new subagents merely because the architecture diagram looks more sophisticated with additional boxes.<\/p>\n<h2>Resume and fork sessions without inventing continuity<\/h2>\n<p>Long-lived tasks are vulnerable to interruption. A research run may stop after gathering documents but before writing the report. A coding task may stop after generating a patch but before testing it. Session resume mechanisms can help continue context, yet a resumed conversation is not automatically a restored operational state.<\/p>\n<p>A reliable resume record states which work products exist, their versions and hashes where useful, which tools were executed, which writes are confirmed and what still needs verification. An interrupted deployment that shows \u201cpending\u201d should not be summarized as deployed. A subsequent run must check the actual destination. This is the same readback principle that matters when a tool times out after performing a write.<\/p>\n<p>Forking a session is another distinct operation. A fork lets investigators pursue alternative hypotheses or design options from a shared starting context, without treating the alternatives as one continuous decision history. This is useful when two debugging paths would otherwise contaminate each other&#8217;s assumptions. Results should be compared through independently observable evidence, such as test outcomes or measured behavior, rather than by asking the forks to vote about which explanation sounds better.<\/p>\n<p>Context carried into a resumed or forked session should be selective. Large duplicated tool outputs can crowd out the constraints that matter most. The article on <a href=\"https:\/\/www.examtopics.info\/blog\/claude-cca-f-context-recovery\/\">context recovery in Claude agents<\/a> examines how to preserve provenance and checkpoints without requiring the model to memorize the entire execution history.<\/p>\n<h2>Design retries around effects, not hopeful messages<\/h2>\n<p>Tools fail in materially different ways. A transient connection error may justify a bounded retry; an authorization failure generally does not; an invalid customer identifier requires corrected input; an ambiguous timeout after a payment call requires reconciliation before repeating the operation. Treating every error as \u201ctry again\u201d can produce duplicate actions or hide a policy violation.<\/p>\n<p>An idempotency key helps a state-changing service recognize a repeated request as the same intended operation. The guarantee still depends on the backing service honoring that key and retaining its result for a documented window. An agent that reuses a key for a different action is not protected. Engineers should also distinguish what the tool reported from what the server ultimately committed: \u201crequest timed out\u201d is not synonymous with \u201ctransaction failed.\u201d<\/p>\n<p>Retry budgets should be small enough that an unavailable service does not trap the conversation indefinitely. Exceeding a budget leads to a structured partial outcome or a <a href=\"https:\/\/www.examtopics.info\/blog\/human-escalation-provenance-claude-cca-f\/\">human escalation<\/a>, not a fabricated success. These rules apply even when Claude could theoretically generate another plan. The user asked for an outcome, not an endless demonstration of agent autonomy.<\/p>\n<h2>Test protocol and recovery failures deliberately<\/h2>\n<p>Build a small test harness for the control flow. One test returns two parallel tool calls with results deliberately delivered in reverse order. A second returns an authorization error for a protected operation. A third reports a successful write but interrupts the model before it acknowledges the receipt. A fourth returns a valid empty search result; a fifth simulates a failed search and checks that the two are not confused.<\/p>\n<p>Record the expected application behavior before running the tests. A correct implementation matches tool results to call IDs, refuses disallowed actions before execution, preserves operation receipts, distinguishes terminal model responses from business success and resumes from durable state. Reviewers should be able to reconstruct a failure without relying on the model to explain its own intentions after the fact.<\/p>\n<p>In an exam scenario, the best answer often introduces a clear boundary: a stop reason instead of parsing ordinary text, a pre-execution gate instead of a warning, a structured subagent return instead of an untraceable summary, or a durable record instead of assumed continuity. The architecture becomes trustworthy when its control flow can be tested by software and reviewed by people independently of Claude&#8217;s fluency.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Understand Claude Agent SDK stop reasons, tool results, policy hooks and resumable sessions through practical architecture decisions.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12,8],"tags":[],"class_list":["post-3925","post","type-post","status-publish","format-standard","hentry","category-ai-data","category-certifications"],"_links":{"self":[{"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/posts\/3925","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/comments?post=3925"}],"version-history":[{"count":2,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/posts\/3925\/revisions"}],"predecessor-version":[{"id":3958,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/posts\/3925\/revisions\/3958"}],"wp:attachment":[{"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/media?parent=3925"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/categories?post=3925"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.examtopics.info\/blog\/wp-json\/wp\/v2\/tags?post=3925"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}