A model can tell an application which tool it would like to call, but that is only a proposal. The application or managed agent runtime still needs to route the call, supply credentials appropriately, validate arguments, execute the operation and return a result. This distinction is the core of safe agent development. When the model is treated as both planner and final authorization authority, a small misunderstanding of a user request can become a real-world write operation.
The AI-103 exam asks candidates to build agents with roles, goals, conversation state and tools. Microsoft Foundry Agent Service supplies managed runtime capabilities, yet developers must design the business boundary around those capabilities. A useful example is a facilities assistant that can find equipment documents, check a maintenance schedule and request a technician visit. Those actions have different risk levels and should not be collapsed into one all-purpose tool.
Choose an agent model based on operational responsibility
A configured prompt agent may be suitable when the main work is conversation plus supported managed tools. A hosted or custom-code agent may be appropriate when the team needs bespoke orchestration, integrations or state handling. The choice affects how much runtime infrastructure and code the application owns. It should follow real deployment and security needs rather than a preference for the most elaborate architecture.
Microsoft’s Foundry Agent Service overview describes prompt and hosted-agent patterns, model connectivity, tools, identity and observability. Feature support evolves by model, region and agent type. Verify the current documentation for the exact combination in a design; do not copy a preview feature from an old example and describe it as universally supported.
Define the work as bounded capabilities
A maintenance assistant might expose three separate operations: find_manual, get_service_window, and request_visit. Each returns a different type of evidence or consequence. A read-only query should not be able to mutate a schedule. A request for a visit should contain a validated asset, location, time window and requesting user identity. The downstream API should enforce those constraints even if the agent submits syntactically valid parameters.
Give tools descriptive names and argument schemas that reflect their real business behavior. Avoid a generic “run operation” function with free-text instructions. Keep side effects explicit: whether the call only reads, creates a draft or commits an external change. The model’s natural-language reasoning is useful for selecting a capability; deterministic code remains responsible for enforcing its contract.
Handle the tool-call loop explicitly
A tool-using exchange can have several steps. The user asks a question, the model proposes a tool call, the runtime executes it, the result is returned, and the model generates a response or another tool request. In a loop, the system must know when to stop. A failed search should not cause the agent to issue the same expensive query indefinitely; a completed write should not be repeated because a network acknowledgement was lost.
Set bounds on tool calls, total execution time, retries and cumulative cost. Use idempotency keys for actions that can be repeated by transport or agent retries. Include a clear tool result that distinguishes “no matching record,” “permission denied,” “temporarily unavailable” and “success.” If the result loses that distinction, the agent may invent a record or treat a denied action as completed. The Foundry tool best-practices guide emphasizes tool descriptions, tracing and input validation.
A service-desk assistant proposes creating a maintenance visit. The downstream API commits the booking, but the response times out before reaching the agent runtime. If the model blindly repeats the call, a second appointment appears. This is not primarily an LLM problem; it is a distributed transaction problem. The request should carry a stable idempotency key derived from the approved proposal, and a repeat execution should return the first booking’s result where the service supports idempotent handling.
The application needs three states that the model should never blur: requested, committed and confirmed, and outcome unknown. An ambiguous timeout belongs in the third state until the system can query a durable booking record. Responding “Your appointment is confirmed” before that check would be misleading. Conversely, announcing failure when the visit was actually created could cause a user to retry manually. Tool-result contracts should include a result class, record identifier when known and a correlation ID for uncertain outcomes.
An approval service must ensure that the model cannot trade one approved proposal for another. An approval for asset M-37 at one time window should not authorize a visit for asset M-38 simply because the tool names are similar. Bind the approval to the exact typed arguments, tenant and requester; validate the binding immediately before commit. The model may explain what was proposed, but only deterministic application code should decide if it matches an authorized operation.
For a release test, simulate the network dropping the success response after the booking database commits. Check that the second call returns the original visit, that the assistant does not invent a second confirmation, and that the trace links proposal, approval, API call and durable outcome. This is a more meaningful agent reliability test than an isolated example where every tool responds instantly.
Make tool errors actionable without exposing secrets
A tool that times out should return a controlled error type and correlation ID. The model can then explain that the system could not confirm the operation. It should not claim “appointment booked” just because it sent a tool request. Do not expose internal stack traces, bearer tokens, raw credentials or sensitive database details in responses, traces or tool-return text.
A useful failure drill is a visit-creation API that commits the appointment but loses the response. On retry, the idempotency key should return the original appointment rather than creating a duplicate. The user-facing agent should report only the state it can verify. This test exposes whether the workflow assumes model fluency is equivalent to a durable transaction.
Restrict external connectors by destination and identity
A tool can cross organizational boundaries: web search, an internal wiki, a ticketing service, an MCP server or a custom API. Review which data may leave the tenant and which identities the tool uses. A service-to-service managed identity may prove that the agent is an authorized workload; it does not automatically prove that the end user may view a particular customer record.
Apply fine-grained access checks in the connector and include user context only through a legitimate delegation mechanism. Don’t place sensitive user tokens inside a prompt as the model’s “instructions.” Treat returned documents or API strings as untrusted task data, not new operating rules. This is especially important when agents retrieve text from external sites that an attacker can edit.
Separate drafts from commitments
For consequential actions, make the agent draft a proposed change first. The proposal should include parameters and any supporting evidence: “Request a maintenance visit for asset M-37 at Building 2 on Tuesday afternoon, based on the current service calendar.” The application can validate asset ownership, allowed scheduling windows and whether the requesting person has authority to commit the action.
Human approval should apply to the exact operation and parameters. If the model later changes the asset or date, the original approval should no longer authorize the new request. Keep an audit record linking the proposal, source evidence, approval and final tool result. This creates useful transparency without making every harmless read-only lookup require a confirmation.
Test agent state and conversation continuity
State is more than a transcript. It may include resolved identifiers, retrieved evidence versions, outstanding approval proposals and tool results. Decide what must persist after a restart and what should expire. An agent should not reuse a previously approved action after the user changes context or the target record changes, even when its memory of an earlier conversation seems relevant.
Try switching users and tenants during a development test. Session IDs must not grant cross-tenant access. Test a stale account lookup and a resumed conversation where the original permission has been revoked. The correct behavior is to re-check authoritative data and authorization at the point of action. Conversation memory is helpful context, not a substitute for current state.
Instrument each proposed and executed action
For an incident, operators need to distinguish the model’s chosen tool from the actual API call and its outcome. Record tool name, normalized argument keys, principal, version, authorization decision, duration, result class and trace ID. Apply data-minimization rules to sensitive values. A visible history of “model chose tool A” is not enough to know whether tool A executed successfully.
Use representative test cases for correct tool selection, unnecessary tool use, unsupported operations, invalid arguments and prompt injection in tool results. Evaluate not only the final answer but the sequence of calls. A high-quality text response can conceal that an agent queried the wrong customer or spent ten unnecessary steps. Trace-level assessment reveals mistakes a user-facing answer-only score may miss.
Build one safe service-desk agent end to end
As a practical AI-103 exercise, start with a read-only asset lookup and verify access filtering. Add a service-calendar tool with a typed result. Then add a write tool that creates a draft request with an idempotency key and a separately authorized commit endpoint. Deliberately return an empty lookup, a permission denial and an ambiguous timeout. Make the assistant report each accurately.
Finally, test malicious instructions hidden inside the returned equipment manual. The agent may quote technical facts from the manual but must not run actions the manual tells it to run. When that boundary holds under realistic error conditions, the system has moved beyond a convincing demo toward an agent that can safely participate in a real application.