Speech changes the assumptions of an AI application. A typed request arrives as a complete piece of text; a spoken request may arrive in fragments, include background noise, and depend on the timing of interruptions. A voice response is also consumed sequentially, so a long tool delay or verbose error message feels very different from an extra paragraph in a chat window. Developers must design recognition, synthesis, translation, turn-taking and tool behavior as one interactive system.
The text-analysis portion of AI-103 includes speech-to-text, text-to-speech, translation and multimodal audio interactions. A useful design example is a multilingual service hotline that identifies the caller’s issue, retrieves a policy, and proposes a service visit only after verifying the account. Speech improves accessibility and convenience, but it cannot weaken the identity or authorization controls applied to a text request.
Select an audio workflow by the user experience
Speech-to-text converts audio into text that an application can interpret or search. Text-to-speech turns approved text into audio. Speech translation combines recognition and translation, sometimes with synthesis in another language. A live voice agent also manages partial transcripts, interruptions, latency and audio playback. Decide which of these experiences is actually required before selecting services.
An offline transcription job may process a recorded meeting and produce a timestamped transcript after several minutes. A customer speaking to a virtual assistant expects far more immediate feedback. An accessibility reader may need predictable pronunciation and control over playback speed. The input modality alone does not determine the architecture; timing, user control and the consequences of mistakes do.
Prepare audio and consent for collection
Before sending audio to a service, define recording consent, permitted retention and who can access the transcript. Microphone recordings may contain voices of people other than the intended user, background account details or accidental private conversations. Do not treat a conversation as a harmless text string after it has been transcribed; its privacy obligations survive the conversion.
Audio format, sampling, channel separation and background noise influence recognition quality. A low-quality phone channel can obscure similar-sounding product identifiers. Test accents, speaking rate, background voices and silence. If the system cannot reliably identify an important number, ask the caller to confirm it rather than turning an uncertain recognition into a database update.
Distinguish partial and final transcripts
In streaming recognition, an interim hypothesis can change as more audio arrives. A partial transcript might interpret “cancel my…” before the caller completes “cancel my appointment reminder, not my appointment.” Triggering a consequential tool from the interim words would be a serious design failure. Choose when the application considers a spoken intent complete and when it must request clarification.
Separate display behavior from action behavior. The interface can show provisional words with a visible status while waiting for a finalized segment. The action layer should use confirmed intent and validated identifiers. Even a final transcript can be wrong; the caller may need to repeat a serial number or approve a summary of the requested change. Turn management is part of correctness, not just a user-interface detail.
Use speech synthesis for clear, controlled responses
A generated text answer may include a long source citation or code block that is cumbersome when spoken aloud. For a voice agent, create an output style that conveys the essential fact first, uses short sentences and offers a way to request more detail. Do not strip mandatory safety or financial conditions merely to make the response shorter. Spoken responses should avoid reading hidden tokens or private identifiers that would not be appropriate in a shared environment.
Text-to-speech voice choice and speed affect comprehension and accessibility. Test names, units, dates, numbers and abbreviations in the languages the application supports. A model may pronounce an account number in a confusing way; a deterministic speech formatting layer can group digits or spell them according to context. An audio response is useful only when users can understand and verify the relevant detail.
Treat translation as an additional uncertainty layer
If a caller speaks one language and a backend tool expects another, errors can arise during recognition, translation or final interpretation. Keep the original transcript and the translation linked with timecodes and a controlled trace ID where permitted. Validate critical business terms and quantities. A translation may read naturally while changing a negation or condition that affects the decision.
Use terminology guidance for domain-specific words, such as equipment names, medication labels or service-plan tiers. For ambiguous requests, confirm the interpreted meaning in the caller’s own language before acting. Do not infer that a positive sentiment label proves consent to a change; authorization requires a deliberate, interpretable approval event.
Integrate real-time voice agents with tools cautiously
A voice agent that consults a knowledge base and schedules appointments has the same trust boundaries as a text agent: read-only retrieval, authenticated records, typed tool arguments and approval before consequential writes. It also has stricter latency requirements. A slow connector can leave awkward silence, and repeated model turns may cause the caller to interrupt or restate the request.
Use short, informative progress messages when appropriate, and decide what happens if the caller speaks while a tool is running. The system should not execute two copies of the same write simply because the user repeated an utterance. Bind pending actions to transaction IDs and ask for confirmation of the final details. Foundry Agent Service documents voice-based agent options, but availability and model-region support should be checked at implementation time.
Consider a multilingual insurance hotline. The caller says “don’t file the claim yet,” but noise obscures the first word, and a partial speech-recognition hypothesis arrives without the negation. If the agent treats streaming partial transcripts as authorization, it may submit a claim before the final corrected transcript appears. For sensitive actions, wait for a stable utterance, confirm the interpreted intent and bind any final submission to the caller’s verified identity and explicit approval.
Divide the voice workflow into capture, transcription, intent interpretation, evidence lookup, proposed action and spoken confirmation. Each boundary should have its own error and latency budget. A network interruption while producing speech does not necessarily mean the underlying claim transaction failed; the application must check the durable result before telling the user to try again. Similarly, a time-limited confirmation prompt cannot reuse a prior approval after the user changes the beneficiary or request amount.
Translation introduces another layer of uncertainty. If the caller speaks one language and the service record uses another, preserve the original utterance or a privacy-compliant audit reference, plus the interpreted language and translation version. Route medical, legal or high-impact terms to a controlled glossary and perform human review when a misunderstanding could cause material harm. A speech pipeline that accurately transcribes words can still misinterpret the business operation those words authorize.
End-to-end tests should include interruptions, accents, mixed-language terms, background television, repeated caller IDs and someone other than the verified account holder speaking. Measure false action rates and correction handling alongside word error rate and synthesis quality. A pleasant voice is not evidence of a trustworthy agent; reliable voice applications make the relationship between what the caller said and what the system actually did auditable.
Keep audio safety, identity and telemetry distinct
A user should not gain access to a confidential record by asking for it aloud instead of typing. The speech pipeline may transcribe the request, but the same downstream access checks must apply. Audio may also carry prompt-injection attempts through playback or recorded media that the agent is supposed to summarize. Treat transcribed source content as data; do not elevate its instructions into the assistant’s operating policy.
Record recognition error classes, language selection, tool latency, time to first audible response, interruptions and task completion. Privacy controls may require storing aggregate measures rather than raw audio. The system should distinguish “no speech detected,” “transcription unclear,” “tool denied access” and “speech synthesis unavailable” so troubleshooting does not collapse every failure into a generic voice-service error.
Evaluate by language, channel and consequence
Create a test set with realistic accents, mobile network quality, noisy environments, translated technical words and interrupted sentences. For transcription, compare words or key entities against a reference; for translation, compare intent and critical terminology; for speech synthesis, evaluate clarity and pronunciation; for the complete agent, measure whether the correct authorized task was performed without duplicate actions.
Include cases where the correct outcome is to ask the caller to repeat or spell an identifier. The application should not guess a customer’s identity from an imperfect transcript simply to sound confident. Measure tail latency and error recovery, not only recognition accuracy on clean studio audio. A system that performs perfectly in a quiet room but fails on ordinary phone calls has not passed a realistic readiness test.
Create a safe multilingual hotline lab
For Microsoft AI-103, implement a prototype that accepts a spoken support request, recognizes the language, captures a transcript, retrieves an approved policy and reads back a short answer. Add an optional read-only service-status lookup. Test unsupported language, a deliberately unclear order number and a network timeout. Observe whether the caller gets an understandable explanation rather than fabricated progress.
Then introduce an appointment proposal that requires a separate confirmation. Replay the same utterance after an ambiguous tool timeout and verify idempotency. This experiment teaches how speech, models, tools and business authorization interact. Voice AI becomes reliable when people can correct it, understand its uncertainty and retain control of actions taken on their behalf.