A visual model can describe a photograph, answer a question about a diagram, identify a region in an image or summarize a video segment. That evidence-reading role differs from image and video generation, which creates or edits media and requires a different fidelity and provenance review. These capabilities are useful, but they are not interchangeable with trusted measurement. A caption that says a worker appears to wear safety equipment does not prove the equipment satisfies a regulation, and a fluent summary of a video does not guarantee that the system inspected every frame relevant to the user’s question.
The computer-vision domain of AI-103 includes multimodal understanding, captions, grounded visual question answering, accessibility descriptions, video interpretation and responsible use. Developers should design input preparation, model selection, confidence handling and evaluation for the actual business task rather than selecting a model solely because it can accept an image.
Start with the decision the visual system must support
Imagine an inventory app that photographs a damaged component. One task is to produce a helpful caption for a human technician. Another is to read the serial number. A third is to decide whether to order a replacement part. Those tasks have different evidence requirements. Caption generation may tolerate approximate descriptive wording; a serial number needs character-level accuracy; replacement decisions require product compatibility and business rules beyond the image itself.
Write an expected result schema for each. A caption might contain object, visible condition and uncertainty. A serial-number extractor should return normalized characters, crop or region reference and a review flag. A replacement recommendation should additionally identify a verified catalog record and approval state. Avoid moving directly from uncertain visual inference to a costly business action.
Prepare images so important evidence survives
Input resolution, lighting, orientation, cropping and compression determine what a model can see. A small safety label photographed from a distance may be unreadable even if the overall scene is clear. A tight crop improves some extraction tasks but can remove context needed to interpret a sign or machine component. Store the original image and any derived crops with traceable identifiers where data policy permits.
Test with blurry images, rotated receipts, partially obscured objects and different mobile cameras. The correct outcome is sometimes “cannot read the text” or “insufficient view,” not a fabricated number. If a model returns a serial number, compare it with a checksum or known product format when available. For high-impact uses, show the source crop to a reviewer instead of hiding uncertainty behind a polished response.
Separate OCR, visual reasoning and entity extraction
OCR identifies visible characters. Visual reasoning describes scene relationships, objects and meaning. A multimodal extraction workflow may combine both, but they should be evaluated separately. A model can correctly recognize that a photograph contains a pressure gauge while misreading the number on its dial. A layout-aware document service may extract printed table cells more consistently than a general-purpose caption model.
For a maintenance form, use OCR or content extraction to identify written values and preserve their positions. Use visual understanding where a photograph of damage or a schematic requires scene interpretation. Apply deterministic validation against asset records before using the result as an operational decision. Microsoft’s Content Understanding documentation describes multimodal extraction across images, documents, audio and video; it does not remove the need for domain-specific verification.
Generate captions that reflect evidence and accessibility needs
An accessible image description should communicate important visual information without inventing intent, identity or details not visible. For a chart, that may include axis labels, trend and notable comparison; for an application screenshot, it may include the critical control and visible error message. A short alt-text field and a long description serve different user needs. Do not fill every image with keyword-heavy alt text that does not help a person understand it.
Test generated descriptions with subject-matter experts and accessibility users where possible. Check whether the output omits an essential label, confuses foreground and background, or repeats nearby text unnecessarily. When a diagram conveys an exact numeric value, extract and verify the number instead of having the model describe it approximately. Human review remains appropriate for regulated or legally consequential descriptions.
Understand video as a sampled sequence with time
A video combines image frames, timing, motion, scene transitions and sometimes audio. A static frame may miss the event that matters, such as a valve moving from closed to open. The application must choose how video is segmented and which frames or moments are analyzed. Sampling every few seconds may be adequate for a broad summary and inadequate for a brief safety incident.
Preserve timecodes and segment identifiers with findings. If a user asks what happened before an alarm, the system needs enough temporal context to compare the relevant frames. A result such as “warning light appeared at 02:14” should link to a reviewable clip rather than only a generated narrative. Validate that the media analysis service and chosen model actually support the input format, size and region required by the workload.
A warehouse incident video is often summarized as “a worker entered a restricted area,” but that claim depends on temporal evidence. A single frame showing a person near a door may not prove entry, and two frames sampled far apart may omit the relevant movement entirely. Define the question before selecting frames: identify the moment of crossing, the relevant restricted-zone boundary and the time uncertainty introduced by sampling. Return evidence with source timestamps so a reviewer can inspect the original sequence.
For a safety-assistance workflow, first preserve the original video and its access controls. Produce a sequence of candidate frames or segments and ask the visual model for observations rather than definitive conclusions about intent. Distinguish “the person appears behind the marked line at 00:24” from “the person deliberately violated policy.” The latter is an interpretation that may require broader context and human judgment. If the camera view or resolution is insufficient, the system should record uncertainty rather than invent a movement path.
Text inside a frame is also untrusted input. A sign in the video might contain a QR code, an address or instructions directing a model to ignore its operating rules. OCR can be useful for labels, but extracted strings must remain data rather than commands to tools. If a video-analysis assistant has permission to send security notifications, route proposed alerts through a policy check rather than allowing text shown in the frame to trigger an arbitrary API call.
Quality evaluation should include clips with partial occlusion, low light, identical-looking objects, time discontinuities and misleading captions. Check source timestamp accuracy, whether the claimed action is visible and whether the system declines to make unsupported inferences. Visual accuracy is a claim about evidence over space and time—not about how detailed the caption sounds.
Handle visual prompt injection as untrusted content
An image can contain printed text that tells an AI assistant to ignore its original task or disclose private data. That text may be useful as evidence—for example, a photographed instruction manual—but it must not become a new system policy. Treat recognized image text as lower-trust content. Restrict downstream tools and authorization independently of whatever instruction-like content appears in a photo.
Build tests with an otherwise legitimate product label that includes a hidden or overt instruction to export records. The application should extract the relevant label data and ignore the command. A second test can include sensitive information on an unrelated background screen; determine whether the application should crop, redact or decline to process that region. Safety should apply to input collection and data retention as well as to model output.
Distinguish the failure types your operators will see
A blank caption may result from unsupported file format, authorization failure, a blocked safety category, a model or regional restriction, or an image that contains no recognizable subject. An inaccurate caption may stem from low resolution or model inference. A video analysis that times out may be an ordinary long-running-job issue rather than evidence that the clip is invalid.
Log the input media type, processing stage, model/analyzer version, result status, duration and a safe source ID. Avoid copying raw media into every trace. During a failed run, operators should be able to decide whether to resubmit, change preprocessing, request a clearer image or escalate to a reviewer. Don’t return a false “no damage found” merely because analysis failed.
Evaluate against task-specific reference examples
For captions, have experts judge whether the important visible facts were included and whether unsupported claims were introduced. For text reading, compare exact characters and numeric units. For visual question answering, record expected evidence regions and whether answers remain grounded in them. For video, check event timing and missed segments. A generic image-quality score is not enough to validate these different requirements.
A Microsoft AI-103 preparation lab can use a small controlled dataset of receipts, machine labels, charts and short video clips. Create cases where the correct output is uncertainty. Inspect mistakes in the original media and the model response side by side, and decide whether preprocessing, service choice or downstream validation is responsible. That is the practical skill behind multimodal understanding: knowing which visual claims are supportable and what must still be verified.