INSIGHTS
AI & Data

AWS AIF-C01: Foundation Models & Generative AI

In this article
  1. Understand what a foundation model contributes
  2. Select models according to capability and operating constraints
  3. Use prompt design to make the task explicit
  4. Ground responses with retrieval when knowledge must be current
  5. Use embeddings for semantic retrieval and similarity
  6. Add tools and agents only when the application needs action
  7. Apply guardrails without confusing them with correctness
  8. Evaluate model behavior before and after release
  9. Operate generative AI as a governed application system

Foundation models change the way applications are built because much of the model capability already exists before the application team begins. Instead of collecting a labeled dataset for every task, developers can use a general model for generation, summarization, classification, extraction, conversation, code, or embeddings and then adapt the application through prompts, retrieval, tools, fine-tuning, or other techniques. The engineering challenge moves from “How do we train a model from scratch?” toward “How do we select, ground, constrain, evaluate, and operate a model for this specific use case?”

The current AIF-C01 exam guide explicitly includes generative AI fundamentals, foundation-model applications, responsible AI, and security/governance. Amazon Bedrock is central to that AWS application layer, but the durable concepts matter beyond any one service.

Understand what a foundation model contributes

A foundation model is trained on broad data and has enough capacity to support many downstream tasks. Large language models are one important type, but foundation models can also work with images and other modalities. The model provides general learned capability; the application supplies instructions, context, user input, business data, and sometimes tools that allow the model to perform a useful task.

This explains why a simple prompt can produce useful text without task-specific training. It also explains the limits. The model does not automatically know the organization’s current policies, private documents, customer state, or authorization rules. Those must be supplied through controlled context or external systems. A model can generate plausible language even when its knowledge is incomplete, so fluency is not evidence of correctness.

General machine-learning concepts still apply: model behavior is probabilistic, evaluation requires representative data, and the production system includes much more than the model artifact.

Select models according to capability and operating constraints

Amazon Bedrock provides access to multiple foundation models through a managed service. Model selection should consider task quality, supported modalities, context-window needs, latency, throughput, cost, model-provider characteristics, supported Regions, and any requirements around customization. The model with the highest score on a public benchmark may not be the best model for a company’s actual prompts.

Create a small evaluation set that represents the difficult cases in the use case. Compare candidate models on the same prompts and scoring criteria. For a support assistant, those criteria might include factuality, correct citation of source material, refusal behavior, tone, and response time. For extraction, structured-output accuracy may matter more than conversational quality.

A model decision should be revisitable. Providers release new versions and capabilities, and application needs change. Keep prompts and evaluation cases portable enough to compare alternatives without rewriting the entire system.

Model access is also an application dependency. Verify availability in the target Region, quotas, supported inference features, and any provider-specific conditions before committing to a design. A model that works in a developer sandbox but is unavailable in the required production Region is not a production option.

Use prompt design to make the task explicit

Prompt engineering is the application’s way of specifying the role, task, context, constraints, and output format expected from the model. Strong prompts tell the model what decision is being supported and what evidence it may use. They can include examples, formatting rules, refusal criteria, and instructions to distinguish known information from uncertainty.

Avoid making the prompt a hidden policy engine. Authorization, transaction limits, compliance rules, and irreversible business actions should be enforced by deterministic controls outside the model. The prompt can explain those rules and guide behavior, but it should not be the only barrier protecting a sensitive operation.

Version prompts like code when they materially affect production behavior. A one-line wording change can alter response style, tool selection, or classification boundaries. Tie prompt versions to evaluation results so teams can explain why a change was promoted.

Separate system instructions from user content and retrieved data so the application can reason about trust. Text retrieved from an external document should not silently gain the same authority as the system policy. Prompt injection becomes easier to manage when the application knows which content is instruction, which is evidence, and which is untrusted user input.

Ground responses with retrieval when knowledge must be current

Retrieval-augmented generation combines a model with external knowledge. Amazon Bedrock Knowledge Bases can retrieve relevant content from configured data sources and use the retrieved context when generating a response. This pattern is valuable when the application must answer from private or frequently changing information rather than relying on the model’s pretraining alone.

RAG quality depends on the entire retrieval pipeline: document selection, permissions, chunking, embeddings, indexing, query transformation, retrieval ranking, and the final prompt. A poor result can come from retrieving the wrong passage even when the model behaves perfectly. Bedrock supports RAG evaluation so teams can examine retrieval and generated-response quality rather than treating the pipeline as a single black box.

Grounding also improves explainability when the application can show the source material used for an answer. It does not eliminate hallucination, so the product should still define what happens when evidence is missing or contradictory.

Define citation behavior as part of the product contract. If an answer claims to be grounded, the application should preserve which chunks were retrieved and which sources were presented to the model. This gives users and operators a path to verify the answer and helps engineers distinguish a retrieval failure from a generation failure.

Use embeddings for semantic retrieval and similarity

Embeddings convert content into numeric vector representations that capture semantic relationships. They enable applications to retrieve material that is conceptually similar even when the query and source do not share exact keywords. This is a common building block for RAG, recommendation, semantic search, deduplication, and clustering.

Embedding design includes choices about model, dimensionality, chunk size, metadata, vector store, filtering, and update strategy. Business metadata can be as important as vector similarity. A support assistant may need to filter by product, region, customer entitlement, or document status before ranking semantic matches.

Do not evaluate embeddings only with a few obvious examples. Build queries where terminology differs, where multiple documents are similar, and where an authoritative document must outrank an outdated one. Retrieval quality is a measurable system property, not a side effect of choosing a model.

Plan for index freshness. Policies, product documentation, and customer records change. A retrieval system that updates weekly may be unacceptable for information that changes hourly. Define ingestion latency, deletion behavior, and re-embedding triggers so stale knowledge does not remain searchable after the source has changed.

Add tools and agents only when the application needs action

A conversational model that answers questions is different from an agent that can choose tools and perform operations. Amazon Bedrock Agents can orchestrate interactions with tools or knowledge sources, which makes workflows more powerful but also increases the security consequences of model decisions. The architecture must define what the agent is allowed to do, under which identity, and which actions require confirmation.

Design tool interfaces to be narrow and explicit. A tool that performs “any database operation” creates a large blast radius, while a tool that retrieves an order by authorized customer ID is easier to reason about. Validate inputs and outputs outside the model, apply least privilege to the execution identity, and record the tool calls for investigation.

For teams moving toward production GenAI engineering, the approved AIP-C01 destination represents a deeper application-development layer than the foundational AIF-C01 view. The distinction matters: knowing what an agent is conceptually is not the same as operating one safely.

Apply guardrails without confusing them with correctness

Amazon Bedrock Guardrails can evaluate user inputs and model responses against configurable policies such as content filters, denied topics, sensitive-information filters, word filters, and other safeguards. Guardrails can be used with inference, agents, knowledge bases, and flows. They provide a consistent safety layer across supported models.

Guardrails are not a guarantee that an answer is factually correct. Content safety, privacy, grounding, authorization, and factuality are separate concerns. A response can be harmless but wrong, or factually correct but reveal sensitive information to an unauthorized user. Each risk needs an appropriate control.

Test guardrails with adversarial and borderline examples. Overly aggressive filters can block legitimate users, while permissive settings may miss harmful inputs. AWS recommends ongoing testing as underlying safeguard models evolve, so guardrail configuration should be treated as a monitored production control.

Evaluate model behavior before and after release

Generative AI evaluation should begin with task-specific acceptance criteria. Automated metrics can measure some properties, while human review remains important for qualities such as helpfulness, tone, business appropriateness, and nuanced factuality. Amazon Bedrock evaluation features support model and RAG assessment, including programmatic and human-based approaches.

Keep evaluation datasets representative and versioned. Include routine cases, hard cases, known failure modes, sensitive prompts, ambiguous questions, and examples where the correct behavior is to refuse or ask for clarification. Re-run the suite when models, prompts, retrieval, guardrails, or tools change.

Traditional data-science foundations are still relevant: define what success means, separate development from evaluation, avoid cherry-picking examples, and measure the system against real use rather than impressive demos.

Operate generative AI as a governed application system

Production architecture includes identity, network boundaries, encryption, logging, cost controls, quotas, data classification, model access, prompt management, retrieval stores, monitoring, and incident response. Foundation models are powerful components, but the application still has normal cloud responsibilities. An AWS security-services view remains relevant because GenAI systems must be monitored and governed like other workloads.

Observability should capture enough information to troubleshoot quality and security without indiscriminately logging sensitive prompts or responses. Track model, prompt version, latency, token use, tool calls, retrieval behavior, guardrail interventions, failures, and user feedback according to privacy requirements. Cost and quality often change together, so teams need both views.

Capacity planning matters too. Token limits, requests per minute, model concurrency, retrieval throughput, and downstream tool quotas can all become bottlenecks. Load tests should use realistic prompt sizes and multi-step agent behavior, because a conversational demo rarely reflects the concurrency and context length of production traffic.

Foundation models reduce the amount of model training required for many applications, but they do not reduce the need for engineering judgment. The strongest AWS GenAI systems choose models deliberately, ground them in trusted data, constrain actions, apply layered safeguards, measure behavior, and preserve the operational evidence needed to improve the system over time.

They also plan for graceful degradation. If retrieval is unavailable, a model may need to refuse rather than answer from general knowledge. If a tool fails, the agent should return a safe error instead of inventing a result. If a preferred model is throttled, a fallback model should be evaluated before it is placed in the failover path. Reliability policy is part of GenAI quality.

Model portfolios also need change control. When a newer model becomes available, compare quality, latency, context limits, safety behavior, and cost against the production baseline before changing the default.

Filed under AI & Data