INSIGHTS
Cloud Computing

ISACA CRISC: Technology Risk in Cloud and AI Programs

In this article
  1. Start with the business use case and dependency map
  2. Apply shared responsibility as a risk model, not a slogan
  3. Treat identity and privilege as primary cloud risk drivers
  4. Assess data risk across cloud and AI lifecycles
  5. Model AI-specific failure modes without treating AI as a separate universe
  6. Control infrastructure and policy through automation
  7. Address concentration, lock-in, and resilience explicitly
  8. Manage supplier and software supply-chain risk
  9. Monitor emerging risk and keep the program governable

Cloud platforms and artificial intelligence programs can accelerate product delivery, analytics, automation, and new customer experiences, but they also change the way technology risk is created and controlled. Risk professionals need to understand not only individual security weaknesses but also concentration, service dependencies, data use, model behavior, identity, software supply chains, provider responsibilities, and the pace at which managed services evolve.

The current CRISC outline includes emerging technologies, enterprise architecture, operations, data lifecycle management, resilience, security principles, and risk response. That combination is well suited to cloud and AI because the risk often spans governance, technology, suppliers, and business use rather than fitting into one technical control domain. A risk-based approach should preserve business value while making accountability and residual exposure explicit.

Start with the business use case and dependency map

Cloud and AI risk should be assessed in the context of what the program is trying to achieve. A customer-facing AI assistant, internal productivity tool, fraud model, data platform, and cloud-hosted ERP system can have very different consequences even when they use similar infrastructure.

Dependency mapping should identify cloud providers, regions, identity services, data stores, APIs, model providers, open-source components, SaaS integrations, monitoring systems, and human approval points. The most important risk may sit outside the component that receives the most technical attention.

Criticality also depends on reversibility. A pilot that can be shut down safely is different from a production workflow that becomes the only way to approve loans or dispatch field staff. Risk treatment should consider what happens if the capability is unavailable, produces wrong results, or must be withdrawn quickly.

Program owners should be able to describe which business objectives depend on the technology, which failures would matter, and who has authority to accept exposure. Without that connection, technical teams may optimize controls without knowing what level of risk the business can tolerate.

Risk mapping should include the control plane as well as the production workload. Billing accounts, organization policies, identity tenants, CI/CD systems, secrets managers, model registries, and observability platforms can determine whether teams can safely operate or recover the service. These shared components often create broader impact than one application.

Apply shared responsibility as a risk model, not a slogan

Cloud services divide operational responsibility between provider and customer, and the split changes by service model. The provider may secure the physical infrastructure and managed platform while the customer remains responsible for identity, configuration, data, workload behavior, network exposure, and how users consume the service.

Risk assessments should identify inherited controls and customer-controlled settings explicitly. A provider certification can support assurance for controls the provider operates, but it does not prove that the customer’s tenant, identities, encryption choices, logging, or application design are secure.

Managed AI services create additional responsibility boundaries. The provider may operate model hosting while the customer decides which data enters prompts, which tools or plugins an agent can invoke, what outputs are trusted, how harmful responses are handled, and whether sensitive information is retained or exposed through telemetry.

The controls discussed in cloud security engineering are most useful when mapped to the exact responsibility boundary. Risk treatment should close customer-side gaps instead of assuming the provider’s control environment covers every layer.

Responsibility mappings should be reviewed by service type. Infrastructure, platform, SaaS, serverless, managed databases, and hosted AI shift control boundaries differently. A template that assumes the same customer responsibility for every service can leave important gaps or duplicate controls that the provider already operates.

Treat identity and privilege as primary cloud risk drivers

Cloud control planes concentrate authority. A compromised administrator, automation role, service principal, or federation trust can modify networks, data access, logging, encryption settings, or workloads across large portions of the environment. Risk assessment should therefore evaluate identity pathways as carefully as host vulnerabilities.

Least privilege should include both human and machine identities. Workload identities, API tokens, CI/CD roles, and service accounts often operate continuously and can accumulate access over time. Their permissions, trust conditions, credential lifetime, and ability to assume other roles should be visible.

Emergency and break-glass access requires separate governance because it intentionally bypasses normal controls. Accounts should be strongly protected, monitored, periodically tested, and limited to conditions where ordinary identity services are unavailable or compromised.

Identity risk also extends across providers. Federation between workforce identity, cloud platforms, SaaS applications, and development systems can create transitive trust. A change in one identity provider may affect many services, so concentration and recovery should be part of the scenario.

Privileged change to identity policy should be monitored as a high-risk event. Creating a new trust, weakening conditional access, changing token lifetime, or granting broad service-role permissions can alter the effective attack surface instantly without deploying new application code.

Periodic entitlement review should focus on effective privilege, not just assigned roles. Nested groups, inherited permissions, cross-account trusts, and automation that can assume higher authority may create access that is not obvious from a simple user list. Effective-access analysis is especially important in large cloud estates.

Assess data risk across cloud and AI lifecycles

Cloud and AI programs can create new copies of data through staging areas, feature stores, prompts, embeddings, model training, logs, backups, caches, and analytics exports. Risk assessment should follow data through the lifecycle rather than assuming the primary database is the only sensitive location.

Classification and purpose limitations matter. Data that is acceptable for one analytics use may be inappropriate for model training or a third-party AI service. Teams should know whether personal, confidential, regulated, or intellectual-property data can be used, retained, transferred, or exposed in outputs.

Encryption protects storage and transport but does not solve misuse by authorized services. Access control, minimization, retention, lineage, and monitoring remain necessary. The risk question is who or what can act on the data and how that authority is constrained.

AI introduces additional integrity concerns. Training-data poisoning, prompt injection, manipulated retrieval sources, and unvalidated model outputs can affect decisions even when confidentiality controls are strong. Risk scenarios should include incorrect or adversarial behavior, not only data theft.

Data provenance is increasingly important in AI programs. Teams should know where training, retrieval, and evaluation data originated, whether usage rights are clear, how quality was assessed, and whether sensitive content can persist in derived artifacts. Weak provenance can create legal, quality, and security risk even when storage controls are strong.

Model AI-specific failure modes without treating AI as a separate universe

AI risk includes security, privacy, bias, reliability, explainability, misuse, model drift, and dependence on external models or tooling. The concerns summarized in AI security risks should be translated into scenarios tied to the organization’s actual use case rather than copied into a generic AI risk register.

Human oversight should be designed around consequence. Low-risk drafting assistance may need simple review, while a model that influences hiring, financial decisions, safety, or security actions may require stronger approval, validation, logging, and escalation.

Model performance can degrade when data, users, adversaries, or business conditions change. Monitoring should look for drift, unexpected outputs, policy violations, and changes in the distribution of inputs. A model that was acceptable at launch may become risky without any code change.

AI components should still follow ordinary technology governance: ownership, change management, supplier review, incident response, access control, data governance, resilience, and retirement. Novel technology changes implementation details, but it does not remove the need for accountable management.

AI agents introduce an additional authority problem because generated output can trigger actions in other systems. Tool access should be constrained by least privilege, validation, transaction limits, and human approval where consequences are high. Risk assessment should consider what the agent can do, not only what it can say.

Evaluation should be tailored to intended use. Accuracy, groundedness, harmful-content behavior, privacy leakage, prompt-injection resistance, and tool-use safety may all matter differently depending on the application. A single benchmark score cannot represent every business risk, so testing should reflect realistic tasks and failure consequences.

Control infrastructure and policy through automation

Cloud environments change quickly enough that manual review alone cannot provide timely assurance. Infrastructure as code, policy as code, configuration scanning, and continuous compliance can enforce or detect expected settings at deployment and during operation.

The distinction between automation and orchestration in infrastructure as code matters because risk can arise from both the individual control and the workflow that coordinates many changes. A secure module can still be deployed through an overly privileged pipeline or into an unapproved environment.

Control automation should be governed like production software. Source control, code review, testing, approved modules, restricted deployment authority, and traceable releases help prevent a flawed policy from propagating across the entire estate.

Terraform security practices illustrate another risk: secrets or privileged configuration can be embedded in code, state, logs, or pipeline variables. Automated delivery improves consistency only when credentials and sensitive outputs are also controlled.

Automated guardrails need an exception path that is fast enough for legitimate emergencies but controlled enough to prevent permanent bypass. Exceptions should be logged, reviewed, time-limited, and visible in risk reporting so delivery pressure does not silently erode the policy baseline.

Address concentration, lock-in, and resilience explicitly

Cloud and AI services can create concentration risk when many critical workloads depend on one provider, region, model vendor, identity service, or data platform. High availability inside one provider does not necessarily protect against a provider-wide control-plane failure or strategic dependency.

Risk treatment may include multi-region design, portable data, tested exports, alternate providers, contractual commitments, local fallback, or the ability to degrade gracefully. True multi-provider architecture can be expensive, so the organization should choose resilience mechanisms based on business impact rather than assuming diversification is always practical.

Lock-in is not automatically a defect. Managed services may create substantial value. The risk question is whether management understands the dependency, can monitor it, and has a credible response if pricing, service quality, legal terms, product direction, or availability changes materially.

Recovery exercises should test provider and identity dependencies as well as application restoration. If a cloud account is inaccessible or an AI provider is unavailable, teams need to know how decisions, data access, communications, and critical operations continue.

Cloud exit planning should include configuration and identity, not only data export. Recreating network policy, encryption keys, access models, deployment pipelines, and operational runbooks may take longer than copying stored data. The organization should understand which capabilities are portable and which are deeply provider-specific.

Manage supplier and software supply-chain risk

Cloud and AI programs depend on providers, model vendors, open-source packages, container images, build actions, APIs, data suppliers, and sometimes fourth parties that are invisible to the end customer. Each dependency can introduce security, availability, legal, and integrity risk.

Due diligence should be proportionate to criticality. Teams can review provider assurance, security practices, ownership, resilience, contractual commitments, breach history, data handling, and supply-chain tiers. High-impact services should receive deeper assessment than low-risk tools.

Software provenance also matters. Organizations should know where critical images, packages, and models originate and how updates are approved. A trusted deployment pipeline can reduce risk only if the artifacts entering it are themselves controlled.

CISA contributes an assurance perspective by asking whether provider and supply-chain controls are evidenced rather than assumed. Risk functions should use audit results, incidents, and monitoring to update the treatment of important dependencies.

The ISACA certifications perspective is useful because cloud and AI dependencies cross audit, risk, and security management. A provider can be technically strong while still creating unacceptable concentration, contractual, privacy, or continuity risk for a specific organization.

Model and package updates should trigger controlled reassessment when behavior or dependencies can change materially. Pinning versions, validating release notes, testing in representative environments, and keeping rollback options can reduce the risk of silent functional or security changes entering production through automated updates.

Monitor emerging risk and keep the program governable

Cloud providers release new capabilities continuously, and AI services evolve even faster. Governance should include a way to identify material new features, evaluate changed responsibility boundaries, and decide whether existing controls still apply before adoption becomes widespread.

Risk indicators might include privileged-role growth, public exposure, unapproved services, data-policy exceptions, model-quality degradation, supplier incidents, concentration, unsupported components, or exceptions approaching expiry. The selection should match the scenarios that matter to the business.

CISM is relevant because security managers must translate technical change into program priorities, resources, and executive decisions. Risk reporting should explain what changed, why the business should care, what treatment is underway, and what residual exposure remains.

Strong technology risk management does not try to slow every cloud or AI initiative. It creates enough visibility, ownership, evidence, and decision discipline that the organization can innovate without losing track of what it depends on, what could fail, and who is accountable when conditions change.

Risk review should distinguish experimentation from scaled production. Sandboxed pilots can allow rapid learning with limited data and authority, while moving into production should trigger stronger requirements for monitoring, resilience, security testing, data governance, and ownership. This staged model supports innovation without treating every prototype as enterprise infrastructure.

Retirement is part of technology risk as well. Models, services, APIs, and cloud features can become unsupported or strategically obsolete. Programs should track end-of-life dates, deprecation notices, migration effort, and data-retention obligations so dependency does not become a crisis when a provider withdraws a capability.

Filed under Cloud Computing