Governance · 4 minute read
SOC 2 for AI Products: What Changes When Models Are Involved
SOC 2 for AI products applies the AICPA Trust Services Criteria, security, availability, processing integrity, confidentiality, and privacy, to the components AI introduces: model providers as subservice organizations, prompts and retrieval as data flows, agents as privileged non-human identities, and evaluation as the evidence of processing integrity, with controls that operate consistently across a Type II observation period.
Customers of AI products ask for a SOC 2 report, and vendors discover that the report's criteria say nothing about models, prompts, or agents. Auditors nevertheless ask about all of them, because the Trust Services Criteria apply to whatever a system contains. This guide maps the criteria to AI components and lists the evidence that satisfies them, so preparation is deliberate rather than reactive. It complements the AI security checklist and the vendor perspective in the AI vendor due diligence whitepaper. It is general guidance, not audit or legal advice; engage your auditor early.
How do the criteria map to AI components?
| Criterion | AI component | What auditors examine |
|---|---|---|
| Security | Agent identities, tool permissions, gateway, sandboxing, provider access | Provisioning, least privilege, reviews, revocation, logging |
| Availability | Model routing, fallbacks, provider incidents, kill switch | Redundancy, monitoring, incident response, tested failover |
| Processing integrity | Prompts, retrieval, evaluation, change gates | Specifications, evaluation evidence, regression gates, monitoring |
| Confidentiality | What reaches models and logs; provider terms; redaction | Classification, redaction, retention, contracts |
| Privacy | Personal data in prompts, retrieval, traces | Minimization, consent basis, redaction, access, retention |
How are model providers handled?
As subservice organizations. Obtain their assurance reports, review the scope for the services you use, map the complementary user-entity controls they expect you to operate (what you send, retention configuration, access to your account), and document vendor management: due diligence, contract terms on training and retention, and monitoring for changes. Self-hosting removes the subservice relationship and adds your own availability and security controls to scope. The choice is discussed in the private AI for regulated industries whitepaper.
What does security evidence look like for agents?
Agents are privileged non-human identities. Evidence: the agent registry with owners and permissions; provisioning records showing scoped principals; periodic access reviews that include agents; revocation and kill-switch test records; gateway logs tying actions to agent and delegated user; sandbox escape-test results. The model is in non-human identities for AI agents and AI agent kill switch design.
Where does evaluation become audit evidence?
Processing integrity asks whether the system processes completely, accurately, and timely as authorized. For AI components that means: specifications defining intended behavior; regression-gate records showing every prompt, model, and retrieval change was evaluated before release; production sampling results; and monitoring with alerts. The discipline is AI regression testing, and its records are the evidence.
How are confidentiality and privacy evidenced?
Data classification; gateway policies routing categories to approved deployments; redaction configuration at prompt assembly and log storage; retention settings; access control on traces; provider contract terms; and, for privacy, the lawful basis and minimization design. The controls are set out in LLM data loss prevention and how to build a PII redaction pipeline.
What about availability?
Routing with tested fallbacks, provider-incident runbooks, the kill switch with rerouting, and monitoring on model and tool latency and errors. Evidence: routing configuration, failover test records, incident records. The platform side is in the LLM gateway architecture whitepaper.
What does the evidence set look like?
| Control | Evidence produced automatically |
|---|---|
| Agent access provisioning | Registry entries with scopes, owner, and approval record |
| Periodic access review | Quarterly attestation including every agent identity |
| Change management for prompts and models | Version history, regression-gate results, approver, release record |
| Processing integrity monitoring | Sampling scores, alert history, review-queue outcomes |
| Confidentiality and privacy | Gateway policy configuration, redaction test results, retention settings |
| Availability | Failover test logs, provider-incident records, kill-switch test timings |
| Incident management | Postmortems with actions and closure |
| Vendor management | Provider reports, complementary control mapping, annual review |
Where evidence is produced by the platform as a side effect of operation, the audit is a reporting exercise. Where it must be assembled by hand, it is a scramble.
How do you prepare for a Type II?
- Inventory AI components and data flows; keep it current in the registry.
- Map each to criteria and name the control and its owner.
- Embed controls in the platform so they operate without reminders: gateway policy, regression gates, brokered credentials, scheduled reviews.
- Capture evidence continuously: logs, gate records, attestations, test results, postmortems.
- Run a readiness assessment against the mapping before the observation period.
- Keep change management disciplined during the period; every prompt change is a change.
What are the common gaps?
- Agents missing from access reviews.
- Prompt changes outside change management.
- No evaluation evidence, so processing integrity is asserted.
- Traces storing raw sensitive data.
- Provider complementary controls unmapped.
- Manual controls that lapse mid-period.
How does FISTA Solutions help?
FISTA Solutions builds AI products and platforms with the controls and evidence capture SOC 2 examines designed in, gateway policy, agent identities, regression gates, redaction, and traces, as part of its AI enablement practice, and helps product teams map components to criteria through forward deployed engineers working with security and compliance. Every AI agent FISTA delivers produces the audit evidence by default. FISTA has delivered 150+ projects for 50+ companies across 12+ countries with 99.9% uptime.
To run a readiness review for your AI product, message FISTA on WhatsApp, or read enterprise AI security for the control baseline.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Does SOC 2 have specific requirements for AI?
No. The Trust Services Criteria are technology-neutral. Auditors apply them to whatever the system includes, and an AI product includes model providers, prompts and retrieval, agents with system access, and evaluation processes. The work is mapping those components to the criteria and producing evidence that the controls operate.
02How are model providers treated?
As subservice organizations, like a cloud host. You obtain and review their assurance reports, define complementary user-entity controls you are responsible for, such as what data you send and how you configure retention, and document the vendor management process. Self-hosted models remove the subservice organization and add your own operational controls.
03What evidence do auditors ask for on AI components?
Inventories of models, agents, and data flows; access provisioning and review records for agent identities; change management records with evaluation gates for prompt and model changes; evaluation and monitoring results as processing-integrity evidence; redaction and retention configuration; incident records; and vendor assurance reports with complementary control mapping.
04How do you prepare for a Type II?
Make the controls operate automatically and capture evidence continuously: gateway logs, regression-gate records, access-review attestations, and incident postmortems, over the whole observation period. Controls that depend on someone remembering to run them fail Type II; controls embedded in the platform pass. This is general guidance, not audit advice.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.