Accéder au contenu

Nouveau : outils IA gratuits — X-Ray votre site web ou obtenez un blueprint IA en 60 secondes.

ASTACKRA
Lancer un projet

ASTACKRA Insights

Data Privacy and Security in AI Systems: What SOC 2 and HIPAA Mean for Your Build

By ASTACKRA 6 min read

Published 3 October 2026

Every conversation about an AI build eventually turns to compliance, and it usually happens later than it should — after architecture decisions are made, sometimes after a vendor contract is signed. Security and privacy requirements have real implications for how an AI system has to be built, not just whether it’s allowed to be deployed, and treating them as a checklist to satisfy at the end produces systems that need expensive rework once legal or a customer’s security team actually reviews the design. SOC 2 and HIPAA come up constantly in these conversations, and they get confused with each other often enough that it’s worth being precise about what each one actually is.

SOC 2 Is an Audit of Your Controls, Not a Law

SOC 2 isn’t a regulation — nobody is legally required to have it. It’s a voluntary attestation, produced by an independent auditor, that your organization has specific controls in place around security, availability, processing integrity, confidentiality, and privacy, depending on which Trust Services Criteria you scope the audit to cover. Most B2B software companies pursue it because enterprise customers increasingly require it before they’ll sign a contract — it’s become a de facto market requirement even though it’s formally optional.

A SOC 2 Type I report attests that your controls are designed appropriately as of a specific point in time. A SOC 2 Type II report, which carries considerably more weight with sophisticated buyers, attests that those controls actually operated effectively over a review period, typically six to twelve months. Type II is what most enterprise customers mean when they ask if you “have SOC 2,” and it requires sustained operational discipline, not a one-time documentation exercise — you can’t retroactively produce a Type II report after the fact, because it certifies behavior over a period that has to have already happened under those controls.

What SOC 2 Actually Requires of an AI System

For an AI build specifically, SOC 2 compliance touches several areas that have direct architectural consequences. Access controls need to extend to whoever and whatever can reach the model, the training or fine-tuning data, and any stored outputs — including service accounts and the AI system’s own access to internal systems, which is easy to overlook when the focus is on human user access. Logging and monitoring need to cover not just application behavior but the AI-specific actions: what data was retrieved for a given response, what tools an agent called, what decisions it made on its own versus what required human approval.

Data handling policies need to account for where training or reference data lives, how long it’s retained, and under what conditions it can be deleted — a question that gets genuinely complicated when a model has been fine-tuned on sensitive data, since “deleting” a record from a database is straightforward but ensuring a fine-tuned model hasn’t memorized that record is a much harder technical problem that SOC 2 auditors are increasingly asking about directly.

HIPAA Is a Federal Law With Real Legal Exposure

HIPAA is categorically different from SOC 2: it’s U.S. federal law, it applies specifically to protected health information, and noncompliance carries real legal and financial consequences rather than just a lost sales opportunity. If an AI system touches PHI in any way — processing patient records, assisting with clinical documentation, handling insurance or billing information tied to an identifiable patient — HIPAA’s requirements apply regardless of whether the organization building the system is a healthcare provider itself or a vendor serving one.

The practical requirement most relevant to AI builds is the Business Associate Agreement: any third-party vendor that touches PHI on behalf of a covered entity needs a signed BAA in place, which creates real constraints on which AI infrastructure you can use. Not every API provider or cloud service offers a BAA, and using one that doesn’t — even for a seemingly minor piece of the pipeline, like a logging service or a third-party embedding model — creates a compliance gap, no matter how solid the rest of the architecture is. This is a common failure point: a team builds a compliant-looking system around a core AI provider that does offer a BAA, then pipes PHI through an auxiliary service that doesn’t, and nobody notices until a security review catches it.

Where the Two Overlap and Where They Diverge

SOC 2 and HIPAA share enough conceptual ground — access controls, encryption, audit logging, incident response — that organizations handling PHI often pursue both together, and much of the underlying infrastructure work serves both. But they diverge in important ways. SOC 2 is about demonstrating your controls are sound to whoever’s asking, commercially. HIPAA is about a specific set of federal legal obligations tied to a specific category of data, enforced with actual penalties, including for willful neglect that goes beyond the technical requirements into questions of whether reasonable diligence was exercised at all.

An organization can be fully HIPAA compliant without ever pursuing SOC 2, if it never deals with enterprise buyers who require the attestation. And an organization can hold a clean SOC 2 Type II report while still being out of HIPAA compliance, if it’s handling PHI without the specific safeguards and BAAs HIPAA requires — SOC 2’s general security controls don’t automatically satisfy HIPAA’s more specific requirements around PHI handling.

What This Means for Architecture, Concretely

The architectural implications show up in a handful of specific places. Data residency and processing location matter — if PHI is involved, every service in the pipeline needs a BAA, which rules out some otherwise-attractive tooling choices and should be checked before the stack is chosen, not after. Retention policies need to be designed in from the start, including for any logs, cached responses, or fine-tuning data that might contain regulated information, since retroactively scrubbing data from a system that wasn’t built to track where sensitive information ended up is a much harder project than designing retention correctly the first time.

Access logging needs to capture AI-specific events, not just traditional application access — which records a given agent retrieved, what it did with that information, and who approved any action it took, at a level of granularity that satisfies both a SOC 2 auditor’s questions about system behavior and HIPAA’s requirements around PHI access tracking. And encryption requirements extend through the full pipeline — in transit, at rest, and often within the processing environment itself depending on how sensitive the data is — which has real performance and architecture implications for how a retrieval system indexes and queries sensitive documents.

Compliance as a Design Constraint, Not a Final Review

The expensive version of this is discovering compliance requirements after a system is built, when the fix means re-architecting which services touch sensitive data, renegotiating vendor contracts to get BAAs in place, or rebuilding logging and retention systems that were never designed to support an audit. The cheap version is treating compliance requirements as design constraints from the first architecture conversation, the same way you’d treat a performance requirement or a budget constraint — something that shapes the build from day one rather than something checked after the fact.

This doesn’t mean every AI project needs SOC 2 or HIPAA compliance baked in regardless of use case — a lot of internal tooling and low-stakes automation genuinely doesn’t touch regulated data and doesn’t need this level of rigor. But knowing which category your project falls into, and designing accordingly from the start, is a decision worth making explicitly and early rather than assuming either way.

We’ve documented our own approach to these questions in detail on our trust center, and if you’re scoping a build that touches sensitive data and want the compliance requirements factored into the architecture from the beginning, start a project conversation and we’ll walk through what your specific data and regulatory situation actually requires.

Related

Continuer la lecture

Tous les éclairages

Étape suivante

Dites-nous ce qui ralentit votre entreprise.

Décrivez le workflow, le site web, le parcours client ou le système que votre équipe a dépassé. Vous n’avez pas besoin d’un cahier des charges technique — nous définirons avec vous la bonne première phase.

Lancer un projet hello@astackra.com
  • Livraison remote-first sur plusieurs fuseaux horaires
  • Périmètre, jalons et décisions écrits
  • AI contrôlée par l’humain, compatible NDA

Studio remote-first d’IA, de software et d’automatisation — cadrage, conception et livraison pour des équipes du monde entier.

Nous concevons des systèmes d’IA et des logiciels sur mesure qui automatisent les opérations, relient les équipes et créent un levier business durable.

Systèmes d’IA, logiciels sur mesure, SaaS, automatisation des workflows, intelligence documentaire et ingénierie de produits digitaux pour des entreprises en croissance partout dans le monde.

Une technologie complexe. Une exécution élégante.

ASTACKRA · Studio de systèmes & software