AI systems can now produce compliance artifacts that are structurally complete, terminologically correct, and by most surface-level assessments indistinguishable from what a trained analyst would produce. TARAs, safety cases, traceability matrices, hazard analyses. All of them.
This should be good news. Instead, it surfaces a question that was easier to ignore when humans were doing the work: what, exactly, is being verified?
The discomfort is justified. But it points in the wrong direction. The problem is not that AI generates compliance documentation. The problem is that AI makes visible a failure mode that was already structural: compliance processes that optimize for artifacts, not for the properties those artifacts were supposed to encode.
A traceability link created the day before an audit is structurally identical to one created during development. The matrix reads green either way. The audit may pass. But a link created to satisfy a KPI does nothing for the safety or security of the product. The artifact and the reality it was meant to reflect have decoupled.
Compliance theater predates AI by decades. AI doesn’t introduce this failure mode. It removes the friction that made it tolerable.
The Artifact Is Not the Argument
Under ISO 21434, a Threat Analysis and Risk Assessment is not valuable because it documents what someone thought about security. It is valuable because it encodes a demonstrable property: given these assets, these threat agents, these attack paths, these mitigations: the residual risk meets the defined threshold. The document is the carrier. The property is what matters.
The same logic applies to traceability. A traceability matrix does not prove requirements coverage by existing. It proves coverage because the links it encodes are formally consistent with the artifacts they reference, and can be verified as such. Remove that formal consistency, and the matrix is a spreadsheet with checkboxes.
Human-generated compliance artifacts always had this problem, because the conditions for maintaining them were never favorable. Keeping traceability synchronized with actual development state is expensive and produces nothing the customer can ship. The implicit prioritization is rational: the product ships, the documentation catches up, the audit gets managed.
The practical goal was not to build a provably safe product. It was to convince an external reviewer that you did.
AI changes the economics, not the incentive structure. Compliance documentation that previously took weeks can now be produced in hours. Volume turns out to be the most effective defense against scrutiny, not because reviewers are careless, but because inconsistencies in a large artifact set become genuinely difficult to locate.
Volume is not evidence. It is cover.
When we introduced automation features in itemis ANALYZE, a traceability tool built for compliance in regulated domains, the usage pattern was immediate and unambiguous. Customers used automation as a replacement for manual traceability work from the first week. Nobody needed to be convinced. The incentive to eliminate compliance overhead is absolute, and it expresses itself the moment an alternative exists.
The Closed Loop
The obvious response is to use AI for the review. Which means AI checking AI-generated documentation. Internal consistency is proven, nothing else. An auditor who accepts this output is not verifying safety or security. They are verifying that the artifacts agree with each other.
That is not compliance. That is a closed loop with a rubber stamp at the end.
The problem is not that AI is involved on both sides. It is that both sides operate on the same thing: documentation that has been decoupled from the actual product. A better reviewer does not fix this. A smarter generator does not either. The closed loop is a structural failure, and it requires a structural fix.
An agent that drives and monitors a defined process is a different thing. It does not generate compliance output. It executes a process: the steps, the rules, the required outputs at each stage, in the right order, with the right inputs, verifying that the formal properties hold throughout. Human involvement concentrates on two things: defining the process, and specifying what the output must demonstrate.
The artifacts that emerge from this are not descriptions of compliance. They are evidence that the process ran.
KPIs and metrics become what they were always supposed to be: not numbers to adjust before an audit, but proof that the process was executed correctly. The auditor’s role doesn’t disappear. It changes: from reading documents to inspecting process execution. From “does this look right?” to “did the process run, and does the output reflect it?” That is a more rigorous standard, not a weaker one.
The document is the byproduct. The process is the proof.
Novel threat scenarios, domain-specific risk judgments, decisions about acceptable residual risk: these remain human responsibilities. What an agent removes is the overhead that currently prevents engineering judgment from being applied where it actually matters. That is not a reduction in rigor. It is the condition for rigor becoming possible.
For a vendor building compliance and traceability tooling, this shift does not change the product roadmap. It changes what the product fundamentally is.
A tool that helps engineers manage compliance is a productivity surface. Its value is measured in time saved, errors avoided, workflows streamlined. When agents become the primary operator, that value proposition disappears. Agents don’t need a better interface.
What agents need is a formal model they can execute against. A precise definition of what compliance requires, encoded in a structure that is deterministic, auditable, and complete enough to run a process on. The graph model, the coverage logic, the semantic consistency checks: these are not features that accelerate human analysts. They are the specification of what correct compliance looks like.
The distinction between a traceability graph that stores links and one that encodes governed semantics is precisely what determines whether an agent can reason over it. A graph without explicit, versioned semantics is not a substrate an agent can execute against. It is a reporting artifact.
A vendor who has formalized that knowledge owns the process. A vendor who built an interface does not.
Traceability was always the connective tissue between the product and its compliance obligations. Maintained by humans under time pressure, it decoupled. Maintained by agents running a defined process, it reflects the actual state of the product continuously, not just before an audit.
That is the purpose compliance documentation was designed to serve. The question is whether the underlying model is formal enough to support it.
Requirements traceability in practice: How end-to-end traceability works in regulated projects, which tools have proven themselves, and when the RTM approach reaches its limits: Requirements Traceability at itemis →