AI Airlocks
Validate a proposed action before granting it authority to change anything.
Context.
A support assistant proposes closing a customer case. It has read the conversation and produced a plausible summary. Closing the case changes a shared record and may trigger a message to the customer. The proposal and those effects belong on opposite sides of a boundary.
The same boundary matters when receiving retrieved documents, accepting another service's result, or releasing generated text. Each is an input to somebody else's system. Treat it according to the authority it will receive, even when it came from your own infrastructure.
Problem.
A well-formed answer can still be wrong, unauthorised, or stale. Checking JSON does not establish that the customer agreed to close the case. Asking a model to be careful does not remove its database credentials. Recording the damage afterwards does not prevent it.
The difficult question is where validation becomes binding. A check in an orchestration library is easy to bypass. A check before a long approval queue can be obsolete by the time the action runs. The service that owns the effect must refuse work that has not passed the required checks.
Forces.
- Structural validity, domain correctness, suspicious behaviour, and permission are separate questions. Passing one does not answer the others.
- Human review and model review consume time. A deadline cannot silently turn an unanswered question into approval.
- Services fail between accepting a command and reporting its result. Retrying must not repeat the effect.
- Domain teams know their business rules; shared governance sets constraints those teams cannot weaken.
- Evidence and approvals must survive an incident without exposing every sensitive input to every operator.
Pattern.
Place an airlock at each trust boundary and immediately before each governed effect. It takes a proposal, trusted context, and a versioned policy. It returns an explicit verdict: allow, reject, or pending review. Only allow can proceed. A validated State can record that review is pending; it cannot imply permission to execute.
Use the five layers from the Foundation canon. Schema validation checks shape and required fields. Semantic validation checks domain relationships against authoritative records. Heuristics flag suspicious patterns. Optional model re-evaluation supplies another assessment where useful. Mutation gating checks authority for the exact operation about to run.
Keep their responsibilities visible. A second model can miss the same error as the first. A keyword rule can miss an equivalent destructive operation. Neither replaces the capability check at the mutation service. Destructive actions require human authorisation; ordinary permitted actions may receive explicit approval from policy.
This descends from boundary input validation, type checking, and explicit control of side effects. The OWASP Input Validation guidance distinguishes syntactic checks from semantic checks. An airlock applies that distinction to AI proposals and adds the authority needed to carry them out. Runtime validation remains necessary when static types end at a network boundary.
Implementation.
Consider a policy for closing support cases. This illustrative YAML describes decisions an implementation must enforce; it is not an executable security product. The timing values are starting budgets for this example, to be measured under load.
policy: support.case-close
version: 3
owner: support-operations
input_contract: CloseCaseProposal.v2
schema:
reject_unknown_fields: true
semantic:
require: [case_exists, customer_resolution_evidence]
expected_revision: required
heuristics:
flags: [bulk_closure, missing_customer_reply]
on_flag: human_review
model_review:
enabled: false
mutation:
capability: case.close
approval: explicit_policy_or_human
destructive_actions: human_only
bind_to: [tenant, actor, action, target, payload_digest]
bind_versions: [case_revision, policy_version]
approval_ttl_seconds: 120
idempotency_key: required
execution:
validator_deadline_ms: 150
on_timeout: pending_review
on_dependency_failure: pending_review
recheck_authority_at_commit: true
persist_intent_before_effect: true
record_receipt: true
The assistant submits case 8472 at revision 19, the resolution evidence, and a proposed closing summary. The airlock parses the proposal, loads the case through a trusted reader, and checks that the evidence belongs to this case. Any provenance identifier supplied by the model must resolve to an authorised source; inventing a convincing identifier earns no trust.
If a human is needed, store the candidate separately and return pending review. Present the reviewer with the exact target, evidence, and proposed effect. After approval, issue a short-lived authorisation bound to those values. A changed summary, target, or case revision requires fresh validation. Authenticate the approver and verify the authorisation server-side; a model-supplied approval field is just another claim.
At commit, the case service verifies current authority and compares revision 19 inside the same transaction that closes the case. It stores the idempotency key and a durable result with that change. Duplicate delivery returns the existing result. If the response is lost, the caller asks for that result instead of inventing a new key.
For an external effect, persist the authorised intent before dispatch and reconcile uncertain results using the receiver's receipt or idempotency mechanism. If neither exists, stop automatic retries and ask an operator to resolve the uncertainty. A compensating action is a new authorised action; it cannot unsend a message.
Record candidate, verdict, policy version, approval, and receipt as separate lineage events. Required checks finish before release. Optional quality analysis may run later, but it cannot retroactively make an unsafe release safe. If mandatory evidence cannot be durably recorded, keep the proposal pending.
Consequences.
The system gains a place to explain why an action was allowed and a place to stop it. Domain failures become explicit outcomes instead of accidental writes. A compromised model loses the ability to grant itself permission, provided every route to the effect uses the boundary.
The costs are real. Trusted reads, durable records, and authorisation checks add latency even without another model call. Review queues need staffing and expiry rules. Candidate retention and validation evidence consume storage. Validators need owners, migrations, and incident support; false alarms create organisational friction when teams disagree about acceptable risk.
Measure queue age, rejection reasons, duplicate suppression, and time spent in each layer. Reserve capacity for review and bound retries. Airlocks reduce particular failure paths; they do not certify truth, remove all malicious behaviour, or repair an insecure execution environment.
Anti-patterns.
The decorative checkpoint. Validation runs in a helper while another credential can write directly. Put enforcement at the resource owner and remove the bypass.
Approval by exhaustion. A timeout, unavailable reviewer, or exhausted retry budget becomes allow. Preserve pending or reject, and expose the operational failure.
The reusable blessing. An approval authorises any later payload for the same case. Bind it to the exact proposal and relevant versions, then check again at commit.
Test.
Submit a valid proposal, change the target record during approval, replay the authorised request twice, and interrupt the validator. The stale proposal must stop, the duplicate must produce at most one effect, and the interruption must produce no unauthorised effect. For every result, retrieve the policy, evidence, verdict, and execution receipt without relying on the model's account.
Related.
AI Airlocks supply the governed checkpoints in Vertical AI Foundation and address the failure paths discussed in The AI Mainframe Trap. They validate the return path from Control Plane vs. Execution Plane, record their decisions in Causal Lineage, and remain mandatory within a Thin Boundary.