Airlock checks what an AI agent sends to the cloud before it leaves your machine. Why it exists, how it works, and how it is measured.
When an AI agent works on your private documents, what actually leaves your machine?
More than the prompt. Every planning turn re-sends what the agent has read so far. Tool calls carry arguments, and those arguments are replayed in the history. Search queries go to a search API, usually a different company from your model provider.
Airlock is an egress firewall for AI agents. It checks every hop locally before it leaves and restores the answer on the device. This is Part 1 of the series: why I built it, what it does, how it works, and how it is measured. Part 2 has the measured results. This post was written as part of a Nebius × NVIDIA Global AI Hackathon entry.
The problem: more exits than the prompt

People and companies avoid cloud AI mainly because of leakage: names, customer data, health details and credentials pasted into a prompt. Local-only models avoid the leak, but they are much weaker than frontier models.
Agents widen the problem. Documents the agent reads go back to the cloud model on every planning turn, and the agent's own tool calls are sent again as part of that history. Search traffic goes somewhere else entirely. Most privacy tooling sits between you and the model. The search engine and the tools an agent calls usually sit outside that line.
Short queries are not automatically safe. A query like <ORG_1> layoff list team lead how to respond contains no name, and it still tells the search provider that someone at a company is on a layoff list.
What already exists, and what does not
Masking a prompt and restoring the answer is a solved problem, and I want to be clear about that up front.
- PasteGuard is an OpenAI-compatible proxy that uses GLiNER, restores answers including streams, and supports coding agents.
- Kiji Privacy Proxy from Dataiku masks with a quantized DistilBERT model and restores the answer.
- LiteLLM's Presidio guardrail does the same job at the gateway.
- PAPILLON (NAACL 2025) already uses a local model to rewrite queries before they reach an API model, and it measures leakage.
So splitting the work between a small local model and a large cloud model is not a new idea either.
The gap is in the other channels. MosaicLeaks (Gurung et al., 2026) shows that an observer can infer private document contents from a deep-research agent's search queries alone. Its fix is retraining the agent, which a user cannot put in front of an existing one. AWS Bedrock Guardrails documents that it does not inspect tool-call arguments or tool results. The READMEs of the open-source proxies above do not mention scanning tool calls or search queries. I read their documentation, not their code.
I also did not find a product that publishes a leak rate measured on the bytes that actually left the machine, repeated across runs, next to baselines. Research papers do measure leakage, so Airlock is not the first to measure it. The narrower aim is a product-level measurement with an open harness, against baselines run under the same conditions.
The thesis: identity versus situation
The goal is narrow on purpose. The cloud can know the problem. It shouldn't know whose problem it is.
That splits private information into two kinds. Identity is names, organizations, phone and account numbers, secrets, and quasi-identifiers: attributes that each look harmless but together single a person out. Situation is what the model needs to read in order to help: a diagnosis, a severance clause, an amount. Hide the situation too and the answer stops being useful.
So Airlock removes or generalizes identity and tries to keep the situation. At the default balanced level, amounts, lab values and diagnoses stay in the prompt, while a city like "Tulsa" becomes "a city in Oklahoma".
Using it means changing one base_url. Point any OpenAI-compatible app at http://127.0.0.1:8787/v1, and your scripts, notebooks, editors and agents go through Airlock.
How it works

A chat request passes five local steps before anything leaves.

- Detect. Regex, entropy and vault matching find secrets and identifiers first and mask them in the copy the local model reads, so the model never receives raw secrets. NVIDIA Nemotron-3-Nano-4B then proposes semantic spans (names, organizations, health details, quasi-identifiers) with JSON-schema-constrained output, and deterministic Korean rules back it up where the small model is weak. A model span whose text does not occur in the original is discarded, so a hallucinated ID never becomes a redaction.
- Substitute. Identity values become typed placeholders such as
<PERSON_1>in a local vault, and quasi-identifiers get a generalization. The same original always maps to the same placeholder within a conversation or agent run. - Gate. The final outbound JSON is walked string by string. If a vault original, a declared term, a canary or a high-confidence secret pattern is still present, the request is blocked with HTTP 422.
- Upstream. The exact bytes that passed the gate go to Nemotron 3 Ultra on Nebius Token Factory.
- Rehydrate. Placeholders in the answer, including tool-call arguments and streamed deltas, are mapped back to the originals on your machine.
Every request leaves an audit record with the outbound payload verbatim and detections as keyed hashes, so you can check afterwards exactly what the cloud saw.
Optionally, NVIDIA GLiNER-PII runs as a second local proposer. It has high recall and over-masks, most of all in Korean, so a span it finds alone has to pass a deterministic type check and an adjudication call to the local Nano model before it counts.
Why the gate is code
The local model proposes. It never decides. Nemotron-3-Nano-4B samples at temperature 0.6, so the same input can produce different spans on different runs. Writing "never send secrets" into its prompt shifts what it tends to do and guarantees nothing. A prompt is a prior, not a gate. A model never gets to say "this is fine."
The gate is plain code on the final bytes, and "present" covers variants: letter case, full-width and look-alike letters, separated Hangul jamo, inserted spaces, numbers spelled out in Korean or English, and values hidden in JSON strings, percent-encoding or base64 up to two layers deep. Content Airlock cannot inspect, such as images, is blocked. If the local model is down, times out or returns invalid JSON, the request is blocked too. Airlock fails closed.
The guarantee is narrow and checkable. A misbehaving local model can cause over-masking or a block. It cannot cause a string Airlock knows is sensitive to be sent.
Agent mode and search intent

In agent mode, every hop that leaves the machine is treated as an egress event, gated and audited.
| Hop | Destination | What happens before it leaves |
|---|---|---|
| Planning turn | Nemotron 3 Ultra | The whole history, including documents read and earlier tool calls, goes through the detector, the vault and the gate |
| Search query | Tavily | The search-intent guard: rewrite locally, gate, judge, retry once, then block that one search |
| Local tools | Nowhere | Reading documents runs on your machine |
| Final answer | Nowhere | Rehydrated locally |
Documents are shown to the model under neutral ids such as doc-1, because file names often describe the situation.
For search queries, the risk is rarely a string the gate already knows. It is intent. So Nemotron-3-Nano-4B rewrites each query toward a generic, information-seeking one, given the user's question and the documents read so far as the thing it must not reveal. In a live run, the agent asked for Tessellate Health AI ambient clinical documentation Denver competitors funding valuation, and Tavily received ambient clinical documentation market valuation competitors.
The rewritten query then passes the gate, and an intent judge decides whether it still reveals a private situation. The default judge is the same local Nano model; NVIDIA Nemotron 3.5 Content Safety can run instead in custom-policy mode, or both together. A rejected rewrite is retried once. If that fails too, only that search is withheld, and the agent is told and continues. A judge that errors counts as a flag.
Which NVIDIA and Nebius pieces do what
The design in one line: the small local model does not solve the task. It only decides what is private. The big cloud model does the reasoning.
| Role | Component | Where it runs | Why |
|---|---|---|---|
| Detector, search rewriter, default intent judge | NVIDIA Nemotron-3-Nano-4B | Your machine, llama.cpp | Small enough for a laptop, follows a JSON schema reliably, and sees private text, so it stays local |
| Second PII proposer (optional) | NVIDIA GLiNER-PII | Your machine | High recall; the PII backend of NeMo Guardrails |
| Reasoning over the sanitized request | NVIDIA Nemotron 3 Ultra | Nebius Token Factory | Frontier-class open model that keeps placeholders intact, including in Korean |
| Search intent judge (optional) | NVIDIA Nemotron 3.5 Content Safety | Your machine | Custom-policy judgment on search queries |
| Web search | Tavily | Cloud API | Receives only queries that were rewritten and passed the gate |
Token Factory is on the execution path of every allowed request, through an OpenAI-compatible API, and Nemotron 3 Super is used once as a fallback if Ultra is unavailable. Token Factory may store prompts and outputs unless Zero Data Retention is enabled, and I recommend enabling it. Airlock's guarantee does not depend on it: whatever the provider keeps is the sanitized payload shown in the audit log.
How it is measured

I did not want the thesis to stay a claim, so measurement was part of the design from the start.
- Dataset: 243 synthetic cases, 122 Korean and 121 English. Every value is invented.
- Baselines: raw pass-through, regex only, Presidio with Korean recognizers, and NVIDIA GLiNER-PII alone, in the same harness.
- Attacker: Nemotron 3 Ultra reads only the outbound payloads and tries to recover identity items and the private situation.
- Main metric: linkable disclosure, the share of the 98 situation-sensitive cases where the attacker recovers an identity item and also infers the private situation. It counts the cases where an observer could say who has which problem.
- Usefulness: a filter that hides everything is easy to build and useless to work with, so a judge scores usefulness and distortion against reference answers, next to over-redaction and latency.
- Repetition: the detector samples, so one run is not evidence. Every system runs 3 passes, the server is reset before every pass, and per-pass detect time confirms the passes were independent.
- Agent scenarios: 16 fictional private-document scenarios, 8 Korean and 8 English, run with an unguarded agent and with Airlock. The attacker reads all hops once and the Tavily queries alone once, the MosaicLeaks threat model.
The measurement is done. Across all 243 cases, Airlock cuts linkable disclosure from 84.7% to 13.9% against raw pass-through. What that costs, what still links, and a mistake I made in the repeated runs are in Part 2.
Limitations and threat model
Airlock is designed to stop accidental disclosure of identifying or secret strings to a cloud LLM provider or a web search provider, through an app you control. It does not cover everything.
- Detector misses. A value every detector misses, and that you did not declare, is sent. The gate only enforces what is known.
- Meaning and situation. Airlock removes strings, not meaning. A rare combination of attributes can still identify someone, and the provider can still learn that someone has this situation.
- Linked queries. Several generic queries in a row can suggest the topic, and Tavily sees timing and your network address. In my own agent scenarios, unguarded search queries were already fairly generic, so the search guard closed a small leak, not a large one.
- Inbound content. Prompt injection in search results and documents is not checked, though an injected instruction still cannot get known strings past the gate.
- Your machine. The vault stores originals in plain text locally, and local detection adds seconds to each request.
Try it
The code and the evaluation harness are open source under Apache 2.0 at github.com/tristan-kkim/airlock.
git clone https://github.com/tristan-kkim/airlock && cd airlock
uv sync && scripts/local_model/serve.sh
uv run airlock serve
The local UI shows what you typed, what the cloud saw, and the rehydrated answer side by side. A hosted demo with fictional scenarios goes live in October and stays free and open until judging ends on December 15. Its detection runs on the demo server, so do not paste real personal data there; for real privacy, run Airlock locally.
Next: the measured results
Part 2, First Measured Results from Airlock, covers the baseline comparison, the trade-off, what Korean costs, and the passes that turned out not to be independent.
Founder of Cortexys, leading AX consulting, corporate AI training, and AI agent development.
Need an AI solution?
See custom AI development services at cortexys.team.

