We measured Airlock, an egress firewall for AI agents, on 243 synthetic cases in 3 independent passes. Linkable disclosure fell from 84.7% to 13.9%.
The numbers in this post were updated in part 3 after further hardening: https://blog.cortexys.team/airlock-how-the-numbers-improved-en
This is Part 2 of the Airlock series. Part 1, Airlock: Egress Firewall for AI Agents, covers why I built it and how it works.
When an AI agent works on your private documents, the prompt is only one of the things that leave your machine. Every planning turn re-sends what the agent has read, tool calls carry arguments, and search queries go to a search API. Airlock is an egress firewall for AI agents that sits on those exits. I built it for the Nebius × NVIDIA Global AI Hackathon, and this post covers its first measured results, including a mistake I made along the way.
Before a request leaves the machine, regex and entropy detectors, NVIDIA Nemotron-3-Nano-4B, Korean rules and an NVIDIA GLiNER-PII ensemble find identifiers and quasi-identifiers locally. A local vault swaps them for placeholders such as <PERSON_1>, and a deterministic gate written in plain code checks the exact outbound bytes and blocks the request if a known sensitive string is still there. Only what passes reaches Nemotron 3 Ultra on Nebius Token Factory, and the answer is restored on the device.
Masking proxies are not new: PasteGuard, Kiji and LiteLLM with Presidio already mask and restore. What I did not find in the products I surveyed was a leak rate measured on the bytes that actually left the machine, repeated across runs, next to baselines run under the same conditions.
How it was measured
- Dataset: 243 synthetic cases, 122 Korean and 121 English. Every value is invented.
- Repetition: every system ran 3 passes. For Airlock, the server was reset before every pass.
- Attacker: Nemotron 3 Ultra reads only the outbound payloads and tries to recover identity items and the private situation.
- Main metric: linkable disclosure is the share of the 98 situation-sensitive cases (finance, health, quasi-identifier, intent search) where the attacker recovers at least one identity item and also infers the private situation. It could say who has what.
- Usefulness and distortion: a Nemotron 3 Ultra judge compares each answer with a reference answer from the raw prompt.
- Baselines: raw pass-through, regex only, Presidio with Korean recognizers, and NVIDIA GLiNER-PII alone, the PII backend of NeMo Guardrails. Each system ran alone on an Apple M3 Pro.
Results and the trade-off
Airlock numbers are the mean and sample standard deviation over 3 independent passes. Baseline outputs were identical across passes and were scored once.
| System | Linkable disclosure | Identity leak | Usefulness 1-5 | Utility ratio | Distortion | Benign masked |
|---|---|---|---|---|---|---|
| raw (pass-through) | 84.7% | 98.5% | 4.52 | 1.05 | 7.7% | 0.0% |
| regex only | 79.6% | 92.7% | 3.91 | 0.87 | 13.9% | 7.4% |
| Presidio + ko/en spaCy + KR recognizers | 53.1% | 76.7% | 3.01 | 0.66 | 33.7% | 66.7% |
| NVIDIA GLiNER-PII alone | 9.2% | 18.9% | 2.69 | 0.59 | 24.5% | 59.3% |
| Airlock, placeholders (default) | 13.9% ± 0.6 | 8.6% ± 0.7 | 4.16 ± 0.03 | 0.95 ± 0.01 | 15.2% ± 1.4 | 11.1% |
Linkable disclosure drops from 84.7% with raw pass-through to 13.9% ± 0.6 through Airlock. Usefulness goes from 4.52 to 4.16, and the utility ratio stays at 0.95. Distortion is 15.2%, against 7.7% for the reference answers judged in the same run.
GLiNER-PII alone links less than Airlock, 9.2% against 13.9%. It gets there by masking 59.3% of benign prompts, and its answers score 2.69 with 24.5% distortion. The model was trained on English only; in one Korean case it labeled an age PASSWORD and a company RELIGIOUS_BELIEF. Presidio lands elsewhere on the curve: 53.1% linkable and 76.7% identity leak, while still masking 66.7% of benign prompts and scoring 3.01. Mask too little and identities leak; mask too much and the answers stop being useful.
Airlock has its own costs. 3 of 27 benign cases get a mask, and local overhead is p50 4.0 s and p95 8.8 s per request.
What still links, and what Korean costs
Every remaining linkable disclosure is a quasi-identifier case. There were 0 in finance, health and intent search, and 13.7 of 27 quasi-identifier cases per pass. 9 cases were linkable in all 3 passes and 18 in at least one. What links is a combination of attributes that each look harmless:
- roles and titles kept as context, such as
pediatric cardiologist - cohort and rank markers, such as
only one who voted against - small places, schools and sites, such as
Tillbrook High - rare personal facts, such as a
Korean-speakingfather aged 61
Closing these means generalizing one more attribute, which trades against the context the answer needs. Outside research points the same way: AURA, an agentic re-identification attack with web search, reports 13 to 21 of 27 interviews re-identified after Presidio.
Korean costs more.
| Metric | Korean | English |
|---|---|---|
| Identity leak | 10.6% ± 1.0 | 6.5% ± 0.6 |
| Linkable disclosure | 14.0% ± 3.5 | 13.9% ± 2.4 |
| Utility ratio | 0.86 ± 0.02 | 1.04 ± 0.04 |
| Over-redaction | 10.0% ± 1.9 | 4.3% ± 1.1 |
| Benign masked | 23.1% | 0.0% |
Linkable disclosure is about the same in both languages, with a wide spread on about 50 cases each. Every benign mask is Korean, and two benign Korean searches are blocked in every pass. Identity leaks outside the quasi-identifier category came from Korean adversarial encodings, such as a phone number in Korean numerals, plus one Korean name sent verbatim.
The passes that were not independent
The first final run reset the server only before pass 1, so passes 2 and 3 reused the detections Airlock had cached in memory. The local detector samples at temperature 0.6, and the cache removed that sampling. An identity leak of 9.7% ± 0.0 looked like stability. It measured nothing: it was one detection scored three times.
The signature was in the data. Local detect time p50 by pass was 3,412 / 462 / 475 ms in the cached run and 3,463 / 3,409 / 3,375 ms with a reset before every pass.
| Metric | Cached run | Independent run |
|---|---|---|
| Linkable disclosure | 15.3% ± 1.0 | 13.9% ± 0.6 |
| Identity leak | 9.7% ± 0.0 | 8.6% ± 0.7 |
| Cases linkable in every pass | 14 | 9 |
| 3-pass mean overhead p50 / p95 | 1,629 / 5,734 ms | 4,004 / 8,771 ms |
The means moved by about a point, within the spread of a single detection. The cache got the spread, the persistence and the latency wrong. The real spread on identity leak is 0.7 points, and the mean latency was understated by 2.4 s at p50. Cases linkable in every pass went from 14 to 9, so part of the apparent persistence was the cache repeating itself.
Counting identical payloads would not have caught this reliably. Nemotron-3-Nano-4B repeats many outputs exactly even when it really re-runs: 130 of 231 cases (56.3%) were byte-identical across the independent passes. The harness now warns on a drop in detect time, and on identical payloads only above 85%.
The lesson: repeated runs must be independent. If a component samples, reset its state before every pass and check independence with a number before you report a standard deviation.
Reproducing it, and the limitations
The code and harness are at github.com/tristan-kkim/airlock. With the local detector model, Nebius Token Factory and Tavily keys, and the GLiNER-PII weights in place:
git clone https://github.com/tristan-kkim/airlock && cd airlock
uv sync --extra gliner
eval/final_protocol.sh --commit 201341b --gliner on --substitution placeholder --passes 3 \
--reuse-committed --publish --systems "raw regex presidio_ko gliner_pii airlock" --yes
The independent run took 1 h 52 min. Scoring cost $3.66, and the estimated total including upstream answers is about $5.05. Full tables are in eval/results/FINAL.md and eval/results/COMPARISON.md.
The limitations:
- One attacker, so a lower bound. No reasoning, no retrieval, no background knowledge. A stronger adversary would link more, so compare systems with each other, not against zero.
- Synthetic data. The quasi-identifier and intent cases were written by the same project that built the detector, and rules were tuned on a 72-case dev split inside the 243 cases, so these are not held-out numbers.
- Judge noise of about ±0.05 in utility ratio and ±3 points in distortion. Baselines carry one draw of attacker and judge sampling, not three.
- Airlock removes identity, not topic. The private situation was still inferred in 72.8% of cases, and a value every detector misses is sent.
This post was written as part of a Nebius × NVIDIA Global AI Hackathon entry.
Founder of Cortexys, leading AX consulting, corporate AI training, and AI agent development.
Need an AI solution?
See custom AI development services at cortexys.team.

