Airlock is submitted to the Nebius × NVIDIA Global AI Hackathon. The three-minute video, the public demo URL, and the agent-mode evaluation rerun with the GLiNER ensemble on (linkable disclosure 81.2% to 8.3%), with all three runs side by side.
Airlock, an egress firewall for AI agents, is submitted to the Nebius × NVIDIA Global AI Hackathon. This post records the three deliverables and the agent-mode measurement that was rerun right before submission. Part 3 promised that full rerun; what it showed is that one evaluation setting decides most of the result.
What was submitted
- The demo video (2:55, English narration and captions): https://youtu.be/vZgG8ndU9Ac
- The public demo: https://airlock-egress.fly.dev
- The repository (Apache-2.0): https://github.com/tristan-kkim/airlock
The video is a recording of Airlock running locally on a Mac: the local Nemotron-3-Nano-4B detecting, Nemotron 3 Ultra on Nebius Token Factory answering over the masked text, the gate inspecting the outbound bytes and blocking, and the agent reading two private documents and rewriting its Tavily queries on the device into generic ones. The browser was driven by Playwright while recording, and the narration is a synthetic voice. Every name, company, number and document on screen is fictional.
The final agent-mode measurement
The agent-mode evaluation runs 16 fictional scenarios (8 Korean, 8 English) in two modes, unguarded and Airlock, three passes each. An attacker model that reads only what left the machine tries to name an identity fact and infer the private situation; a run where it does both counts as a linkable disclosure.
It was run twice before submission: first in the default configuration (GLiNER ensemble off), then with the GLiNER ensemble on. Each row compares the two modes inside one run.
| Run | GLiNER ensemble | Linkable disclosure, unguarded → Airlock | Identity leak (scanner), Airlock | Facts recovered, Airlock | Rubric utility, unguarded / Airlock | Guarded runs without an answer |
|---|---|---|---|---|---|---|
| 2026-10-02, on | on | 81.2% → 8.3% ± 3.6 | 16.7% ± 3.6 | 20.5% ± 2.6 | 4.38 / 4.58 | 1 of 48 |
| 2026-10-02, off | off | 93.8% → 31.2% ± 6.2 | 43.8% ± 10.8 | 27.8% ± 3.3 | 4.62 / 4.35 | 3 of 48 |
| 2026-09-15 (the part 3 numbers) | off | 83.3% → 17.2% ± 8.2 | 33.3% ± 9.5 | 23.6% ± 3.2 | 4.38 / 4.50 | 3 of 48 |
Three things came out of it.
First, agent mode should run with the GLiNER ensemble on. With it off, English names inside documents depend on the 4B model alone, and that is a sampling outcome: fed one scenario's offer letter ten times, the detector proposed the addressee's name twice and returned an empty list eight times. That is why the two runs without the ensemble came out at 17.2% and 31.2%; the local model's span proposal (prompt and call) did not change between them, and three passes over 16 scenarios cannot separate the two numbers. With the ensemble on, the scanner found an identity fact in 8 of 48 guarded runs and the attacker linked who to what in 4.
Second, the unguarded column moves between runs too. It has no detector in its path, yet its linkable disclosure ranged from 81.2% to 93.8%, because the attacker and the graders are a sampled LLM. So the comparison is between the two columns of one run, never between runs.
Third, the English blocks are gone and what remains is Korean. The 3 guarded runs that ended blocked without an answer in part 3 were all English, caused by a public search result page that contained a word the detector had vaulted. After the fix, 0 of 24 English guarded runs were blocked in either run, and English rubric utility is 5.00 in both modes. The one run without an answer with the ensemble on is a Korean lawsuit scenario: the model's own search query repeated a phrase the detector had vaulted as a quasi-identifier, and the model's own text still fails closed.
What still leaks with the ensemble on is the code name of the deal in both M&A memos; it sits in the document title and went out in every pass. One Korean name and one neighbourhood went out in one pass each. Situation facts (amounts, dates, the other party) are recovered at 44.7%, against 74.8% unguarded. The cloud has to know the problem to help; what Airlock removes is the identity. The situation was inferred in 85.4% of unguarded and 79.2% of guarded runs.
Utility favoured the guarded side: rubric 4.58 against 4.38, pairwise utility ratio 1.08 ± 0.11, distortion 16.7% against 27.1%. Korean answers vary a lot in both modes: 10 unguarded and 5 guarded Korean answers repeated one clause dozens of times, misstated a statute or a date, or ignored the documents and answered a generic case. Latency is 388 s median per guarded run against 18 s unguarded; with the ensemble off it was 242 s.
The ensemble-on run was interrupted when the laptop went to sleep and lost the network. The five runs that ended with a connection error record the outage, so they were removed, and the harness was resumed on the same output directory to complete pass 2 and pass 3. Services are built fresh per pass, so the passes stay independent. The details are in eval/results/FINAL.md in the repository.
The public demo
The demo runs on one Fly.io machine. Each visitor gets an isolated in-memory session; the vault and the audit log are never written to disk and expire after 30 minutes. The hosted demo's detector is Nemotron-3-Nano-30B-A3B on Token Factory, which differs from the local install's 4B model: input reaches Token Factory before masking, so a banner says not to paste real personal data and points to the local install for real protection. Zero Data Retention is on for the demo account. Requests, input size and a daily budget are capped, and when the budget is spent or a live run fails, the same input is shown as a recorded run, labelled as such.
Next
The submission is in; judging runs from November 2 to December 15. The demo stays up through that window. The next measured targets are document titles, which are not masked today, and Korean common nouns that the detector sometimes vaults as quasi-identifiers and that then block the model's own query.
Series and repository
Part 1 Airlock, an egress firewall for AI agents, part 2 the first measured results, part 3 how the numbers improved. Every number can be recomputed from eval/results/FINAL.md and README.md in https://github.com/tristan-kkim/airlock.
Founder of Cortexys, leading AX consulting, corporate AI training, and AI agent development.
Need an AI solution?
See custom AI development services at cortexys.team.

