How Airlock's linkable disclosure went from 13.9% to 7.5% on the same 243 cases: the strategies tried, the settings compared, and which choices were adopted.
Part 2 reported the first independent measurement of Airlock, an egress firewall for AI agents: on 243 synthetic cases, linkable disclosure fell from 84.7% with raw pass-through to 13.9%, every remaining link was a quasi-identifier case, and Korean performed worse than English. This post covers what happened between that run and the current one, which measures 7.5% on the same 243 cases with the same harness: which strategies were tried, which settings were compared, and which choices were adopted and why. Ablation tables for each change are included. It was written as part of a Nebius × NVIDIA Global AI Hackathon entry.
Results at the end of part 2
The independent run at 201341b gave linkable disclosure 13.9% ± 0.6, identity leak 8.6% ± 0.7, usefulness 4.16 with a utility ratio of 0.95, distortion 15.2%, over-redaction 7.1% and a benign mask rate of 11.1%. Two weaknesses stood out. The 9 cases linkable in every pass and the 18 linkable in at least one were all quasi-identifier prompts, where a title, a cohort year and a small place, each harmless alone, singled a person out. And Korean performed worse: identity leak 10.6% against 6.5% in English, every benign mask was Korean, and two benign Korean searches were blocked in every pass.
Both weaknesses share the same trade-off. Generalize one more attribute and fewer cases link, but distortion and over-redaction rise. The strategies below are attempts to reduce linking without raising distortion or over-redaction.
Measurement procedure

- Dev and test split. Seeded (20260914) and stratified by category and language: 72 dev and 171 test cases. Rules, thresholds and variant choices were made on dev only, and reports quote the test column. Because the final 243 cases include the dev split, the final numbers are not a held-out estimate.
- Pass independence. After the cache bug in part 2, every run resets the server before every pass and checks independence with the per-pass detect time (3,463 / 3,409 / 3,375 ms p50 in an independent run, against 3,412 / 462 / 475 ms in the cached one).
- The attacker is a lower bound. Attacker and graders run at temperature 0. Attacking the pass-1 payloads three times gave linkable disclosure 8.2 / 8.2 / 8.2% with the same 8 cases each time, so the spread in the final tables comes from the detector's sampling. One attacker with no reasoning or retrieval; a stronger adversary would link more.
- Judge noise. The judge scores answer pairs, so the same cached reference answers scored 4.35, 4.40 and 4.38 across three runs. Differences within about ±0.05 in utility ratio or ±3 points in distortion are not read as changes.
Whether the judge could be cheaper was measured too. Each scoring role was re-run on a stratified sample with Nemotron 3 Super, Lightning and Nano, and a candidate qualified only if every headline rate stayed within ±3pp of Ultra (±0.03 for the utility ratio) and no pair of systems that Ultra separates by 10pp or more changed order.
| Role | Candidate | Largest move from Ultra | Verdict |
|---|---|---|---|
| Attacker | Lightning | situation inferred −17.5pp | rejected |
| Attacker | Nano | quasi re-identification +12.5pp | rejected |
| Grader | Nano | situation inferred +21.9pp, kappa 0.19 | rejected |
| Utility judge | Super | distortion −15.1pp | rejected |
| Distortion confirmation | Lightning | −3.5pp on a 144-item retest, kappa 0.81 | rejected |
| Grader | Ultra retest | situation inferred +1.6pp, kappa 0.89 | the reference noise |
Every candidate moved at least one headline number beyond the tolerance, so Ultra was kept for all four roles and scoring cost was reduced through protocol size instead: baselines produce identical output across passes, so they are scored once and reused by payload hash. The calibration cost $0.51; scoring the final run cost $3.69, about $5.1 with the upstream answers, in 2 h 19 min.
Strategy 1: recall from an ensemble, precision from agreement

NVIDIA GLiNER PII alone has high recall and low precision: on the 243 cases it links in 9.2% but masks 59.3% of benign prompts and scores 2.69 for usefulness. It was trained on English only; in qid-ko-01 it labeled the age 마흔다섯 PASSWORD and the company RELIGIOUS_BELIEF. So its spans pass three filters. A GLiNER span that overlaps one from regex, the Korean rules or Nano is accepted by agreement; agreement can only widen a mask. A GLiNER-only span must have the shape of its label: SECRET needs entropy or a pattern, phone and ID numbers need digits, PERSON needs a surname-initial Hangul name or capitalized Latin words. Only spans that pass go to Nemotron 3 Nano 4B in one batched call per request, at temperature 0 with a JSON schema, asking whether each is private to the user in this sentence. The threshold is 0.4 because on dev the raw scores barely separated true from false spans anywhere between 0.3 and 0.6.
The compared settings were variants B and B1. B adjudicates every label and accepts two-syllable Hangul names. B1 sets AIRLOCK_GLINER_AGREEMENT_ONLY=city,date_of_birth, so cities and birth dates need agreement, and AIRLOCK_GLINER_KO_NAME_MIN_SYLLABLES=3, so a Hangul name needs three syllables on GLiNER's word alone. The reason was dev precision: GLiNER-only cities were wrong in 4 of 8 accepted spans, the only benign dev false positive was 2024-01-01 read as a birth date, and 2 of 3 two-syllable Hangul "names" were common words.
| Test split, 171 cases | A (GLiNER off) | B (every label adjudicated) | B1 (default) |
|---|---|---|---|
| Leak rate | 26.3% | 8.6% | 11.2% |
| Canary leak | 14.1% | 2.0% | 1.0% |
| Quasi re-identification | 38.9% | 19.4% | 27.8% |
| Benign masked | 0.0% | 42.1% | 21.1% |
| Benign blocked | 0.0% | 15.8% | 10.5% |
| Benign masked ko / en | 0.0% / 0.0% | 44.4% / 40.0% | 33.3% / 10.0% |
B1 leaks 2.6 points more than B on test and masks benign prompts half as often; on dev its benign mask rate goes from 12.5% to 0.0%. The choice was made on dev, and no second variant was run. Over all 243 cases, the attacker's full recovery of planted values falls from 9.9% with GLiNER off to 1.0% with B1. Quasi re-identification improves less, because GLiNER has no label for roles or uniqueness cues; that problem moves to strategy 2.
The trade-off is latency. GLiNER inference is p50 457 ms and p95 635 ms, in parallel with Nano. Adjudication was needed in 58 of 206 chat requests (28%) and then took p50 1,706 ms; B adjudicated 78 of 205 at p50 2,196 ms. GLiNER alone was faster on MPS (p50 324 ms against 456 ms on CPU), but inside Airlock, while Nano decodes on Metal, MPS slowed to p50 1,165 ms and CPU stayed at 467 ms, so CPU is the default.
Strategy 2: unlinkability instead of redaction
The ensemble detected individual values but not attribute combinations. So the target changed: instead of removing strings, make sure the attributes that leave the machine do not point at fewer than k people.
The first change linked employer, org unit and role and generalized them together: 동해누리정밀(주) 품질보증팀 박성훈 팀장 leaves as <ORG_1> 품질 부서 <PERSON_1> 관리자, and the Payments Platform team at Halcyon Freight Systems as the engineering team at <ORG_1>. Amounts, lab values and common diagnoses stay in the prompt, and a generalization is accepted only if the original entails it.
| Test subset, 90 cases | B1 | Org-unit generalization |
|---|---|---|
| Linkable disclosure | 16.7% | 8.3% |
| Situation inferred | 55.0% | 72.5% |
| Over-redaction (171 cases) | 15.5% | 7.8% |
| Utility ratio | 0.91 | 0.87 |
| Distortion | 16.2% | 27.6% |
Linking and over-redaction halved, and situation inference rose, which is the design. In agent mode, on 4 scenarios, 0 of 19 identity facts reached the cloud, where team names such as 품질보증팀 and Payments Platform had survived every pass before. Distortion rose.
Next, that distortion was classified by mechanism from the stored payloads and answers. Of the 14 answers distorted under org-unit generalization but not B1, 6 had a payload identical to B1's (judge noise) and 4 were the upstream model embellishing around a diagnosis that had been kept. Across all 21 distortions: 8 placeholder misreads (a user's value called "a placeholder", something computed from a token, a full number restored under "last 4 digits"), 1 masked task value (dose times 7시, 11시, 15시, 19시), 1 diagnosis reworded by the local model, 5 embellishments and 6 upstream errors with an adequate payload. In 6 of the 8 misreads the token had the wrong type: two companies as <PERSON_n>, a job title as <ID_NUMBER_1>.
Fixes were made per mechanism. The placeholder note sent upstream now carries a legend built from the token types already in the payload (<ORG_n> is an organization's name; <SECRET_n> is a password, key, token or connection string, exactly as the user wrote it) and forbids calling a value a placeholder or computing anything from a token. The types are already visible in the payload, so the legend reveals nothing. Declared terms get a type from their surface instead of defaulting to PERSON, a number after "last 4 digits" is rehydrated as its last four digits, and at the balanced level what the task operates on stays: schedules, lab values, labeled amounts, diagnoses. On dev, distortion fell from 35.7% to 6.7%, over-redaction from 9.9% to 6.0%, and linkable disclosure stayed at 13.3%. The one distortion the org-unit rule caused (품질관리팀 read back as 품질 부서) was kept, because that rule stops team names leaking in agent mode.
The third change scored the combination itself. Quasi-identifier attributes fall in seven categories: place, named institution, cohort, rare fact, role, age and family structure. The first four weigh 2, the last three weigh 1, and only the most specific attribute per category still in the outbound text counts. When the sum reaches k (3 at balanced, 2 at strict), the most identifying attributes are coarsened until the score is below k. The tables are deterministic and entailed.
| Attribute | Original | Coarsened |
|---|---|---|
| Small place | 경북 새내군 | 경북의 한 군 지역 |
| Cohort year | 2022년 행정고시 | 2020년대 행정고시 |
| Specialty | pediatric cardiologist | physician |
| Rank | 수석 | 상위권 |
| Small committee | five-member planning commission | local board |
An attribute whose content word appears elsewhere in the message is what the question is about and is kept, and only a message about a person is scored. No local model call is involved. On dev, identity leak fell from 8.9% to 2.4%, distortion from 20.7% to 11.9%, and the utility ratio moved from 1.01 to 1.05. Linkable disclosure stayed at 4 of 30, where one case is 3.3 points, so a single attacker pass could not show a gain. A local scan of the 9 cases linkable in every pass of part 2 could: the old code reached the case's k in 16 of 18 passes, the new code in 3 of 18. In the final 3-pass run, the scanner's quasi re-identification went from 28.7% to 11.3% and the persistent linkable set from 9 cases to 4.
One attempt was reverted. Telling the search rewriter to keep Korean queries in Korean kept private details in one dev pass, blocked two searches and let a private organization reach a query. The prompt is unchanged.
Strategy 3: placeholders against surrogates
Whether an identity value becomes a typed placeholder such as <PERSON_1> or a realistic, format-preserving fake value was decided by a rule set before the measurement: keep placeholders unless surrogates are clearly better on linkable disclosure without worse distortion.
| Metric | Surrogate e0ee6aa † | Placeholder 201341b | Placeholder eb311ed |
|---|---|---|---|
| Linkable disclosure | 12.6% | 13.9% | 7.5% |
| Identity leak | 8.3% | 8.6% | 3.1% |
| Situation inferred | 54.0% | 72.8% | 71.6% |
| Usefulness | 4.08 | 4.16 | 4.15 |
| Distortion | 19.4% | 15.2% | 14.9% |
† is the cached run from part 2, so its spread is not measured. Placeholders in that same cached run linked 15.3% with 14.7% distortion, so on that detector surrogates linked 1.3 points less and distorted 4.2 points more. The drop in situation inference to 54.0% was not read as a privacy win: it means the cloud model understood the problem less. The 90-case test subset of the org-unit generalization showed the same pattern: surrogates leaked slightly less identity (8.3% against 9.0%) but linked more (11.1% against 8.3%) and distorted more (32.9% against 27.6%). The surrogate distortions had distinct mechanisms: a card "ending in" the surrogate's digits, a surrogate name transliterated into Korean in a translation task, a pronoun that did not match the surrogate's gender. Those were fixed afterwards but not rerun. So placeholders stayed the default. Surrogates are appropriate where the task operates on the value itself, such as input the user asks to decode or compute with; AIRLOCK_SUBSTITUTION=surrogate remains available and needs its own independent run before any claim about the current detector.
Strategy 4: detect on normalized text, not only gate on it
The gate knew about variants from the start: it compares after folding full-width characters, composing separated Hangul jamo, removing inserted spaces and rewriting numbers spelled in Korean. But the gate only enforces strings Airlock already knows. A name written as 김 민 지 that no detector recognized went out unmasked, and a vaulted value that survived in a variant form blocked the whole request: blocked rather than masked. At 201341b the Korean adversarial cases leaked names and card numbers, adv-ko-07 in every pass and adv-ko-13, adv-ko-06, adv-ko-02 and adv-ko-14 in some.
The current version runs detection on a normalized view as well. The regex rules and the Korean name rules see a copy where spaced or dotted syllables are joined, separated jamo composed, Korean, Hanja and English numerals rewritten as digits, and full-width forms and homoglyphs folded; hits map back to raw offsets, so the value is masked before the local model sees it. At eb311ed none of those five cases leaked in any pass. What remains is one pass each of adv-ko-08, a spaced resident registration number whose synthetic value fails the checksum so only Nano can catch it, and adv-ko-14.
| Case | At 201341b | Cause | Fix |
|---|---|---|---|
ben-ko-08, ben-ko-13 | blocked in every pass | a GLiNER-only organization (KTX, 불국사) was masked, the rewrite kept the public name, the gate blocked it | a GLiNER-only organization or place in a request with no person cue is the topic and stays |
ben-ko-09 | masked | 장영실이 masked as a person | listed public figures are dropped |
adv-ko-07, adv-ko-14 | over-redacted | the job title 선임연구원 typed as ORG or PERSON | a job title alone is never a PERSON span |
A precision problem from a demo recording was fixed with a structural check. The Korean resignation scenario sent the search <PERSON_8> 수급 자격 <PERSON_7>법: the detector had proposed 실업급여 (unemployment benefit) and 고용보험 (employment insurance) as people, and the only shape check on a model PERSON span was "fewer than four digits". Now a Hangul PERSON span must be 2 to 4 syllables per name, start with a surname syllable, and not end in a syllable that ends nouns but never given names (증, 법, 사), so 자격증, 근로기준법 and 실업급여 fail; a short exclusion list covers legal terms that pass every structural check (위로금, 배우자). Ages and birth years are generalized for their slot: 34살입니다 becomes 30대입니다, born in 1992 becomes born in the 1990s, and the "I'm age" seen in an earlier recording is gone. Korean over-redaction on dev fell from 6.5% to 3.5%.
In the final run, Korean identity leak went from 10.6% to 2.6% and benign masked from 23.1% to 0.0%; Korean now leaks less than English. The utility gap remains: utility ratio 0.88 against 1.03, distortion 17.6% against 12.1%. The reference answers themselves score higher in Korean (4.58 against 4.17), so part of the gap comes from the reference scores.
Agent mode: before and after, and a new defect
Agent mode uses the same detection entry point as chat. Masking is recomputed for the whole history on every turn, and a value detected in any slot of a turn is masked in every slot before the gate decides. The first org-unit generalization attempt blocked every run at its third turn because GLiNER read the tool-argument key doc_id as an HTTP cookie; object keys of tool-call arguments are now never rewritten. Each search query is rewritten by Nano toward a generic query, gated on the exact Tavily payload, judged for intent, retried once, and only then blocked on its own.
| Metric | Unguarded | Airlock d2ff994 | Airlock eb311ed |
|---|---|---|---|
| Private facts recovered, all hops | 82.3% ± 2.8 | 41.3% ± 4.2 | 23.6% ± 3.2 |
| Identity leak (scanner) | 95.8% ± 3.6 | 77.1% ± 3.6 | 33.3% ± 9.5 |
| Linkable disclosure | 83.3% ± 3.6 | 68.8% ± 6.2 | 17.2% ± 8.2 |
| Situation inferred (no anchor needed) | 87.5% ± 6.2 | 93.8% ± 6.2 | 87.1% ± 6.9 |
| Answer utility (rubric, 1 to 5) | 4.38 ± 0.29 | 4.71 ± 0.18 | 4.50 ± 0.49 |
| Runs finished | 100% | 100% | 93.8% ± 6.2 |
| Latency per run, median | 24 s | 367 s (GPU shared) | 216 s |
Strategy 2 was also effective on document input. With team names and employers gone, linkable disclosure fell from 68.8% to 17.2% while the situation is inferred in about 87% of runs in both modes. 93.2% of queries were rewritten, and the attacker reading only the Tavily queries inferred the situation in 0.0% of runs against 16.7%. The blind pairwise judge preferred the guarded answers (utility ratio 1.09, distortion 14.6% against 27.1%), because many unguarded answers state wrong dates, amounts or jurisdictions after long contexts. The remaining 23.6% of facts are detector misses, worst in the Korean M&A memo.
The run also exposed a new defect. 3 of 48 guarded runs ended blocked with no answer and scored 1, which is the whole English utility deficit (4.50 against 5.00). Two were pregnancy-en: a search result page contained the health words the detector had vaulted from the personal note, so the gate refused to send that tool result into the history. One was debt-en: the local detector fell into a repetition loop and the turn failed closed. In the d2ff994 run no guarded run was blocked. The gate did what it is designed to do, and a public page about pregnancy is still not a leak.
The stored audit records identified the causes. The two pregnancy-en blocks were the gate and the masker disagreeing about public search text. The English health rule matched "pregnant" on a result page and generalized it to its category, "pregnancy"; the word "pregnancy" on the same page is already its own category, so the generalization check rejected it and it was masked as <HEALTH_1>. The outbound history therefore contained a vault original by construction. On top of that, a positioned rule span masked only the sentence it matched, while the gate matches every occurrence. The debt-en block was the local detector falling into a repetition loop on both samples, local_detector_malformed:repetition, and the turn failing closed.
The fix is merged as 564fe0a. A health term equal to its own category is kept at balanced, a generalization is validated against every original in the request, and a positioned span is substituted at every occurrence. In agent mode, a tool result the gate still refuses after masking is dropped on its own: the agent is told and the turn is rebuilt, instead of blocking the turn. A malformed detector answer is retried with a fresh seed and +0.2 temperature; if that fails too, the turn runs degraded (deterministic spans, the rules and every GLiNER span, all masked) and is marked detector_degraded in the audit. The gate still runs on the result, documents and the question still block, and the chat path stays fail-closed. 10 new tests; 560 pass.
The fix was verified on the two affected scenarios and on pregnancy-ko as a control, 2 passes each in airlock mode, about 36 cloud calls.
| Scenario | Finished / blocked | Rubric utility | Identity facts outbound |
|---|---|---|---|
pregnancy-en | 2 / 0 | 5.00, 5.00 | 0 |
pregnancy-ko | 2 / 0 | 5.00, 5.00 | 0 |
debt-en | 2 / 0 | 5.00, 5.00 | card last 4 4417 |
6 of 6 runs finished, none blocked, rubric utility 5.00 on all 6. The 4417 in debt-en was also in the outbound of the eb311ed run, so it is a pre-existing detector miss rather than a regression; situation facts such as amounts, dates and 임신 11주 are kept by design. The full 16-scenario × 2-mode × 3-pass agent evaluation has not been rerun after this fix: the agent table above and the README numbers are still from eb311ed, and the full table will be re-measured before the October demo and submission.
Final results
Both runs reset the server before every pass, on the same data with the same scoring models, so the differences are detector changes plus sampling noise.
| Metric | 201341b | eb311ed | Change |
|---|---|---|---|
| Linkable disclosure | 13.9% ± 0.6 | 7.5% ± 0.6 | −6.4 pt |
| Identity leak (scanner) | 8.6% ± 0.7 | 3.1% ± 0.3 | −5.5 pt |
| Quasi re-identification (scanner) | 28.7% ± 2.3 | 11.3% ± 2.3 | −17.4 pt |
| Situation inferred | 72.8% ± 1.1 | 71.6% ± 2.8 | unchanged by design |
| Usefulness / utility ratio | 4.16 / 0.95 | 4.15 / 0.95 | unchanged |
| Distortion (reference in the same run) | 15.2% (7.7%) | 14.9% (10.0%) | within noise |
| Over-redaction | 7.1% | 5.3% | −1.8 pt |
| Benign masked / over-blocked | 11.1% / 7.4% | 0.0% / 0.0% | to zero |
| Local overhead p50 / p95 | 4,004 / 8,771 ms | 4,179 / 9,379 ms | +175 / +608 ms |
| By language | 201341b ko | eb311ed ko | 201341b en | eb311ed en |
|---|---|---|---|---|
| Identity leak | 10.6% | 2.6% | 6.5% | 3.6% |
| Linkable disclosure | 14.0% | 6.7% | 13.9% | 8.3% |
| Utility ratio | 0.86 | 0.88 | 1.04 | 1.03 |
| Over-redaction | 10.0% | 7.5% | 4.3% | 3.1% |
Guarded agent runs blocked: 3/48 at eb311ed, 0/6 on the affected scenarios after the fix; the full rerun is pending.
Four principles generalize. Get recall from an ensemble and precision from agreement and deterministic shape checks, and ask a model only the contextual question. Score combinations, not strings, and keep the attribute the question is about. Distortion is fixable only after reading the stored payloads and answers and counting mechanisms; most of it was a token type problem. And because the gate only enforces known strings, variants have to be caught by detection.
Four cases remain linkable in every pass. qid-en-06 is the former CFO of a small Colorado business that closed in January 2025; qid-en-11 an AP chemistry teacher at a Georgia high school who kept goal on the state-champion water polo team; qid-ko-10 인천 연수구 <LOCATION_1> 아파트 관리사무소장; qid-ko-14 a couple in their 30s farming apples in a rural county of Gyeongbuk. The school and the apartment complex are masked and the profession is generalized, but one more attribute (a district, a month, a team fact) keeps the combination at k. Removing them would require giving up some of the utility gained from org-unit generalization and the distortion fixes, and the English quasi-identifier cases are now the weakest category.
The combination scoring rule was tuned against part 2's 9 persistent cases and the 72-case dev split, all inside the 243 cases, so the final numbers are not a held-out estimate. Full tables and the per-change records are under eval/results/ in the repository. To reproduce:
git clone https://github.com/tristan-kkim/airlock && cd airlock
uv sync --extra gliner
eval/final_protocol.sh --commit eb311ed --gliner on --substitution placeholder --passes 3 \
--reuse-committed --publish --systems "raw regex presidio_ko gliner_pii airlock" --yes
uv run python eval/results/s10-hardening/report.py
The series and the repository
Part 1, Airlock: Egress Firewall for AI Agents, covers why it exists, how it works and how it is measured. Part 2, First Measured Results from Airlock, covers the first independent measurement and the cache bug. The code and the evaluation harness are open source under Apache 2.0 at github.com/tristan-kkim/airlock.
Founder of Cortexys, leading AX consulting, corporate AI training, and AI agent development.
Need an AI solution?
See custom AI development services at cortexys.team.

