Airlock: Egress Firewall for AI Agents
Dev

Airlock: Egress Firewall for AI Agents

Back to blog

Airlock checks what an AI agent sends to the cloud before it leaves your machine. Why it exists, how it works, and how it is measured.

Tristan Kim · CortexysAirlockAI agentsprivacyNVIDIA NemotronNebiusTavily한국어로 읽기

When an AI agent works on your private documents, what actually leaves your machine?

More than the prompt. Every planning turn re-sends what the agent has read so far. Tool calls carry arguments, and those arguments are replayed in the history. Search queries go to a search API, usually a different company from your model provider.

Airlock is an egress firewall for AI agents. It checks every hop locally before it leaves and restores the answer on the device. This is Part 1 of the series: why I built it, what it does, how it works, and how it is measured. Part 2 has the measured results. This post was written as part of a Nebius × NVIDIA Global AI Hackathon entry.

The problem: more exits than the prompt

Diagram of the channels through which an agent sends data off the machine. Model calls are covered by existing proxies, while search and tool traffic sits outside that coverage.
Agent egress channels: model calls that existing proxies cover, and the search and tool traffic outside that line

People and companies avoid cloud AI mainly because of leakage: names, customer data, health details and credentials pasted into a prompt. Local-only models avoid the leak, but they are much weaker than frontier models.

Agents widen the problem. Documents the agent reads go back to the cloud model on every planning turn, and the agent's own tool calls are sent again as part of that history. Search traffic goes somewhere else entirely. Most privacy tooling sits between you and the model. The search engine and the tools an agent calls usually sit outside that line.

Short queries are not automatically safe. A query like <ORG_1> layoff list team lead how to respond contains no name, and it still tells the search provider that someone at a company is on a layoff list.

What already exists, and what does not

Masking a prompt and restoring the answer is a solved problem, and I want to be clear about that up front.

  • PasteGuard is an OpenAI-compatible proxy that uses GLiNER, restores answers including streams, and supports coding agents.
  • Kiji Privacy Proxy from Dataiku masks with a quantized DistilBERT model and restores the answer.
  • LiteLLM's Presidio guardrail does the same job at the gateway.
  • PAPILLON (NAACL 2025) already uses a local model to rewrite queries before they reach an API model, and it measures leakage.

So splitting the work between a small local model and a large cloud model is not a new idea either.

The gap is in the other channels. MosaicLeaks (Gurung et al., 2026) shows that an observer can infer private document contents from a deep-research agent's search queries alone. Its fix is retraining the agent, which a user cannot put in front of an existing one. AWS Bedrock Guardrails documents that it does not inspect tool-call arguments or tool results. The READMEs of the open-source proxies above do not mention scanning tool calls or search queries. I read their documentation, not their code.

I also did not find a product that publishes a leak rate measured on the bytes that actually left the machine, repeated across runs, next to baselines. Research papers do measure leakage, so Airlock is not the first to measure it. The narrower aim is a product-level measurement with an open harness, against baselines run under the same conditions.

The thesis: identity versus situation

The goal is narrow on purpose. The cloud can know the problem. It shouldn't know whose problem it is.

That splits private information into two kinds. Identity is names, organizations, phone and account numbers, secrets, and quasi-identifiers: attributes that each look harmless but together single a person out. Situation is what the model needs to read in order to help: a diagnosis, a severance clause, an amount. Hide the situation too and the answer stops being useful.

So Airlock removes or generalizes identity and tries to keep the situation. At the default balanced level, amounts, lab values and diagnoses stay in the prompt, while a city like "Tulsa" becomes "a city in Oklahoma".

Using it means changing one base_url. Point any OpenAI-compatible app at http://127.0.0.1:8787/v1, and your scripts, notebooks, editors and agents go through Airlock.

How it works

Architecture diagram of Airlock: on the local machine, NVIDIA Nemotron-3-Nano-4B, NVIDIA GLiNER-PII, the vault, the deterministic gate and the optional Nemotron 3.5 Content Safety judge; in the cloud, Nemotron 3 Ultra on Nebius Token Factory and Tavily search.
The full system: what runs on your machine, what runs on Nebius Token Factory, and what goes to Tavily

A chat request passes five local steps before anything leaves.

Step-by-step flow of one request: span detection, placeholder substitution, the deterministic gate, the upstream call, and rehydration of the answer.
One request, step by step: spans, placeholders, gate, upstream, rehydrate
  • Detect. Regex, entropy and vault matching find secrets and identifiers first and mask them in the copy the local model reads, so the model never receives raw secrets. NVIDIA Nemotron-3-Nano-4B then proposes semantic spans (names, organizations, health details, quasi-identifiers) with JSON-schema-constrained output, and deterministic Korean rules back it up where the small model is weak. A model span whose text does not occur in the original is discarded, so a hallucinated ID never becomes a redaction.
  • Substitute. Identity values become typed placeholders such as <PERSON_1> in a local vault, and quasi-identifiers get a generalization. The same original always maps to the same placeholder within a conversation or agent run.
  • Gate. The final outbound JSON is walked string by string. If a vault original, a declared term, a canary or a high-confidence secret pattern is still present, the request is blocked with HTTP 422.
  • Upstream. The exact bytes that passed the gate go to Nemotron 3 Ultra on Nebius Token Factory.
  • Rehydrate. Placeholders in the answer, including tool-call arguments and streamed deltas, are mapped back to the originals on your machine.

Every request leaves an audit record with the outbound payload verbatim and detections as keyed hashes, so you can check afterwards exactly what the cloud saw.

Optionally, NVIDIA GLiNER-PII runs as a second local proposer. It has high recall and over-masks, most of all in Korean, so a span it finds alone has to pass a deterministic type check and an adjudication call to the local Nano model before it counts.

Why the gate is code

The local model proposes. It never decides. Nemotron-3-Nano-4B samples at temperature 0.6, so the same input can produce different spans on different runs. Writing "never send secrets" into its prompt shifts what it tends to do and guarantees nothing. A prompt is a prior, not a gate. A model never gets to say "this is fine."

The gate is plain code on the final bytes, and "present" covers variants: letter case, full-width and look-alike letters, separated Hangul jamo, inserted spaces, numbers spelled out in Korean or English, and values hidden in JSON strings, percent-encoding or base64 up to two layers deep. Content Airlock cannot inspect, such as images, is blocked. If the local model is down, times out or returns invalid JSON, the request is blocked too. Airlock fails closed.

The guarantee is narrow and checkable. A misbehaving local model can cause over-masking or a block. It cannot cause a string Airlock knows is sensitive to be sent.

Agent mode and search intent

Agent mode flow: each planning turn to Nemotron 3 Ultra passes the gate, local tools run on the device, and each search query is rewritten locally and judged before it reaches Tavily.
Agent hops, and where the search query is rewritten before Tavily sees it

In agent mode, every hop that leaves the machine is treated as an egress event, gated and audited.

HopDestinationWhat happens before it leaves
Planning turnNemotron 3 UltraThe whole history, including documents read and earlier tool calls, goes through the detector, the vault and the gate
Search queryTavilyThe search-intent guard: rewrite locally, gate, judge, retry once, then block that one search
Local toolsNowhereReading documents runs on your machine
Final answerNowhereRehydrated locally

Documents are shown to the model under neutral ids such as doc-1, because file names often describe the situation.

For search queries, the risk is rarely a string the gate already knows. It is intent. So Nemotron-3-Nano-4B rewrites each query toward a generic, information-seeking one, given the user's question and the documents read so far as the thing it must not reveal. In a live run, the agent asked for Tessellate Health AI ambient clinical documentation Denver competitors funding valuation, and Tavily received ambient clinical documentation market valuation competitors.

The rewritten query then passes the gate, and an intent judge decides whether it still reveals a private situation. The default judge is the same local Nano model; NVIDIA Nemotron 3.5 Content Safety can run instead in custom-policy mode, or both together. A rejected rewrite is retried once. If that fails too, only that search is withheld, and the agent is told and continues. A judge that errors counts as a flag.

Which NVIDIA and Nebius pieces do what

The design in one line: the small local model does not solve the task. It only decides what is private. The big cloud model does the reasoning.

RoleComponentWhere it runsWhy
Detector, search rewriter, default intent judgeNVIDIA Nemotron-3-Nano-4BYour machine, llama.cppSmall enough for a laptop, follows a JSON schema reliably, and sees private text, so it stays local
Second PII proposer (optional)NVIDIA GLiNER-PIIYour machineHigh recall; the PII backend of NeMo Guardrails
Reasoning over the sanitized requestNVIDIA Nemotron 3 UltraNebius Token FactoryFrontier-class open model that keeps placeholders intact, including in Korean
Search intent judge (optional)NVIDIA Nemotron 3.5 Content SafetyYour machineCustom-policy judgment on search queries
Web searchTavilyCloud APIReceives only queries that were rewritten and passed the gate

Token Factory is on the execution path of every allowed request, through an OpenAI-compatible API, and Nemotron 3 Super is used once as a fallback if Ultra is unavailable. Token Factory may store prompts and outputs unless Zero Data Retention is enabled, and I recommend enabling it. Airlock's guarantee does not depend on it: whatever the provider keeps is the sanitized payload shown in the audit log.

How it is measured

Evaluation design: synthetic Korean and English cases, four baselines, an attacker model that sees only outbound payloads, usefulness and distortion judges, and independent repeated passes.
The evaluation design: data, baselines, the attacker, the judges and independent passes

I did not want the thesis to stay a claim, so measurement was part of the design from the start.

  • Dataset: 243 synthetic cases, 122 Korean and 121 English. Every value is invented.
  • Baselines: raw pass-through, regex only, Presidio with Korean recognizers, and NVIDIA GLiNER-PII alone, in the same harness.
  • Attacker: Nemotron 3 Ultra reads only the outbound payloads and tries to recover identity items and the private situation.
  • Main metric: linkable disclosure, the share of the 98 situation-sensitive cases where the attacker recovers an identity item and also infers the private situation. It counts the cases where an observer could say who has which problem.
  • Usefulness: a filter that hides everything is easy to build and useless to work with, so a judge scores usefulness and distortion against reference answers, next to over-redaction and latency.
  • Repetition: the detector samples, so one run is not evidence. Every system runs 3 passes, the server is reset before every pass, and per-pass detect time confirms the passes were independent.
  • Agent scenarios: 16 fictional private-document scenarios, 8 Korean and 8 English, run with an unguarded agent and with Airlock. The attacker reads all hops once and the Tavily queries alone once, the MosaicLeaks threat model.

The measurement is done. Across all 243 cases, Airlock cuts linkable disclosure from 84.7% to 13.9% against raw pass-through. What that costs, what still links, and a mistake I made in the repeated runs are in Part 2.

Limitations and threat model

Airlock is designed to stop accidental disclosure of identifying or secret strings to a cloud LLM provider or a web search provider, through an app you control. It does not cover everything.

  • Detector misses. A value every detector misses, and that you did not declare, is sent. The gate only enforces what is known.
  • Meaning and situation. Airlock removes strings, not meaning. A rare combination of attributes can still identify someone, and the provider can still learn that someone has this situation.
  • Linked queries. Several generic queries in a row can suggest the topic, and Tavily sees timing and your network address. In my own agent scenarios, unguarded search queries were already fairly generic, so the search guard closed a small leak, not a large one.
  • Inbound content. Prompt injection in search results and documents is not checked, though an injected instruction still cannot get known strings past the gate.
  • Your machine. The vault stores originals in plain text locally, and local detection adds seconds to each request.

Try it

The code and the evaluation harness are open source under Apache 2.0 at github.com/tristan-kkim/airlock.

git clone https://github.com/tristan-kkim/airlock && cd airlock
uv sync && scripts/local_model/serve.sh
uv run airlock serve

The local UI shows what you typed, what the cloud saw, and the rehydrated answer side by side. A hosted demo with fictional scenarios goes live in October and stays free and open until judging ends on December 15. Its detection runs on the demo server, so do not paste real personal data there; for real privacy, run Airlock locally.

Next: the measured results

Part 2, First Measured Results from Airlock, covers the baseline comparison, the trade-off, what Korean costs, and the passes that turned out not to be independent.

김지우 (Tristan Kim)

Founder of Cortexys, leading AX consulting, corporate AI training, and AI agent development.

Need an AI solution?

See custom AI development services at cortexys.team.

Request a consultation