Version 1.0.0
Where the data goes
Where each category of data lives, which flows exist between components, and how to verify it yourself.
Lemniscate runs entirely inside your perimeter, and no flow leaves it. This page describes where each category of data lives, which flows carry it between components, and how you verify it yourself.
Flows between components
Section titled “Flows between components”Four flows are allowed, all internal to the perimeter (security white paper V3, section 3.3). Each carries a known category of data.
| Flow | What it carries |
|---|---|
| Workstation to gateway | the developer’s instruction, their validation decisions, the diffs to present |
| Gateway to execution host | the session’s control channel: working copy of the repository, commands, results |
| Gateway to inference engine | the context sent to the model, for the duration of the request |
| Gateway to directory and SIEM | authentication and log export |
Four flows are forbidden by construction: from the sandbox to any network, from the workstation to the execution host or the engine, from any component to the Internet, and from the execution host to any component. The full matrix, with the detail of each row, is in Architecture.
The data map
Section titled “The data map”Six categories of data, each with a known location and duration (security white paper V3, section 8.3).
| Category | Where it lives | Duration |
|---|---|---|
| Repository source code | the workstation, and the working copy in the sandbox; no index, no derived copy | the copy is destroyed with the sandbox |
| Context sent to the model | between the gateway and the inference engine | the duration of the request; not retained by default |
| Session content | the database, on your side, if you enable its retention | separate retention and encryption, which you set |
| Usage metadata | the database, on your side: user, date, action type, files, model | retained |
| Secrets and credentials | excluded from the context, detected at the gateway | not logged in clear text |
| Identities and groups | your directory, where they are read | that of your directory |
The only data persisted by default is usage metadata. Session content is retained only if you decide to. By default, the log contains neither requests nor code: it creates no copy of the source code.
Code is read by exploring the repository on demand. Lemniscate builds no index of the source code, and neither trains nor fine-tunes any model on your code (security white paper V3, section 11.3).
All data stays on your side. You therefore handle GDPR data subject requests yourself, with your own procedures. The vendor receives no data, and the product contains no mechanism that would allow it to.
Between the session and the model
Section titled “Between the session and the model”Everything sent to the model passes through the gateway (security white paper V3, section 8.2).
- The inference engine listens only on the inference server’s internal network.
- Your network filtering allows only the gateway host to reach it.
- The extensions and the sandboxes do not know its address.
- The gateway detects secrets before transmission.
Secret detection relies on known patterns (infrastructure provider keys, access tokens, private keys, connection strings) and on entropy-based detection. The policy sets the response: block, mask or alert. Blocking is the default.
In the sandbox, the agent sees only the working copy of the repository, minus the files designated by the context exclusions: environment files, secret directories, certificates. The developer’s home directory, their keys, their forge credentials and their environment variables are not mounted there; see Execution guardrails.
When the inference engine is your own, the provenance of the model weights remains your responsibility. The white paper recommends the safetensors format, which excludes code execution at load time, and fingerprint verification of the weights.
What is not sent
Section titled “What is not sent”No component has a destination outside your perimeter (security white paper V3, sections 1.3 and 11.3).
| What is not sent | Detail |
|---|---|
| Product telemetry | no usage events, no error reports |
| Version check | no call to find out whether an update exists |
| License check | no call to the vendor |
| Startup probe | no call when modules load |
Neither the services nor the extensions update themselves: an update is verified
and then applied by your team. The allowAnonymousTelemetry setting, which the workstation
configuration accepts, opens no transmission in the on-premise artifact: the
code capable of sending is not there.
The first two rows of the table, drawn alongside the flows that do exist:
Three things are visible in this drawing.
- These flows do not stop at a boundary: they stop at the source. The inert layer is inside the workstation artifact, not at its exit; that is why it is drawn inside. The cross is not a firewall: there is no flag to turn back on, and no network rule to write to prevent them.
- The workstation has a single destination, the gateway. The firewall rule for workstations fits on one line.
- No arrow leaves the perimeter. The inference engine is on your network, and the gateway is the only thing that reaches it.
The build profile
Section titled “The build profile”The absence of transmission does not come from a setting: it comes from the way
the artifact is built. The product is built according to a deployment profile,
chosen at build time rather than at runtime. The artifact installed in your
perimeter is built with the on-premise profile. The profile is read from the
build’s --profile= argument or, failing that, from the LEMNISCATE_DEPLOYMENT_PROFILE variable. An
unknown value makes the build fail.
The on-premise profile is the default: running the build without specifying a
profile produces this artifact. A configuration oversight therefore cannot produce
an artifact that sends data.
In the on-premise artifact, the network egress layer is inert: it is not
disabled by a flag that could be turned back on. The inert implementation is the
one the code imports by default, so there is nothing to run to obtain it. This
commitment corresponds to an absence of code, which can be confirmed during an
audit.
How to verify it
Section titled “How to verify it”Two verifications complement each other: observing your network, and inspecting the delivered artifact.
On your network
Section titled “On your network”The white paper suggests three checks that you run with your own probes (security white paper V3, section 12.1):
- observe the network, and confirm that no flow exists outside the flow matrix;
- attempt network egress from a sandbox;
- attempt a read outside the repository copy.
On the delivered artifact
Section titled “On the delivered artifact”The verification targets the delivered artifact rather than the source code: a script inspects the built bundle and fails if any of the symbols or libraries it watches for remains in it. It applies four checks: forbidden symbols in the text of the artifacts, the entries in the build files, the inventory of deployable units, and boundary calls at module load time. The last two read the sources, because a call written on the caller’s side does not appear in the bundle’s dependency list.
npm run build --prefix extensions/clinpm run esbuild-base --prefix extensions/vscode(cd gui && NODE_OPTIONS=--max-old-space-size=6144 npx vite build)npm run build --prefix llm-gatewaynpm run build --prefix gateway-adminThe check inspects six units: the five above and the JetBrains plugin
(intellij). Without the --composant option, it inspects all of them. A
component that has not been built is not deemed compliant: it is reported as
missing, and the check fails. An artifact built with another profile is also
reported, as indeterminate.
node scripts/verify-onprem-artifact.mjs \ --composant cli --composant vscode --composant gui \ --composant llm-gateway --composant gateway-adminThis command, the one used in continuous integration, names five components. The same check runs on every proposed change to the product, with these five components; the JetBrains plugin is inspected the same way, outside continuous integration. A regression that reintroduced an outgoing call in these five components cannot be merged without making this check fail.
The tutorial Verify that no data leaves walks through this verification step by step.
What this diagram does not show
Section titled “What this diagram does not show”- The workstation’s other egress boundaries.
core/egressis Node code, so it covers neither the interface displayed in the IDE nor the command line program. Those two have their own boundary, inert inon-premisein the same way. The drawing shows one; there are three. - The JetBrains plugin is the fourth, and it works differently. It is written in
Kotlin: module substitution, which is a JavaScript mechanism, does not reach it.
Its equivalent is a build profile (
-PprofilDeploiement,cloisonneby default) that selects a set of sources, and the result is the same: with the air-gapped profile, the error reporting library is not in the delivered archive. The artifact check opens this archive, down to the.jarfiles it contains, and inspects it like the five other units. - The scope of the artifact check. It attests to the absence of a closed list of named libraries and symbols. The absence of outgoing flows is confirmed on your network, through the checks in the previous section.
- The directory and the SIEM, which the gateway authenticates against and exports logs to. They are inside your perimeter, at addresses you set.
- The gateway’s own network configuration: which authorities it recognizes, which certificate it presents. That is the subject of The chain of trust.