Skip to content
Version 1.0.0

Where the data goes

Where each category of data lives, which flows exist between components, and how to verify it yourself.

Lemniscate runs entirely inside your perimeter, and no flow leaves it. This page describes where each category of data lives, which flows carry it between components, and how you verify it yourself.

Four flows are allowed, all internal to the perimeter (security white paper V3, section 3.3). Each carries a known category of data.

FlowWhat it carries
Workstation to gatewaythe developer’s instruction, their validation decisions, the diffs to present
Gateway to execution hostthe session’s control channel: working copy of the repository, commands, results
Gateway to inference enginethe context sent to the model, for the duration of the request
Gateway to directory and SIEMauthentication and log export

Four flows are forbidden by construction: from the sandbox to any network, from the workstation to the execution host or the engine, from any component to the Internet, and from the execution host to any component. The full matrix, with the detail of each row, is in Architecture.

Six categories of data, each with a known location and duration (security white paper V3, section 8.3).

CategoryWhere it livesDuration
Repository source codethe workstation, and the working copy in the sandbox; no index, no derived copythe copy is destroyed with the sandbox
Context sent to the modelbetween the gateway and the inference enginethe duration of the request; not retained by default
Session contentthe database, on your side, if you enable its retentionseparate retention and encryption, which you set
Usage metadatathe database, on your side: user, date, action type, files, modelretained
Secrets and credentialsexcluded from the context, detected at the gatewaynot logged in clear text
Identities and groupsyour directory, where they are readthat of your directory

The only data persisted by default is usage metadata. Session content is retained only if you decide to. By default, the log contains neither requests nor code: it creates no copy of the source code.

Code is read by exploring the repository on demand. Lemniscate builds no index of the source code, and neither trains nor fine-tunes any model on your code (security white paper V3, section 11.3).

All data stays on your side. You therefore handle GDPR data subject requests yourself, with your own procedures. The vendor receives no data, and the product contains no mechanism that would allow it to.

Everything sent to the model passes through the gateway (security white paper V3, section 8.2).

  • The inference engine listens only on the inference server’s internal network.
  • Your network filtering allows only the gateway host to reach it.
  • The extensions and the sandboxes do not know its address.
  • The gateway detects secrets before transmission.

Secret detection relies on known patterns (infrastructure provider keys, access tokens, private keys, connection strings) and on entropy-based detection. The policy sets the response: block, mask or alert. Blocking is the default.

In the sandbox, the agent sees only the working copy of the repository, minus the files designated by the context exclusions: environment files, secret directories, certificates. The developer’s home directory, their keys, their forge credentials and their environment variables are not mounted there; see Execution guardrails.

When the inference engine is your own, the provenance of the model weights remains your responsibility. The white paper recommends the safetensors format, which excludes code execution at load time, and fingerprint verification of the weights.

No component has a destination outside your perimeter (security white paper V3, sections 1.3 and 11.3).

What is not sentDetail
Product telemetryno usage events, no error reports
Version checkno call to find out whether an update exists
License checkno call to the vendor
Startup probeno call when modules load

Neither the services nor the extensions update themselves: an update is verified and then applied by your team. The allowAnonymousTelemetry setting, which the workstation configuration accepts, opens no transmission in the on-premise artifact: the code capable of sending is not there.

The first two rows of the table, drawn alongside the flows that do exist:

Votre périmètrePoste, artefact construit en profil on-premise

Télémétrie produit

core/egress, implémentation inerte

Vérification de version

Consigne et validations

Passerelle

Bac à sable, sans réseau

Moteur d'inférence

Three things are visible in this drawing.

  • These flows do not stop at a boundary: they stop at the source. The inert layer is inside the workstation artifact, not at its exit; that is why it is drawn inside. The cross is not a firewall: there is no flag to turn back on, and no network rule to write to prevent them.
  • The workstation has a single destination, the gateway. The firewall rule for workstations fits on one line.
  • No arrow leaves the perimeter. The inference engine is on your network, and the gateway is the only thing that reaches it.

The absence of transmission does not come from a setting: it comes from the way the artifact is built. The product is built according to a deployment profile, chosen at build time rather than at runtime. The artifact installed in your perimeter is built with the on-premise profile. The profile is read from the build’s --profile= argument or, failing that, from the LEMNISCATE_DEPLOYMENT_PROFILE variable. An unknown value makes the build fail.

The on-premise profile is the default: running the build without specifying a profile produces this artifact. A configuration oversight therefore cannot produce an artifact that sends data.

In the on-premise artifact, the network egress layer is inert: it is not disabled by a flag that could be turned back on. The inert implementation is the one the code imports by default, so there is nothing to run to obtain it. This commitment corresponds to an absence of code, which can be confirmed during an audit.

Two verifications complement each other: observing your network, and inspecting the delivered artifact.

The white paper suggests three checks that you run with your own probes (security white paper V3, section 12.1):

  • observe the network, and confirm that no flow exists outside the flow matrix;
  • attempt network egress from a sandbox;
  • attempt a read outside the repository copy.

The verification targets the delivered artifact rather than the source code: a script inspects the built bundle and fails if any of the symbols or libraries it watches for remains in it. It applies four checks: forbidden symbols in the text of the artifacts, the entries in the build files, the inventory of deployable units, and boundary calls at module load time. The last two read the sources, because a call written on the caller’s side does not appear in the bundle’s dependency list.

Fenêtre de terminal
npm run build --prefix extensions/cli
npm run esbuild-base --prefix extensions/vscode
(cd gui && NODE_OPTIONS=--max-old-space-size=6144 npx vite build)
npm run build --prefix llm-gateway
npm run build --prefix gateway-admin

The check inspects six units: the five above and the JetBrains plugin (intellij). Without the --composant option, it inspects all of them. A component that has not been built is not deemed compliant: it is reported as missing, and the check fails. An artifact built with another profile is also reported, as indeterminate.

Fenêtre de terminal
node scripts/verify-onprem-artifact.mjs \
--composant cli --composant vscode --composant gui \
--composant llm-gateway --composant gateway-admin

This command, the one used in continuous integration, names five components. The same check runs on every proposed change to the product, with these five components; the JetBrains plugin is inspected the same way, outside continuous integration. A regression that reintroduced an outgoing call in these five components cannot be merged without making this check fail.

The tutorial Verify that no data leaves walks through this verification step by step.

  • The workstation’s other egress boundaries. core/egress is Node code, so it covers neither the interface displayed in the IDE nor the command line program. Those two have their own boundary, inert in on-premise in the same way. The drawing shows one; there are three.
  • The JetBrains plugin is the fourth, and it works differently. It is written in Kotlin: module substitution, which is a JavaScript mechanism, does not reach it. Its equivalent is a build profile (-PprofilDeploiement, cloisonne by default) that selects a set of sources, and the result is the same: with the air-gapped profile, the error reporting library is not in the delivered archive. The artifact check opens this archive, down to the .jar files it contains, and inspects it like the five other units.
  • The scope of the artifact check. It attests to the absence of a closed list of named libraries and symbols. The absence of outgoing flows is confirmed on your network, through the checks in the previous section.
  • The directory and the SIEM, which the gateway authenticates against and exports logs to. They are inside your perimeter, at addresses you set.
  • The gateway’s own network configuration: which authorities it recognizes, which certificate it presents. That is the subject of The chain of trust.