Version 1.0.0
Route a developer's IDE to the gateway
Point to the gateway in a workstation's configuration: it is the workstation's only path to the model.
The workstation does not reach the inference engine: it only knows the gateway’s address, and the gateway authenticates, applies policy, logs, and is the only component that reaches the engine (security white paper V3, sections 3.3 and 8.2). This page describes the workstation configuration that points to the gateway.
The purpose of this detour is described in The role of the gateway.
What the workstation declares
Section titled “What the workstation declares”A model in the Lemniscate configuration points to the gateway through its apiBase
and authenticates to it with the developer account’s key.
In ~/.lemniscate/workstation.yaml:
models: - name: Code principal (passerelle) provider: openai model: any apiBase: https://passerelle.interne:6001/code-principal/v1 apiKey: sk-lemniscate-<la-clé-du-développeur> roles: - chatField by field:
apiBasecarries the gateway address and the endpoint name:https://<hôte>:<port>/<nom-de-l-endpoint>/v1. The first path segment names the endpoint; what follows is relayed as-is to the inference engine. The address useshttps://: traffic between components is encrypted with TLS (security white paper V3, section 6.2).providernames the dialect the IDE speaks, that is, the shape of the requests and the responses. The gateway reaches the engine through an OpenAI-compatible API: the value isopenai. The endpoint, on the gateway side, decides which engine is called.modelis ignored. The gateway rewrites the model in the request body with the one the endpoint names. Any value works; a value that looks like a real model name misleads the next person who reads the file.apiKeyis the account key on the gateway. It is only valid against the gateway, which strips the caller’s authentication header before relaying. If the engine requires a key, the gateway reads it from an environment variable in its own container; that key is not sent to the workstation.rolesischat,summarize,applyandeditwhen absent. Name them to restrict this model to one use.
Port 6001 is the gateway’s default listening port; its PORT variable
redefines it.
To apply this routing to a single project, put the same models block in a
.lemniscate/models/ file at the root of the repository concerned.
The full list of keys this file accepts is in the reference.
Check the endpoint name
Section titled “Check the endpoint name”This is the most common mistake, and the gateway names it: an unknown endpoint
gets a 404 whose body lists the endpoints that exist.
{ "error": "Unknown endpoint: code-principale", "available": ["code-principal", "completion-rapide"]}Other refusals to recognize:
| Response | What it means |
|---|---|
401 | key missing or unknown to the gateway |
403 | the endpoint is disabled, or no role on the account carries the right to call it |
500 naming a variable | the engine expects a key whose variable is missing from the gateway environment |
The serverless build profile adds two refusals, tied to spending plans: 403 for
an account with no plan, 402 for an account that has reached its plan’s
monthly cap. They are cleared by different actions:
Cap a customer’s spending. A service that does
not respond at all belongs to
The gateway refuses to start.
The gateway accepts the key under Authorization: Bearer <clé> as well as under
x-api-key: <clé>.
Behind a corporate firewall
Section titled “Behind a corporate firewall”The two halves of the path are configured separately.
From the workstation to the gateway, the IDE goes through a corporate proxy: the
configuration accepts requestOptions with proxy, noProxy, caBundlePath
and clientCertificate, at the root level as well as per model.
From the gateway to the inference engine, inside your perimeter, the call is configured through environment variables on the gateway service:
| Variable | What it sets |
|---|---|
HTTPS_PROXY, HTTP_PROXY | the internal proxy to go through to reach the engine |
NO_PROXY | the destinations to reach directly, with wildcards, suffixes and ports |
NODE_EXTRA_CA_CERTS | an internal certificate authority, added to Node’s store |
UPSTREAM_TLS_CA_FILE | an authority specific to the gateway’s outbound path |
UPSTREAM_TLS_CLIENT_CERT_FILE, UPSTREAM_TLS_CLIENT_KEY_FILE | the client certificate, when the engine requires mTLS |
Lowercase names (https_proxy, http_proxy) are accepted, with the same
precedence as in the rest of the product.
An invalid setting stops startup: a proxy variable that is not an address, or an
unreadable NODE_EXTRA_CA_CERTS, are rejected at launch, with the name of the
offending variable.
If you configure mTLS, the client certificate does not override corporate trust:
the trust store is assembled, and NODE_EXTRA_CA_CERTS still applies.
If the response arrives all at once instead of being written out
Section titled “If the response arrives all at once instead of being written out”A proxy that buffers responses breaks word-by-word display, whatever the gateway does. The developer sees a silent assistant for several seconds, then a response that appears all at once. The total time is normal; its distribution changes.
This is the default behavior of many appliances that inspect content: inspecting means having the whole response before relaying it.
To check this on your side, compare the same call with and without the proxy:
- From a workstation that goes through the proxy, send a request whose response is long.
- Repeat the same request from a workstation or a network that does not go through it.
If the words are written out progressively in the second case and not in the first, the buffering is in the proxy. There is nothing to fix on the product side.
On the other half of the path, between the gateway and the engine, the gateway
measures the arrival rate of long responses. When several responses arrive in a
single burst, it writes this to its error output under the [upstream-buffering] prefix, once,
then every hundred occurrences. Its silence proves nothing: it does not see the
path between the workstation and itself.
The setting depends on the appliance. For example, mitmproxy buffers by default
and relays progressively with --set stream_large_bodies=1. Look for the equivalent
in yours: the notion is called streaming, pass-through or
no buffering.
For support, this is the first question to ask when a customer reports an assistant that is slow to respond while server-side measurements are good.
What this path exposes
Section titled “What this path exposes”To weigh when choosing which network to put the gateway on.
The path from the workstation to the gateway carries the requests and the
developer’s key. It is encrypted with TLS (security white paper V3, section 6.2):
the gateway terminates TLS when you give it TLS_CERT_FILE and TLS_KEY_FILE, and
the workstation’s apiBase then uses https://. Without those two variables, the
gateway listens in cleartext and announces it at startup; a TLS terminator
placed in front of it must then encrypt the workstation traffic.
The engine’s address is data in the database, read from the base_url column
of the declared engine. The console checks the shape of that address (an
absolute URL, in https://, or in http:// with explicit acceptance), not its
destination. Anyone who can write to the database can direct requests to another
address. What protects this column is the console: it requires a role per
route, can delegate authentication to your organization’s directory, and
logs administration actions (granting a role, removing one, revoking a
subject). A shared pair of bootstrap credentials is still tolerated for the
first startup; remove it once roles are in place.
What this page does not describe
Section titled “What this page does not describe”The configuration accepts a provider named lemniscate-proxy, handled separately
in the schema, with its own orgScopeId and onPremProxyUrl fields.
This is not the path described here, nor the way to reach your gateway.
Its implementation points by default to a service hosted by the vendor, whose
existence and operation nothing in the repository attests to; for that reason
the surface inventory classifies it as indeterminate. The
onPremProxyUrl mode is self-contained, and nothing in the repository makes it
possible to write its procedure without inventing it.
On the other legacy surfaces that talk to this same unoperated service: What does not exist.