Skip to content
Version 1.0.0

Route a developer's IDE to the gateway

Point to the gateway in a workstation's configuration: it is the workstation's only path to the model.

The workstation does not reach the inference engine: it only knows the gateway’s address, and the gateway authenticates, applies policy, logs, and is the only component that reaches the engine (security white paper V3, sections 3.3 and 8.2). This page describes the workstation configuration that points to the gateway.

The purpose of this detour is described in The role of the gateway.

A model in the Lemniscate configuration points to the gateway through its apiBase and authenticates to it with the developer account’s key.

In ~/.lemniscate/workstation.yaml:

models:
- name: Code principal (passerelle)
provider: openai
model: any
apiBase: https://passerelle.interne:6001/code-principal/v1
apiKey: sk-lemniscate-<la-clé-du-développeur>
roles:
- chat

Field by field:

  • apiBase carries the gateway address and the endpoint name: https://<hôte>:<port>/<nom-de-l-endpoint>/v1. The first path segment names the endpoint; what follows is relayed as-is to the inference engine. The address uses https://: traffic between components is encrypted with TLS (security white paper V3, section 6.2).
  • provider names the dialect the IDE speaks, that is, the shape of the requests and the responses. The gateway reaches the engine through an OpenAI-compatible API: the value is openai. The endpoint, on the gateway side, decides which engine is called.
  • model is ignored. The gateway rewrites the model in the request body with the one the endpoint names. Any value works; a value that looks like a real model name misleads the next person who reads the file.
  • apiKey is the account key on the gateway. It is only valid against the gateway, which strips the caller’s authentication header before relaying. If the engine requires a key, the gateway reads it from an environment variable in its own container; that key is not sent to the workstation.
  • roles is chat, summarize, apply and edit when absent. Name them to restrict this model to one use.

Port 6001 is the gateway’s default listening port; its PORT variable redefines it.

To apply this routing to a single project, put the same models block in a .lemniscate/models/ file at the root of the repository concerned.

The full list of keys this file accepts is in the reference.

This is the most common mistake, and the gateway names it: an unknown endpoint gets a 404 whose body lists the endpoints that exist.

{
"error": "Unknown endpoint: code-principale",
"available": ["code-principal", "completion-rapide"]
}

Other refusals to recognize:

ResponseWhat it means
401key missing or unknown to the gateway
403the endpoint is disabled, or no role on the account carries the right to call it
500 naming a variablethe engine expects a key whose variable is missing from the gateway environment

The serverless build profile adds two refusals, tied to spending plans: 403 for an account with no plan, 402 for an account that has reached its plan’s monthly cap. They are cleared by different actions: Cap a customer’s spending. A service that does not respond at all belongs to The gateway refuses to start.

The gateway accepts the key under Authorization: Bearer <clé> as well as under x-api-key: <clé>.

The two halves of the path are configured separately.

From the workstation to the gateway, the IDE goes through a corporate proxy: the configuration accepts requestOptions with proxy, noProxy, caBundlePath and clientCertificate, at the root level as well as per model.

From the gateway to the inference engine, inside your perimeter, the call is configured through environment variables on the gateway service:

VariableWhat it sets
HTTPS_PROXY, HTTP_PROXYthe internal proxy to go through to reach the engine
NO_PROXYthe destinations to reach directly, with wildcards, suffixes and ports
NODE_EXTRA_CA_CERTSan internal certificate authority, added to Node’s store
UPSTREAM_TLS_CA_FILEan authority specific to the gateway’s outbound path
UPSTREAM_TLS_CLIENT_CERT_FILE, UPSTREAM_TLS_CLIENT_KEY_FILEthe client certificate, when the engine requires mTLS

Lowercase names (https_proxy, http_proxy) are accepted, with the same precedence as in the rest of the product.

An invalid setting stops startup: a proxy variable that is not an address, or an unreadable NODE_EXTRA_CA_CERTS, are rejected at launch, with the name of the offending variable.

If you configure mTLS, the client certificate does not override corporate trust: the trust store is assembled, and NODE_EXTRA_CA_CERTS still applies.

If the response arrives all at once instead of being written out

Section titled “If the response arrives all at once instead of being written out”

A proxy that buffers responses breaks word-by-word display, whatever the gateway does. The developer sees a silent assistant for several seconds, then a response that appears all at once. The total time is normal; its distribution changes.

This is the default behavior of many appliances that inspect content: inspecting means having the whole response before relaying it.

To check this on your side, compare the same call with and without the proxy:

  1. From a workstation that goes through the proxy, send a request whose response is long.
  2. Repeat the same request from a workstation or a network that does not go through it.

If the words are written out progressively in the second case and not in the first, the buffering is in the proxy. There is nothing to fix on the product side.

On the other half of the path, between the gateway and the engine, the gateway measures the arrival rate of long responses. When several responses arrive in a single burst, it writes this to its error output under the [upstream-buffering] prefix, once, then every hundred occurrences. Its silence proves nothing: it does not see the path between the workstation and itself.

The setting depends on the appliance. For example, mitmproxy buffers by default and relays progressively with --set stream_large_bodies=1. Look for the equivalent in yours: the notion is called streaming, pass-through or no buffering.

For support, this is the first question to ask when a customer reports an assistant that is slow to respond while server-side measurements are good.

To weigh when choosing which network to put the gateway on.

The path from the workstation to the gateway carries the requests and the developer’s key. It is encrypted with TLS (security white paper V3, section 6.2): the gateway terminates TLS when you give it TLS_CERT_FILE and TLS_KEY_FILE, and the workstation’s apiBase then uses https://. Without those two variables, the gateway listens in cleartext and announces it at startup; a TLS terminator placed in front of it must then encrypt the workstation traffic.

The engine’s address is data in the database, read from the base_url column of the declared engine. The console checks the shape of that address (an absolute URL, in https://, or in http:// with explicit acceptance), not its destination. Anyone who can write to the database can direct requests to another address. What protects this column is the console: it requires a role per route, can delegate authentication to your organization’s directory, and logs administration actions (granting a role, removing one, revoking a subject). A shared pair of bootstrap credentials is still tolerated for the first startup; remove it once roles are in place.

The configuration accepts a provider named lemniscate-proxy, handled separately in the schema, with its own orgScopeId and onPremProxyUrl fields. This is not the path described here, nor the way to reach your gateway.

Its implementation points by default to a service hosted by the vendor, whose existence and operation nothing in the repository attests to; for that reason the surface inventory classifies it as indeterminate. The onPremProxyUrl mode is self-contained, and nothing in the repository makes it possible to write its procedure without inventing it.

On the other legacy surfaces that talk to this same unoperated service: What does not exist.