Skip to content
Version 1.0.0

Discover the gateway

Bring up the gateway services on your workstation, route a call through them to a local inference engine, and confirm the call went through the gateway.

The gateway services are three: the gateway, its database and the admin console. The gateway authenticates, applies the project policy, drives the session, logs and revokes; the database and the console serve it (security white paper V3, sections 3.1 and 06). This tutorial brings them up on your workstation to show you the first of these functions: a call to the model that is authenticated, relayed and attributed. It belongs to the Install the complete code assistant path, and it also serves anyone who wants to integrate Lemniscate into their stack.

This page is not a production procedure, and the repository does not contain one. It brings up each piece separately: gateway-db/docker-compose.yml for the database, one image file for the gateway, another for the console. It versions two Kubernetes manifests, durcissement/kubernetes/llm-gateway.yaml and durcissement/kubernetes/gateway-admin.yaml, which set the runtime constraints and nothing else: the image, the secrets, the probes and the resources belong to the customer. No file describes the three together: no composition, no installer, no verified sequence of steps. The only complete versioned deployment path targets one particular host. This page brings the three components up on your workstation, so you can see how they fit together before you operate them.

The setup you end up with is not secure: the gateway listens over plain HTTP, the admin console has a single set of credentials shared by everyone who opens it, and the database passwords come from example files published in the repository. Do not reproduce it anywhere but on your own machine.

In the reference architecture, the gateway connects to the inference engine inside your perimeter through an OpenAI-compatible API, the model is an open-weights model, and no traffic leaves (security white paper V3, sections 3.3 and 3.4). This tutorial follows that path: Ollama, on your workstation, plays the role of the inference server, and serves the open-weights model qwen2.5-coder:14b through its OpenAI-compatible API. No call leaves the machine.

The setup simplifies the reference architecture on one point: the four roles all sit on a single machine. Here is what that changes compared with a deployment.

  • The three gateway services run on a dedicated host, separate from the workstation, the execution host and the inference server. Here, they share your workstation with the engine.
  • Traffic between components is encrypted with TLS, and your network filtering only allows the gateway host to reach the engine (security white paper V3, sections 6.2 and 8.2). Here, the gateway reaches the engine over HTTP on the machine’s local interface, where Ollama listens by default.
  • Identities come from your directory. Here, an admin key and a pair of credentials stand in for them.
  • An agent session runs in a sandbox, on the execution host. This tutorial stops at the model call: it opens no session.

Allow about thirty minutes. You need three terminals: the steps below call them A, B and C, and each one starts at the root of your copy of the repository. Terminal A is used for steps 1 to 5 and then 13 to 15; B and C each carry a service that stays in the foreground.

Terminal A. gateway-db/docker-compose.yml reads a .env file that is not versioned. Only the example is:

Fenêtre de terminal
cd gateway-db
cp .env.example .env

Expected result: a gateway-db/.env file carrying POSTGRES_USER=admin, POSTGRES_PASSWORD=admin_secret and POSTGRES_DB=proxy, along with the development passwords of the two service roles. These passwords are published in the repository: they are only valid for a discovery workstation.

Still in terminal A, in gateway-db:

Fenêtre de terminal
docker compose up -d --wait

Expected result: the command returns once the gateway-db container is healthy, and docker compose ps shows it running, port 6000 on your machine mapped to 5432 in the container. This first startup, and only this one, runs init.sql, which lays down the full schema, then dev-role-passwords.sql, which gives a development password to the two service roles lemniscate_gateway and lemniscate_console. The database is empty: no engine declared, no model, no endpoint, no account. Everything the gateway serves, you declare yourself over the course of this tutorial; this is also what an operator sees on a fresh installation.

init.sql describes the target state of the schema, but it does not maintain the migration registry. Replaying the migrations/ sequence on this database would fail on the very first one, which recreates existing tables. So you record the migrations as applied, without running them:

Fenêtre de terminal
export DATABASE_URL=postgresql://admin:admin_secret@localhost:6000/proxy
./migrate.sh --baseline 029

Expected result: one inscrite sans exécution : … line per migration, then jalon posé jusqu'à 029, followed by the number of records written. Stamping is a one-time operation, reserved for a database whose schema is already up to date.

The stamp number is that of the last migration in the repository: the stamp must cover everything init.sql has already laid down. A stamp that is too low leaves migrations to be applied on a schema that already contains them. The authoritative list is the one printed by ./migrate.sh --status.

Fenêtre de terminal
./migrate.sh

Expected result: base à jour, aucune migration à appliquer. From here on, this command alone applies the migrations written after these.

The gateway built for this setup refuses to start without ADMIN_API_KEY: the database seeds no account, and it is this variable that creates the admin account on first startup.

Fenêtre de terminal
export ADMIN_API_KEY="$(printf 'sk-lemniscate-%s' "$(openssl rand -hex 24)")"
echo "$ADMIN_API_KEY"

Expected result: a line starting with sk-lemniscate- followed by 48 hexadecimal characters. Copy it: terminal C needs it at step 12, and it serves as the call key at step 14.

Terminal B, at the repository root:

Fenêtre de terminal
cd gateway-admin
npm install
LEMNISCATE_DEPLOYMENT_PROFILE=serverless npm run build

Expected result: the build announces Construction de gateway-admin... (profil : serverless), then produces dist/client.css, dist/client.js and dist/server.js. Without this build, the console has nothing to serve.

Still in terminal B, in gateway-admin. ADMIN_USER and ADMIN_PASS are mandatory: without them, the console refuses to start. DATABASE_URL points to the console’s service role, lemniscate_console: the console refuses to start with the database owner account.

Fenêtre de terminal
DATABASE_URL=postgresql://lemniscate_console:console_secret@localhost:6000/proxy \
ADMIN_USER=admin ADMIN_PASS=change-me npm run start

Expected result: Admin UI running at http://localhost:6002, followed by two lines that state what guards the console and which account repository it serves. Open that address; the browser asks for a username and a password, which are the two values above. This pair is the only one in the installation, and anyone who holds it holds the whole console.

The gateway relays calls to an inference engine, which the console calls a Provider. The database contains none: declaring it is your first act as an operator.

Terminal A. First check that the engine responds and that it serves the model:

Fenêtre de terminal
curl -s http://127.0.0.1:11434/v1/models

Expected result: a JSON object whose data list contains an entry with the identifier qwen2.5-coder:14b. If there is no response, start Ollama before continuing.

In the console, under the Endpoints entry in the sidebar, section Providers, expand Add provider and fill in:

  • Slug: moteur-local
  • Name: Moteur local
  • Base URL: http://127.0.0.1:11434
  • API Key Env: leave empty
  • API Key Header: leave empty

The address starts with http://: a checkbox appears below the form. Its text states that requests and responses travel in the clear over this hop, and ends with I accept this for this provider. Tick it, then confirm with Add. Without the checkbox, the console refuses the declaration.

Expected result: a moteur-local line appears in the Providers table, with its address.

Accepting the cleartext hop applies only to this engine, and to this setup where the hop does not leave the machine. In a deployment, the engine is declared in https://.

The two key fields stay empty because Ollama does not ask for one. For an engine that does require a key, API Key Env holds the name of a gateway environment variable, not the key itself.

Same entry, section Models, expand Add model and fill in:

  • Slug: qwen2.5-coder:14b
  • Name: Qwen2.5 Coder 14B
  • Provider: Moteur local
  • Input price and Output price: leave at 0

Confirm with Add.

Expected result: a line appears in the Models table, attached to the moteur-local engine. The slug is the identifier the gateway sends to the engine: it must be the one step 8 read in the engine’s list.

Same entry, section Endpoints, expand Add endpoint:

  • Name: code-principal
  • Model: Qwen2.5 Coder 14B

Confirm with Add.

Expected result: a code-principal line appears in the Endpoints table, active. The endpoint name is the first segment of the URL callers use; it determines both the engine reached and the model used. It names a use, not a model: you can change the model served without changing this name.

Terminal C, at the repository root:

Fenêtre de terminal
cd llm-gateway
npm install
LEMNISCATE_DEPLOYMENT_PROFILE=serverless npm run build

Expected result: npm install completes without error, then the build announces Construction du llm-gateway... (profil : serverless), lists the artifacts produced in dist/ and ends with Construction terminée.

The profile matters here for the same reason as at step 6. The closed artifact, the one npm run build produces with no profile and the one the no-network-egress check inspects, seeds no account from ADMIN_API_KEY and requires a deployment license. That is not the artifact for this setup.

Still in terminal C, in llm-gateway. Both variables are mandatory. DATABASE_URL points to the gateway’s service role, lemniscate_gateway: the gateway refuses to start with the database owner account.

Fenêtre de terminal
DATABASE_URL=postgresql://lemniscate_gateway:gateway_secret@localhost:6000/proxy \
ADMIN_API_KEY=<la clé affichée à l'étape 5> \
npm run start

Expected result: the startup log begins with these lines.

[ADMIN] administrator account "admin" created from ADMIN_API_KEY
[ADMIN] administrator account "admin" put on the unlimited "administrator" plan
[ÉCOUTE] PLAINTEXT — http://localhost:6001. No TLS_CERT_FILE/TLS_KEY_FILE was given, …
[MOTEUR] no client certificate configured (…)

The first says the admin account has just been created, with your key. The second says it is put on the unlimited plan: the gateway answers 403 to an account with no plan.

The next two are warnings, not errors. The third says the port is open and that nothing encrypts this link. On this setup, the step 14 call does not leave the machine. In a deployment, the gateway receives a certificate and listens over TLS (security white paper V3, section 6.2). The fourth says it presents no client certificate to the engine. As long as the [ÉCOUTE] line announces the port, the gateway is listening. Other announcement lines follow, one per subsystem: agentic execution, network egress, documentation, governance, request timeout.

If startup stops on a message instead of these lines, read The gateway refuses to start.

Your account exists, it is on an unlimited plan, and it still has the right to do nothing. These are two separate questions: the plan says what you can spend, the role says what you can do. The gateway grants no role automatically, not even to admin.

Terminal A:

Fenêtre de terminal
cd llm-gateway
DATABASE_URL=postgresql://admin:admin_secret@localhost:6000/proxy \
npm run attribuer-un-role -- admin utilisateur-standard "tutoriel"

Expected result: rôle « utilisateur-standard » attribué à « admin » par « tutoriel ».

Without this step, the call in the next step receives Your identity is recognised, but no role you hold carries the right to call this endpoint. This refusal names what is missing; it is not a key problem.

The third argument is the author of the grant. It is kept with the grant: on a real installation, knowing who gave a right matters as much as knowing who holds it.

Terminal A, the one where ADMIN_API_KEY is still exported:

Fenêtre de terminal
curl -s http://localhost:6001/code-principal/v1/chat/completions \
-H "authorization: Bearer $ADMIN_API_KEY" \
-H "content-type: application/json" \
-d '{"model":"ignore","max_tokens":16,"messages":[{"role":"user","content":"Dis bonjour."}]}'

Expected result: the engine’s JSON response, in OpenAI-compatible format. The model’s text is in choices[0].message.content, and the usage field carries the token counters:

{
"id": "chatcmpl-678",
"object": "chat.completion",
"created": 1791223266,
"model": "qwen2.5-coder:14b",
"system_fingerprint": "fp_ollama",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Bonjour! Comment puis-je vous aider aujourd'hui?"
},
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 33, "completion_tokens": 12, "total_tokens": 45 }
}

The first call waits for the engine to load the model into memory: depending on the hardware, the response can take more than a minute to arrive. Later calls are faster.

Three details of this command are worth reading. The gateway strips the first segment of the URL, the endpoint name, and appends the rest, /v1/chat/completions, to the engine’s base address. The model field in the body is set to ignore: the gateway replaces it with the endpoint’s model, here qwen2.5-coder:14b. The endpoint decides the model, not the caller. The key in the authorization header is your account’s key on the gateway: the gateway checks it, then strips the header before relaying.

15. Confirm the call went through the gateway

Section titled “15. Confirm the call went through the gateway”

Two traces, independent of each other.

Look at terminal C first: the gateway logged the call.

[DEBUG] POST http://localhost:6001/code-principal/v1/chat/completions
[COST] admin: €0.000000 used (plan has no cap)
admin -> code-principal (qwen2.5-coder:14b)

Then, in terminal A, read back what the gateway wrote to the database:

Fenêtre de terminal
psql "$DATABASE_URL" -c "SELECT u.username, r.model, r.input_tokens, r.output_tokens, r.latency_ms FROM request_logs r JOIN users u ON u.id = r.user_id;"

Expected result: one line, and only one: admin, the endpoint’s model, the two token counters reported by the engine, and the duration of the exchange. This table carries the attribution of each relayed call to an account.

A call addressed to the gateway was authenticated, attached to an account, checked against that account’s rights, relayed to the inference engine under the model the endpoint designates, and recorded in the database under its name. The caller only knew the gateway’s address and an endpoint name: not the engine’s address, not the model’s name. The model is an open-weights model served on the machine, and no call left it.

You have also seen the configuration points that have no default value: DATABASE_URL and ADMIN_API_KEY block gateway startup as long as they are missing, and the engine’s address is read from the database, not from the code.

A development workstation targets this same endpoint, using the key belonging to its holder. The dialect is that of the OpenAI-compatible API:

- name: Code principal (passerelle)
provider: openai
model: any
apiBase: http://localhost:6001/code-principal/v1
apiKey: <la clé du porteur>

To shut the setup down: Ctrl+C in terminals B and C, then docker compose down in gateway-db, or docker compose down -v to erase the data as well.

When a startup stops on a message, or when the gateway starts but every request fails, the procedure to follow is in The gateway refuses to start. To understand what going through a gateway brings, read The role of the gateway.

To declare the inference engine in your perimeter instead of Ollama, with its address in https://, follow Declare a model served by the gateway. Connecting a workstation is covered by Your first session in VS Code.