Version 1.0.0
Discover the gateway
Bring up the gateway services on your workstation, route a call through them to a local inference engine, and confirm the call went through the gateway.
The gateway services are three: the gateway, its database and the admin console. The gateway authenticates, applies the project policy, drives the session, logs and revokes; the database and the console serve it (security white paper V3, sections 3.1 and 06). This tutorial brings them up on your workstation to show you the first of these functions: a call to the model that is authenticated, relayed and attributed. It belongs to the Install the complete code assistant path, and it also serves anyone who wants to integrate Lemniscate into their stack.
This page is not a production procedure, and the repository does not contain one. It brings
up each piece separately: gateway-db/docker-compose.yml for the database, one image file
for the gateway, another for the console. It versions two Kubernetes manifests,
durcissement/kubernetes/llm-gateway.yaml and durcissement/kubernetes/gateway-admin.yaml,
which set the runtime constraints and nothing else: the image, the secrets, the probes and
the resources belong to the customer. No file describes the three together: no composition,
no installer, no verified sequence of steps. The only complete versioned deployment path
targets one particular host. This page brings the three components up on your workstation,
so you can see how they fit together before you operate them.
The setup you end up with is not secure: the gateway listens over plain HTTP, the admin console has a single set of credentials shared by everyone who opens it, and the database passwords come from example files published in the repository. Do not reproduce it anywhere but on your own machine.
In the reference architecture, the gateway connects to the inference engine inside your
perimeter through an OpenAI-compatible API, the model is an open-weights model, and no
traffic leaves (security white paper V3, sections 3.3 and 3.4). This tutorial follows that
path: Ollama, on your workstation, plays the role of the inference server, and serves the
open-weights model qwen2.5-coder:14b through its OpenAI-compatible API. No call
leaves the machine.
The setup simplifies the reference architecture on one point: the four roles all sit on a single machine. Here is what that changes compared with a deployment.
- The three gateway services run on a dedicated host, separate from the workstation, the execution host and the inference server. Here, they share your workstation with the engine.
- Traffic between components is encrypted with TLS, and your network filtering only allows the gateway host to reach the engine (security white paper V3, sections 6.2 and 8.2). Here, the gateway reaches the engine over HTTP on the machine’s local interface, where Ollama listens by default.
- Identities come from your directory. Here, an admin key and a pair of credentials stand in for them.
- An agent session runs in a sandbox, on the execution host. This tutorial stops at the model call: it opens no session.
Allow about thirty minutes. You need three terminals: the steps below call them A, B and C, and each one starts at the root of your copy of the repository. Terminal A is used for steps 1 to 5 and then 13 to 15; B and C each carry a service that stays in the foreground.
1. Give the database its credentials
Section titled “1. Give the database its credentials”Terminal A. gateway-db/docker-compose.yml reads a .env file that is not versioned.
Only the example is:
cd gateway-dbcp .env.example .envExpected result: a gateway-db/.env file carrying POSTGRES_USER=admin,
POSTGRES_PASSWORD=admin_secret and POSTGRES_DB=proxy, along with the development
passwords of the two service roles. These passwords are published in the repository: they
are only valid for a discovery workstation.
2. Bring up the database
Section titled “2. Bring up the database”Still in terminal A, in gateway-db:
docker compose up -d --waitExpected result: the command returns once the gateway-db container is healthy, and
docker compose ps shows it running, port 6000 on your machine mapped to 5432 in the
container. This first startup, and only this one, runs init.sql, which lays down the full
schema, then dev-role-passwords.sql, which gives a development password to the two
service roles lemniscate_gateway and lemniscate_console. The database is empty: no
engine declared, no model, no endpoint, no account. Everything the gateway serves, you
declare yourself over the course of this tutorial; this is also what an operator sees on a
fresh installation.
3. Stamp the migration registry
Section titled “3. Stamp the migration registry”init.sql describes the target state of the schema, but it does not maintain the migration
registry. Replaying the migrations/ sequence on this database would fail on the very first
one, which recreates existing tables. So you record the migrations as applied, without
running them:
export DATABASE_URL=postgresql://admin:admin_secret@localhost:6000/proxy./migrate.sh --baseline 029Expected result: one inscrite sans exécution : … line per migration, then
jalon posé jusqu'à 029, followed by the number of records written. Stamping is a one-time
operation, reserved for a database whose schema is already up to date.
The stamp number is that of the last migration in the repository: the stamp must cover
everything init.sql has already laid down. A stamp that is too low leaves migrations to be
applied on a schema that already contains them. The authoritative list is the one printed
by ./migrate.sh --status.
4. Apply the pending migrations
Section titled “4. Apply the pending migrations”./migrate.shExpected result: base à jour, aucune migration à appliquer. From here on, this
command alone applies the migrations written after these.
5. Generate the admin key
Section titled “5. Generate the admin key”The gateway built for this setup refuses to start without ADMIN_API_KEY: the database
seeds no account, and it is this variable that creates the admin account on first
startup.
export ADMIN_API_KEY="$(printf 'sk-lemniscate-%s' "$(openssl rand -hex 24)")"echo "$ADMIN_API_KEY"Expected result: a line starting with sk-lemniscate- followed by 48 hexadecimal
characters. Copy it: terminal C needs it at step 12, and it serves as the call key at
step 14.
6. Build the admin console
Section titled “6. Build the admin console”Terminal B, at the repository root:
cd gateway-adminnpm installLEMNISCATE_DEPLOYMENT_PROFILE=serverless npm run buildExpected result: the build announces
Construction de gateway-admin... (profil : serverless), then produces dist/client.css,
dist/client.js and dist/server.js. Without this build, the console has nothing to
serve.
7. Bring up the admin console
Section titled “7. Bring up the admin console”Still in terminal B, in gateway-admin. ADMIN_USER and ADMIN_PASS are
mandatory: without them, the console refuses to start. DATABASE_URL points to the
console’s service role, lemniscate_console: the console refuses to start with the
database owner account.
DATABASE_URL=postgresql://lemniscate_console:console_secret@localhost:6000/proxy \ ADMIN_USER=admin ADMIN_PASS=change-me npm run startExpected result: Admin UI running at http://localhost:6002, followed by two lines that
state what guards the console and which account repository it serves. Open that address;
the browser asks for a username and a password, which are the two values above. This pair
is the only one in the installation, and anyone who holds it holds the whole console.
8. Declare the inference engine
Section titled “8. Declare the inference engine”The gateway relays calls to an inference engine, which the console calls a Provider. The database contains none: declaring it is your first act as an operator.
Terminal A. First check that the engine responds and that it serves the model:
curl -s http://127.0.0.1:11434/v1/modelsExpected result: a JSON object whose data list contains an entry with the identifier
qwen2.5-coder:14b. If there is no response, start Ollama before continuing.
In the console, under the Endpoints entry in the sidebar, section Providers, expand Add provider and fill in:
- Slug:
moteur-local - Name:
Moteur local - Base URL:
http://127.0.0.1:11434 - API Key Env: leave empty
- API Key Header: leave empty
The address starts with http://: a checkbox appears below the form. Its text states that
requests and responses travel in the clear over this hop, and ends with
I accept this for this provider. Tick it, then confirm with Add. Without the
checkbox, the console refuses the declaration.
Expected result: a moteur-local line appears in the Providers table, with its
address.
Accepting the cleartext hop applies only to this engine, and to this setup where the hop
does not leave the machine. In a deployment, the engine is declared in https://.
The two key fields stay empty because Ollama does not ask for one. For an engine that does require a key, API Key Env holds the name of a gateway environment variable, not the key itself.
9. Declare the model
Section titled “9. Declare the model”Same entry, section Models, expand Add model and fill in:
- Slug:
qwen2.5-coder:14b - Name:
Qwen2.5 Coder 14B - Provider:
Moteur local - Input price and Output price: leave at
0
Confirm with Add.
Expected result: a line appears in the Models table, attached to the
moteur-local engine. The slug is the identifier the gateway sends to the engine: it
must be the one step 8 read in the engine’s list.
10. Create the endpoint
Section titled “10. Create the endpoint”Same entry, section Endpoints, expand Add endpoint:
- Name:
code-principal - Model:
Qwen2.5 Coder 14B
Confirm with Add.
Expected result: a code-principal line appears in the Endpoints table, active.
The endpoint name is the first segment of the URL callers use; it determines both the
engine reached and the model used. It names a use, not a model: you can change the model
served without changing this name.
11. Install and build the gateway
Section titled “11. Install and build the gateway”Terminal C, at the repository root:
cd llm-gatewaynpm installLEMNISCATE_DEPLOYMENT_PROFILE=serverless npm run buildExpected result: npm install completes without error, then the build announces
Construction du llm-gateway... (profil : serverless), lists the artifacts produced in
dist/ and ends with Construction terminée.
The profile matters here for the same reason as at step 6. The closed artifact, the one
npm run build produces with no profile and the one the no-network-egress check
inspects, seeds no account from ADMIN_API_KEY and requires a deployment
license. That is not the artifact for this setup.
12. Bring up the gateway
Section titled “12. Bring up the gateway”Still in terminal C, in llm-gateway. Both variables are mandatory.
DATABASE_URL points to the gateway’s service role, lemniscate_gateway: the gateway refuses to
start with the database owner account.
DATABASE_URL=postgresql://lemniscate_gateway:gateway_secret@localhost:6000/proxy \ ADMIN_API_KEY=<la clé affichée à l'étape 5> \ npm run startExpected result: the startup log begins with these lines.
[ADMIN] administrator account "admin" created from ADMIN_API_KEY[ADMIN] administrator account "admin" put on the unlimited "administrator" plan[ÉCOUTE] PLAINTEXT — http://localhost:6001. No TLS_CERT_FILE/TLS_KEY_FILE was given, …[MOTEUR] no client certificate configured (…)The first says the admin account has just been created, with your key. The second says
it is put on the unlimited plan: the gateway answers 403 to an account with no plan.
The next two are warnings, not errors. The third says the port is open and that nothing
encrypts this link. On this setup, the step 14 call does not leave the machine. In a
deployment, the gateway receives a certificate and listens over TLS (security white paper
V3, section 6.2). The fourth says it presents no client certificate to the engine. As long
as the [ÉCOUTE] line announces the port, the gateway is listening. Other announcement lines
follow, one per subsystem: agentic execution, network egress, documentation, governance,
request timeout.
If startup stops on a message instead of these lines, read The gateway refuses to start.
13. Give your account the right to call
Section titled “13. Give your account the right to call”Your account exists, it is on an unlimited plan, and it still has the right to do nothing.
These are two separate questions: the plan says what you can spend, the role says what you
can do. The gateway grants no role automatically, not even to admin.
Terminal A:
cd llm-gatewayDATABASE_URL=postgresql://admin:admin_secret@localhost:6000/proxy \ npm run attribuer-un-role -- admin utilisateur-standard "tutoriel"Expected result: rôle « utilisateur-standard » attribué à « admin » par « tutoriel ».
Without this step, the call in the next step receives Your identity is recognised, but no role you hold carries the right to call this endpoint. This refusal names what is missing; it is not
a key problem.
The third argument is the author of the grant. It is kept with the grant: on a real installation, knowing who gave a right matters as much as knowing who holds it.
14. Route a call through the gateway
Section titled “14. Route a call through the gateway”Terminal A, the one where ADMIN_API_KEY is still exported:
curl -s http://localhost:6001/code-principal/v1/chat/completions \ -H "authorization: Bearer $ADMIN_API_KEY" \ -H "content-type: application/json" \ -d '{"model":"ignore","max_tokens":16,"messages":[{"role":"user","content":"Dis bonjour."}]}'Expected result: the engine’s JSON response, in OpenAI-compatible format. The model’s
text is in choices[0].message.content, and the usage field carries the token
counters:
{ "id": "chatcmpl-678", "object": "chat.completion", "created": 1791223266, "model": "qwen2.5-coder:14b", "system_fingerprint": "fp_ollama", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "Bonjour! Comment puis-je vous aider aujourd'hui?" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 33, "completion_tokens": 12, "total_tokens": 45 }}The first call waits for the engine to load the model into memory: depending on the hardware, the response can take more than a minute to arrive. Later calls are faster.
Three details of this command are worth reading. The gateway strips the first segment of
the URL, the endpoint name, and appends the rest, /v1/chat/completions, to the engine’s base
address. The model field in the body is set to ignore: the gateway replaces it with the
endpoint’s model, here qwen2.5-coder:14b. The endpoint decides the model, not the caller.
The key in the authorization header is your account’s key on the gateway: the gateway
checks it, then strips the header before relaying.
15. Confirm the call went through the gateway
Section titled “15. Confirm the call went through the gateway”Two traces, independent of each other.
Look at terminal C first: the gateway logged the call.
[DEBUG] POST http://localhost:6001/code-principal/v1/chat/completions[COST] admin: €0.000000 used (plan has no cap)admin -> code-principal (qwen2.5-coder:14b)Then, in terminal A, read back what the gateway wrote to the database:
psql "$DATABASE_URL" -c "SELECT u.username, r.model, r.input_tokens, r.output_tokens, r.latency_ms FROM request_logs r JOIN users u ON u.id = r.user_id;"Expected result: one line, and only one: admin, the endpoint’s model, the two token
counters reported by the engine, and the duration of the exchange. This table carries the
attribution of each relayed call to an account.
What you have established
Section titled “What you have established”A call addressed to the gateway was authenticated, attached to an account, checked against that account’s rights, relayed to the inference engine under the model the endpoint designates, and recorded in the database under its name. The caller only knew the gateway’s address and an endpoint name: not the engine’s address, not the model’s name. The model is an open-weights model served on the machine, and no call left it.
You have also seen the configuration points that have no default value:
DATABASE_URL and ADMIN_API_KEY block gateway startup as long as they are
missing, and the engine’s address is read from the database, not from the code.
A development workstation targets this same endpoint, using the key belonging to its holder. The dialect is that of the OpenAI-compatible API:
- name: Code principal (passerelle) provider: openai model: any apiBase: http://localhost:6001/code-principal/v1 apiKey: <la clé du porteur>To shut the setup down: Ctrl+C in terminals B and C, then docker compose down
in gateway-db, or docker compose down -v to erase the data as well.
When a startup stops on a message, or when the gateway starts but every request fails, the procedure to follow is in The gateway refuses to start. To understand what going through a gateway brings, read The role of the gateway.
To declare the inference engine in your perimeter instead of Ollama, with its address in
https://, follow
Declare a model served by the gateway.
Connecting a workstation is covered by
Your first session in VS Code.