Version 1.0.0
Declare a model served by the gateway
The inference engine, the model and the endpoint, in the order they must exist, and what the console refuses.
A gateway that has just been installed answers 404 to everything: it does not
serve any model yet. This page describes the three declarations that connect it
to your perimeter’s inference engine, the order in which they must exist, and
what the console refuses.
The gateway connects to any engine that exposes an OpenAI-compatible API, with
the open-weights model you have chosen. It is the only component that reaches
the engine: neither workstations nor sandboxes know its address (security white
paper V3, sections 3.4 and 8.2). The console names the engine
Provider; this page says “provider” when it refers to that label.
The examples on this page declare an engine at the address
https://inference.interne:8000, which serves the model qwen2.5-coder, and
publish it under the endpoint code-principal.
Everything happens in the Endpoints tab, present in both build profiles. The
console labels are in English; they are quoted here as you read them on screen.
The order of the three declarations
Section titled “The order of the three declarations”Fournisseur (Provider) → Modèle (Model) → Point d'appel (Endpoint)A model requires the provider that exposes it, an endpoint requires the model it
routes to, and the database refuses a row whose parent is missing. In the
interface, the Provider selector in the model form offers no option as long as
no provider is declared.
To undo, the order is reversed: an endpoint before its model, a model before its provider. The console refuses any deletion that would leave an orphan, and says so.
1. Declare the provider
Section titled “1. Declare the provider”Section Providers, button Add provider.
| Field | Required | What it carries |
|---|---|---|
Slug | yes | The engine’s short identifier, for example moteur-interne. Letters, digits, _ and - only. |
Name | yes | The readable name, shown in lists. |
Base URL | yes | The engine’s absolute address, scheme included, for example https://inference.interne:8000. |
API Key Env | no | The name of the environment variable that carries the engine’s key, not the key itself. |
API Key Header | no | The name of the HTTP header the key must be placed in. |
The two key fields go together. Left empty, the gateway relays without placing a key, which suits an engine that does not ask for one. Filled in both, the gateway reads the variable from its own environment at relay time, and places its value as-is in the named header. The console stores no key: it keeps the name of the variable.
The gateway builds the address it calls by appending to the Base URL whatever the
caller wrote after the endpoint name. A workstation that calls
/code-principal/v1/chat/completions therefore reaches
https://inference.interne:8000/v1/chat/completions. The /v1 segment is written in
only one of the two places: in the workstations’ call address, as here, or at
the end of the Base URL.
What the console refuses
Section titled “What the console refuses”| What you entered | Response |
|---|---|
| An empty required field | 400, slug, name, and base_url are required |
A Slug with a space, a period, an accent | 400, the accepted format is recalled in the message |
A Slug already taken | 409, A provider with this slug already exists |
A Base URL the gateway cannot read | 400, an absolute URL with its scheme is expected |
A scheme other than https:// or http:// | 400, the gateway only relays to these two schemes |
An engine reached over http://
Section titled “An engine reached over http://”The path from the gateway to the engine is TLS-encrypted (security white paper
V3, section 3.3): declare the engine over https://. A Base URL over http:// is
refused by default. The reason is written in the refusal: on that path, the
requests, the responses and the engine’s key travel in clear text, and the
gateway presents no client certificate to the engine; mutual authentication only
applies to engines over https://.
An exception exists, outside the reference architecture: as soon as the address
begins with http://, a checkbox appears in the form, and checking it counts as
accepting the clear-text path, for that provider only. The tutorial Discover
the gateway uses it for an engine running on the same machine as the
gateway.
Two points to know:
- There is no setting that lifts this check for all providers at once. Acceptance is given provider by provider.
- The box is reset after each successful addition: it does not carry over to the next declaration.
Correct a provider address
Section titled “Correct a provider address”The console exposes no way to change the Base URL, the key variable name or the
header of a provider already declared. To correct an address, delete the
provider and declare it again, which means deleting its models and their
endpoints first.
The same applies to models. Only endpoints can be modified.
2. Declare the model
Section titled “2. Declare the model”Section Models, button Add model.
| Field | Required | What it carries |
|---|---|---|
Slug | yes | The model name as the engine serves it. |
Name | yes | The readable name. |
Provider | yes | The provider that exposes it, chosen from the list. |
Input price (nano€/token) | no, defaults to 0 | Price of an input token, in nano-euros, whole number. |
Output price (nano€/token) | no, defaults to 0 | Price of an output token, same unit. |
The Slug is the name the gateway writes in the model field of each relayed
request, for example qwen2.5-coder. It must be the one the engine announces in its
model list.
The two prices are used by the spending plans of the serverless build profile. For a
model served by your own engine, leave them at 0: consumption is counted in
tokens and valued at zero. The unit is the nano-euro per token, that is, a
billionth of a euro.
Possible refusals: 400 if the slug, the name or the provider is missing; 409
if the slug is already taken; 400 with Provider not found if the provider has disappeared in
the meantime.
3. Declare the endpoint
Section titled “3. Declare the endpoint”Section Endpoints, button Add endpoint. Two fields: Name, and Model chosen from the list.
The name accepts letters, digits, _ and -; it is what clients write in
their call address. A name that describes a use, such as code-principal or completion-rapide, remains
valid when you replace the model served.
An endpoint is active as soon as it is created. The form offers no way to create it disabled.
Possible refusals: 400 if the name or the model is missing, 400 if the name
contains anything other than the accepted characters, 409 if the name is
already taken.
At this point, the gateway serves the model. The next step, pointing a workstation at this endpoint, is described in Route the IDE to the gateway.
Working with existing endpoints
Section titled “Working with existing endpoints”Reconnect an endpoint to another model
Section titled “Reconnect an endpoint to another model”In the Model column, click the model name: the cell becomes a selector. The
change is saved the moment you choose, with no validation button and no
confirmation.
This gesture changes the model served to a team without the team touching its configuration: the call address does not move, what sits behind it changes.
Enable or disable
Section titled “Enable or disable”The switch in the Enabled column writes the endpoint’s state. A disabled endpoint
is refused by the gateway before any routing: the calls that reach it reach no
engine. Its row stays visible in the console, greyed out.
Neither the reconnection nor the switch asks for confirmation, and a failure displays no message: the state returns to its original value when the table is refreshed. After either of these two gestures, reload the page to check that the console shows what you expect.
Delete
Section titled “Delete”Deleting an endpoint has two outcomes, and the console chooses:
| Situation | What happens |
|---|---|
| No call has ever been logged for it | The row is deleted. |
| Calls have been logged | The row is kept and disabled. |
The reason for the second case: the usage history references the endpoint.
Deleting it would detach that history, which attributes each call to an account
and which is used, in the serverless profile, to compute spending caps. In both
cases, the endpoint stops accepting calls.
The confirmation box announces both outcomes. The interface does not say which one happened: the row has disappeared, or it is greyed out.
Delete a model or a provider
Section titled “Delete a model or a provider”Both are refused as long as a child remains:
| What you are deleting | Response |
|---|---|
| A model that has endpoints | 409, Cannot delete model: it still has endpoints. Delete them first. |
| A provider that has models | 409, Cannot delete provider: it still has models. Delete them first. |
There is no cascading deletion: nothing disappears unless you asked for it by name.
The required permissions
Section titled “The required permissions”Each of these three families requires its own permission, and reading requires the same one as writing: listing the providers reveals their base addresses and the names of the variables carrying their keys, that is, the deployment topology. The exact permission names and the role that carries them are in the administration console reference.