Version 1.0.0
The lifecycle of an agent session
The stages of an agent session, from instruction to approved diff, when you are consulted, and what bounds the session.
Asking a language model a question gives you an answer. Handing a task to an agent opens a session: the model requests an action, the action runs in a sandbox, its result comes back, and the model requests the next one. A task can take several dozen round trips.
This page describes that session: its stages, what happens in the loop, when you are consulted, and what bounds it. There is no permanent agent: a session is born from an instruction and ends with its sandbox.
The five stages of a session
Section titled “The five stages of a session”A session goes through five stages, each handled by a conventional component rather than by the model (security white paper V3, section 3.2).
- Authentication. You are authenticated against your organization’s directory, and the session is tied to a project.
- Project policy. The gateway loads that project’s policy: active tools, free commands, commands requiring approval, budgets.
- Sandbox opening. The gateway opens a sandbox on the execution host, with a working copy of the repository.
- Agent work. The agent explores the code, modifies the copy and runs the permitted commands. Every exchange with the model goes through the gateway.
- Diff, approval, log. The result comes back to you as a diff, which you accept or reject. The whole session is logged, then the sandbox is destroyed.
The code is read by exploring the repository, with no prior index.
One turn of the loop: the life of a tool call
Section titled “One turn of the loop: the life of a tool call”During the fourth stage, the agent works through tool calls. A “tool” is an action the model can request: read a file, write one, run a command, search the code. Each tool request goes through six possible states, which the product names.
Three properties can be read off this diagram.
A tool does not run from the state where the model is writing. There is an
intermediate state, generated, where the request is complete and nothing has been
done yet. The decision to execute is taken in that state. What the model produces
is a proposal, not an action.
A malformed request is rejected. Arguments arrive from the model in fragments; the IDE parses them once, complete, and strictly. Unparsable text becomes an error returned to the model, not an empty object executed by guessing the intent. The terminal does not follow this rule: it parses the fragments as they arrive.
An error does not stop the session. errored returns the error to the model the way
done returns the result: in both cases the loop resumes, and the model can
correct itself. What does stop the session is described in
What bounds the loop.
What this diagram does not show
Section titled “What this diagram does not show”- Where the action runs. The diagram describes the lifecycle of a request, not where it runs; see the next section.
- Several tools at once. A single turn can carry several requests; the loop resumes only when all of them have left their intermediate states.
- Cancellation coming from elsewhere. Interrupting a response in progress moves
requests still in
generatingorgeneratedto thecanceledstate, without a rejection having been expressed on each one. - Resuming after a rejection. The
continueAfterToolRejectionsetting, off by default, decides whether your rejection restarts the loop (the model receives a rejection message and continues) or stops it. - Compaction. When the history gets too long, it is summarized in the middle of the loop. This does not change the state of a tool call. In the terminal, a compaction consumes one action from the session budget; in the IDE, it does not.
- The terminal moves a tool forbidden by policy to
canceled, not toerrored.
Where the action runs
Section titled “Where the action runs”Every action runs in a sandbox, created for the session and destroyed when it ends (security white paper V3, section 4.2). The sandbox has no network interface, runs unprivileged, on a read-only root filesystem, and shares no state with another session or another project. Exchanges with the model are relayed by the gateway.
The sandbox runs in one of two places (security white paper V3, section 4.4):
- on a dedicated execution host, the reference mode. The sandbox works on a copy of the repository, the workstation carries only the extension, and the agent does not hold the user’s rights on their workstation;
- on the workstation, the variant. The container is unprivileged and without network, the local repository is mounted as its only writable volume, and the workstation’s toolchains are available. This suits workstation baselines that already allow a container engine.
The location is chosen with the LEMNISCATE_AGENT_EXECUTION variable, which accepts two values:
server (the sandbox on the execution host) and workstation (the container on the
workstation, the value used when the variable is absent). An unknown value is
rejected rather than falling back to the default. The choice is made per
perimeter, with the security officer.
What the agent sees in the sandbox, and what it does not carry, is described in Execution safeguards.
What decides to consult you
Section titled “What decides to consult you”Between generated and calling, the project policy decides. It is written by the
administrator, evaluated by the gateway, and not interpreted by the model
(security white paper V3, section 5.1). For each tool, it has three values, from
the most restrictive to the least restrictive:
| Policy | Behavior |
|---|---|
disabled | the tool does not run, and the model is told |
allowedWithPermission | you are consulted before each execution |
allowedWithoutPermission | the tool runs without consulting you |
The default is allowedWithPermission: without a decision from the administrator, no command is
free, and the product asks you. The terminal names the same three values exclude,
ask and allow, with ask by default.
Two properties of this mechanism matter as much as the list.
An evaluation made at execution time cannot loosen the base policy. Some tools
are judged a second time based on what exactly is being asked of them, a terminal
command for example. That second judgment can harden the verdict, not soften it.
A tool set to disabled does not reopen because the requested command looks
harmless. An evaluation that fails yields disabled.
The decision is attributed. Each execution records whether the approval came from a human or from an automation. There is no third value, and no default value.
What reads a command before running it
Section titled “What reads a command before running it”The security boundary is the sandbox, not the list of permitted commands: a permitted interpreter runs anything (security white paper V3, section 4.2). The command analysis described here is a matter of hygiene, and comes on top of confinement.
A terminal command is not judged as a block. It is split up: lines, chains,
pipes. Each piece is judged separately, and the most restrictive verdict among
all the pieces applies. A harmless command followed by a dangerous command is a
dangerous command. A redirection that writes to a file requires your approval, as
does a pipe into an interpreter (sh, bash, python).
Several families are refused with no approval able to unblock them:
- privilege escalation (
sudo,su,doasand their Windows equivalents); - destruction of the system tree: recursive forced deletion of a root or a system
directory, deletion of system files, direct writes under
/dev/, formatting; - opening up permissions (
chmod 777,setuidbit) and changing ownership toroot; - loading kernel modules and modifying the firewall;
evalandexec.
Two details matter in practice. A command whose name comes from a variable is treated as unknown, and therefore requires your approval: the real name is only readable at execution time. A command about which nothing is known requires your approval; it is neither refused nor let through.
In the sandbox, the interpreter is bash, launched without loading a profile.
Environment variables that can carry executable code (BASH_ENV, ENV, SHELLOPTS,
BASHOPTS, ZDOTDIR, and any BASH_FUNC_* variable) are removed before execution. No
environment variable from your workstation is passed in.
Delegating to a sub-agent
Section titled “Delegating to a sub-agent”While working, the agent can hand part of the task to a sub-agent. The sub-agent runs in its own session, with its own context and its own bounds, does not widen the rights of the delegating agent, and returns a report. Its permission requests come back to you. The rules are explained in Delegating to sub-agents, and the procedure in Delegating a task to a sub-agent.
The diff and its approval
Section titled “The diff and its approval”What the session produces comes out through a single channel: a code diff, presented to you, which you accept or reject. No pending write is applied to your repository without this approval (security white paper V3, sections 1.2 and 4.1).
The agent has no account and pushes nothing. The approved diff is applied to your repository, and you integrate it under your own identity, with your usual tools and through your review process (security white paper V3, section 4.3).
The approval request shows the diff or the exact command, and flags files that fall under dedicated control: continuous integration configuration, build scripts, hooks, dependency manifests. See Execution safeguards.
What bounds the loop
Section titled “What bounds the loop”Every session has a planned end. Four budgets bound each session: duration, volume of exchanges with the model, number of commands, and sandbox resources (processor, memory, processes). Exceeding them causes a clean, logged stop (security white paper V3, sections 4.2 and 4.5).
The session budget is evaluated before each call to the model. It counts:
- actions: each tool call consumes one (150 by default);
- time elapsed since the start of the session (45 minutes by default);
- a number of extensions (5 by default), each granting 50 more actions and 15 more minutes.
These values are set with the LEMNISCATE_SESSION_MAX_ACTIONS, LEMNISCATE_SESSION_MAX_MINUTES and LEMNISCATE_SESSION_MAX_EXTENSIONS variables, or with the
execution.sessionMaxActions, execution.sessionMaxMinutes and execution.sessionMaxExtensions configuration keys.
When the budget runs out, in an interactive session, the session asks you for an extension. The message lists what has been consumed and what the next extension grants. In the IDE, the extension is granted by sending a new message. Once the number of extensions is exhausted, the session stops, and you have to open a new one. In the terminal, the budget restarts from zero with each user message, and the stopped session produces a final summary, with no tools.
Outside an interactive session, there is no one to answer: the session stops and exits with code 3; see Execution safeguards.
Two practical consequences:
- the budget does not survive an IDE restart. A bound restored from yesterday’s session would be a false bound;
- an unreadable budget value is rejected rather than falling back to the default: a fatal configuration error in the IDE, an error on the first read of the bounds in the terminal. A setting that loosens or tightens the default is announced at startup.
Commands have their own, independent budget: a maximum duration (10 minutes by
default, LEMNISCATE_COMMAND_MAX_SECONDS) and a maximum output volume (10 MiB by default, LEMNISCATE_COMMAND_MAX_OUTPUT_BYTES). What
exceeds the volume is counted then discarded, not accumulated in memory.
Stopping and revocation
Section titled “Stopping and revocation”Any session can be stopped immediately (security white paper V3, section 4.5).
- You have a one-gesture stop from your environment. It interrupts the work in progress, destroys the sandbox and discards unapproved writes.
- The administrator has a revocation from the console, per session or per user. It takes effect without a restart, it is logged, and it destroys the sandbox.
What stays written down
Section titled “What stays written down”Every session is logged from the initial instruction to the result of each command: authentication, the instruction, files read and modified, actions proposed by the agent, your decisions (approval, rejection, stop), executions in the sandbox with their return code and duration, and the policy in force (security white paper V3, section 7.2). Each bound extension also leaves an entry there. In the terminal, a session frame is written before the first act.
Logs are append-only. Entries are chained: each one carries the digest of the previous one, so modifying a past entry breaks the chain and is detected. Each event is timestamped and carries the identity from the directory. Logs are exported to your organization’s SIEM (security white paper V3, section 7.1).
By default, the log contains neither the text of the instruction, nor the code, nor the command output: it does not create a copy of the source code. Your organization can enable content retention, with separate encryption and retention rules. A detected secret is not logged in clear text.