SubLaneSubLane

Account pool runtime

SubLane guide: account pool runtime.

Use this to investigate busy accounts, cooldowns or conversation routing. For a first look at a failed call, start with Everyday use.

SubLane tracks model-request leases and transient account failures across every pool sharing an account. The administrator can inspect scheduling state from Accounts, change a per-account limit through Scheduling settings, and inspect all completed calls under Administration → All requests. Every signed-in user, including the administrator, has a personal General → Requests view.

Admission and affinity

New accounts default to 30 simultaneous model requests, configurable from 1 to 30. Upgrades preserve each existing account's saved limit; change it in Accounts → Scheduling settings if needed. This is a local SubLane safeguard, not a concurrency limit advertised by the subscription provider. The workspace admits at most 30 gateway operations, including waiting requests, and each member has a separate concurrency limit (10 by default). An account slot is reserved atomically with selection and retained until the response body closes, including HTTP/SSE, Chat Completions, compaction, and WebSocket turns. Cancellation closes/releases it exactly once. Reducing a limit allows existing requests to finish and stops additional admission until capacity is available.

An idle conversation or an open WebSocket connection consumes no model-request slot between turns. A completed response releases its slot when the body closes. After a client cancellation, a Codex stream may continue reading for up to three seconds to collect final usage before releasing its slot; the ten-minute request deadline and process shutdown still apply. If the client leaves a request running without sending cancellation, that request retains its slot until it ends or reaches the request deadline.

A bound conversation keeps its account. When capacity is full, a new request waits up to 10 seconds, with at most 8 waiting requests per pool. A released account slot or an increased concurrency limit wakes waiting requests. A full queue returns HTTP 429 account_queue_full; an expired wait returns HTTP 429 account_wait_timeout, both with Retry-After: 1. Waiting and these local rejections do not interrupt active streams or penalize upstream account health.

Waiting requests count toward the member concurrency and workspace admission limits and consume the member RPM allowance once. They do not reserve allowance or consume an account slot until admitted. Pool access, model policy, member enablement, key enablement/revocation/expiry, account health, subscription quota and allowance are checked again before dispatch. Client cancellation and service shutdown remove waiting requests promptly. Waiting releases the admission lock and holds no database transaction; completed usage is recorded before a released slot becomes available to another request.

Cooldown returns HTTP 429 account_cooling with a bounded Retry-After header and is not queued. No request automatically switches an existing conversation to another account. A sessionless or new conversation can use another eligible account immediately; a bound conversation waits only for its original account. Subscription quota exhaustion, member/workspace limits and allowance admission errors retain their normal guards.

Without Session_id or a nonempty prompt_cache_key, every call gets independent selection and a random upstream correlation ID. It does not read or create an affinity entry. New native selections rotate within each pool among eligible subscription accounts supporting the model. Explicit legacy prefixes limit selection to their provider. Existing explicit default-group conversations retain their historical digest.

The limit covers model requests; account verification, model discovery, and quota reads remain bounded by the existing global admission policy and do not consume model-request slots. These metadata operations do not clear model-request cooldowns.

Example account scheduling and subscription quota view with synthetic values:

Account card showing quota windows and concurrency state

Cooldown and recovery

Cooling down means the account recently returned a transient upstream failure, so SubLane temporarily stops sending new model requests to it. It is not an active-request count: an account can show 0 / 30 while cooling. The retry time is the end of the local safety window, not a guarantee that the provider has recovered. An administrator can clear the local cooldown, but repeated upstream failures can start it again.

ObservationBehavior
Upstream 429Immediately cool down; use integer/date Retry-After, bounded to 1–3600 seconds; default 60 seconds
5xx, upstream/network failure, timeout or interrupted/invalid streamThree consecutive logical failures trigger 30 seconds; subsequent failures back off exponentially up to five minutes
Valid completed or incomplete responseReset failure state when it belongs to a request started after the last failure and cooldown has expired
Client cancellation or downstream write failureRelease capacity without penalizing the upstream account
Authentication/entitlement/input errorsRetain existing credential/request error behavior; do not treat these as transient server failures

After cooldown expires, an account with no in-flight requests may admit one recovery probe. No background generation or provider polling occurs; a normal incoming request performs the probe. A successful probe restores normal concurrency; another failure cools the account again. Older in-flight success cannot clear newer failures, even within the same clock tick.

Clear cooldown also clears the failure streak and advances the runtime revision so older outcomes cannot reinstate it. It does not enable an account or repair authorization. Account settings/cooldowns are stored in SQLite. Native conversation bindings use the existing affinity table with an auto scope and survive restarts. In-flight counters and per-pool cursors are process-local, and restart starts them empty. Run one SubLane process per database. Persistence failures are reported through sanitized application logs; in-memory cooldown still blocks admission if its write fails.

Credential refresh and recovery

Codex credentials are refreshed on demand within two minutes of expiry. If that refresh temporarily fails, requests may continue with the existing access token only while it has more than 30 seconds remaining and has not been rejected by the upstream. Expiry is checked again after the refresh attempt. Claude and Antigravity retain their existing validity requirements.

Failed refresh attempts share an account-specific retry delay: 30 seconds, one minute, two minutes, four minutes, then at most five minutes. A longer OAuth Retry-After is respected, up to one hour. Once the delay expires, the next credential read or model request can retry; there is no additional background polling loop. Other accounts retain their own retry schedules.

If no usable token remains, model requests return 503 account_refresh_failed with Retry-After. Refresh failures do not accumulate model-call cooldown penalties. Successful refresh or reauthorization clears the refresh retry state. Retry delays are process-local, while an upstream rejection is stored with the encrypted credential so restarting cannot make that token eligible for fallback again.

The Codex token endpoint must return an explicit credential failure, such as invalid_grant or refresh_token_reused, to require reauthorization. An unknown HTTP 400/401, proxy response, or shared OAuth client configuration error alone does not mark an account as needing reauthorization. A refreshed token that is still rejected by the model endpoint continues to require reauthorization. Existing conversations keep their account binding throughout recovery.

Quota-aware selection

Codex, Claude and Antigravity model traffic starts shared background quota refreshes for eligible pool accounts, without waiting for provider IO. At most two quota reads run concurrently within the workspace admission slots; duplicate reads coalesce, and existing cooldown/backoff rules apply. There is no general idle polling loop. Managed share resources additionally run bounded retries after request completion to reconcile waiting usage, including after the last request; see reconciliation and pauses. Claude main windows can exclude an exhausted account. Antigravity checks only the exact requested native model, with SDK thinking suffixes normalized; other models remain eligible and account-wide runtime state stays available.

For Codex and Claude, selection reads the persisted normalized main subscription limit in the same short transaction as pool policy and affinity. An explicit denial, a reached-limit flag, or a window at 100% used excludes the account from new selection only while the observation is fresh (two minutes, shortened by a reset). Unknown, stale, future-dated or already-reset data does not mean zero quota. Additional named limits are not assumed to map to models and cannot exclude an entire account. Low-but-positive quota does not alter round-robin weights. A failed refresh retains the last successful observation; it can affect scheduling only until its original freshness deadline, never indefinitely.

A bound conversation whose account or requested Antigravity model is exhausted returns HTTP 403 quota_exhausted. The provider is checked again after the snapshot freshness/reset deadline; reaching that time is not a promise that allowance has recovered. It never moves to another account. Quota rejection does not increase transient failures, and Clear cooldown does not erase subscription quota. Accounts show a distinct quota-exhausted scheduling state.

Successful observations are saved before publication and survive restart. Quota rows carry the existing account lifecycle revision; disable/enable and reauthorization invalidate prior observations, and conditional writes reject older in-flight refreshes. The consolidated schema stores lifecycle revision metadata with quota rows and request diagnostics.

Recent request records

Records cover authenticated model attempts that reach gateway scheduling, including account-capacity/cooldown rejection. Authentication failures, HTTP body/content-type validation failures, global admission rejection, model discovery, quota queries, and WebSocket prewarm do not create model-call records.

Stored metadata is limited to user/key/pool/account IDs, provider, a bounded model identifier, protocol/operation, start time, duration, request correlation ID, nullable first output latency, outcome, sanitized error code, upstream HTTP status, and reported input/output/cached token counts. Prompts, response bodies, raw provider errors, headers, IP addresses, session identifiers and credentials are not recorded. Missing token counts remain null. Counts describe individual observed responses; they are not billing or the resource allowance ledger. Independent daily summaries and member request limits are described in team controls. A stream may have HTTP 200 yet finish with an error, so outcome and upstream status are distinct.

Personal records are filtered by the authenticated session user in SQL before pagination; a client-supplied user ID or scope cannot expand access. Subscription account IDs and names are cleared in the personal API response. Users can filter their own history by outcome, exact model, API key, request ID and time interval, including calls through keys that were subsequently revoked. Administrators can separately filter by member and subscription account. Caller options derive only from retained history inside the same ownership scope. Both views are paginated in batches of 50. At most 5,000 records are retained in a seven-day window. Cleanup runs atomically when recording, and on reading the history; idle data is reclaimed on the next such operation. Monotonic record IDs are not reused after retention cleanup. Deleted account names render as deleted, without deleting historical metadata.

Personal and administrator request tables show an input-token cache hit rate beside the cached-token count: cached input tokens divided by total input tokens, formatted as a percentage with at most one decimal place. This measures the share of input tokens served from the upstream cache, not the share of requests that hit a cache. A reported zero cache count with positive input shows 0%; missing counts, zero input or a cache count exceeding total input show an em dash. The rate is derived from existing request metadata and adds no stored column or upstream request.

Each new model attempt receives a server-generated req_… identifier, ignoring caller-supplied IDs. HTTP responses return it in X-Request-ID; WebSocket turns add request_id to created/completed/incomplete events and gateway errors. The upstream response ID is unchanged. Records from before the upgrade have no correlation ID or first-output measurement. The request table shows a compact ID beneath each start time with a copy action for the complete value. Clicking the start time opens its details with the full ID.

first_token_ms measures gateway processing start to the first nonempty normalized text, reasoning, reasoning-summary or tool-argument delta, including scheduling and upstream waiting. It is not time to the first lifecycle event or client-rendered token. Compaction, streams without an observed delta and older records keep null; the UI calls it First output. Total duration still covers body closure.

Example personal request diagnostics with synthetic records:

Request history showing request IDs, outcomes and latency

API and storage

  • GET /api/accounts/runtime: current scheduling state and model-request counts.
  • PATCH /api/accounts/{id}/limits: set max_concurrency (integer 1–30).
  • POST /api/accounts/{id}/resume: explicitly clear cooldown/failures.
  • GET /api/me/requests?cursor=0&outcome=: current user’s request metadata, for either enabled role.
  • GET /api/requests?cursor=0&account_id=&outcome=: administrator request metadata; outcome may be success, incomplete, error, canceled, or rejected.

Both history endpoints also accept model, request_id, key_id, from (inclusive Unix seconds) and until (exclusive Unix seconds); the administrator endpoint accepts user_id. Retention remains seven days regardless of the requested range. GET /api/me/requests/filters and administrator-only GET /api/requests/filters return bounded caller/key choices with no subscription identity or secret values.

Except for the explicitly personal history endpoints, these routes require an enabled administrator session; mutations retain origin/JSON protections. The consolidated schema creates workspace-scoped runtime and history tables. Back up the stopped database and credentials.key together before replacing a development installation.

Model-specific limits

Explicit structured Antigravity rate-limit responses are narrowed only when the error domain, reason and model match the selected request. Unknown 429 responses retain account-wide cooldown. Model limits persist across restart and belong to the current account lifecycle; reauthorization, disable/enable and manual cooldown clearing invalidate prior state.

A limited model returns 429 model_cooling; the original provider limit is reported as 429 model_rate_limited. Both preserve bounded retry hints. Other models remain usable, and existing conversations retain their original account binding. After cooldown, at most one normal request for that model attempts recovery at a time; no extra inference probe is sent. Older or unrelated successes cannot erase a newer limit. A failed persistence attempt retains the restriction in the current process and is logged as a persistence failure.

Account cards display a limited-model count while overall account scheduling remains available. Scheduling settings → Clear cooldown clears account and model cooldowns, without changing enablement, authorization or subscription quota. Upstream HTTP 403 is reported as upstream_forbidden, rather than a request to reauthorize.