SubLaneSubLane

Native client protocols

SubLane guide: native client protocols.

These endpoints were introduced after v0.1.0-rc.1; that release does not include them. Use a source build until a newer release includes this change.

Codex, Claude and Antigravity subscription accounts are enabled. Claude Messages and Gemini endpoints describe client request formats independently of the upstream provider. SubLane accepts Claude Messages and Gemini generation requests alongside its OpenAI-compatible endpoints. Incoming protocol and subscription provider are independent: requests still select an eligible account reporting the requested model within the API key's pool. The public CLIProxyAPI executors handle provider translation; native Claude requests do not pass through an intermediate Responses representation.

Endpoints and authentication

ProtocolEndpointCredentials
Claude MessagesPOST /v1/messagesx-api-key or Authorization: Bearer
Gemini JSONPOST /v1beta/models/{model}:generateContentx-goog-api-key, Authorization: Bearer, or key query parameter
Gemini SSEPOST /v1beta/models/{model}:streamGenerateContent?alt=sseSame as Gemini JSON

Use a personal SubLane API key. If multiple credential locations are supplied, they must agree; an invalid Authorization header cannot fall back to another key. Browser session cookies are not gateway credentials. Prefer headers to query-string keys because other proxies may log URLs. SubLane does not forward client keys, cookies or query strings to providers.

Use GET /v1/models with Bearer authentication to discover permitted native model IDs. For Gemini requests, the model in the URL is authoritative; a body model or stream field cannot override the route. streamGenerateContent returns SSE; alt may be omitted or set to sse.

Claude example

Set SUBLANE_URL to the instance origin and SUBLANE_API_KEY to your personal key. Replace YOUR_MODEL_ID with a model from the catalog:

Claude example · 1
curl "$SUBLANE_URL/v1/messages" \
  -H "x-api-key: $SUBLANE_API_KEY" \
  -H 'anthropic-version: 2023-06-01' \
  -H 'content-type: application/json' \
  --data '{"model":"YOUR_MODEL_ID","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'

Add "stream": true for named Messages SSE events. Native request content includes system blocks, tools, tool results, cache controls and thinking configuration. Claude-format responses preserve content blocks and thinking signatures, subject to the selected executor's provider-specific handling. Errors use the Anthropic envelope and named SSE error events. request-id and X-Request-ID identify the corresponding request record.

Gemini example

Gemini example · 2
curl "$SUBLANE_URL/v1beta/models/YOUR_MODEL_ID:generateContent" \
  -H "x-goog-api-key: $SUBLANE_API_KEY" \
  -H 'content-type: application/json' \
  --data '{"contents":[{"role":"user","parts":[{"text":"Hello"}]}]}'

For streaming, use :streamGenerateContent?alt=sse and curl --no-buffer. Responses contain native candidates, parts and usageMetadata; Antigravity's internal response envelope is removed by the SDK. Tool calls and thought signatures stay in their Gemini representation. Errors use Google-style error.code, error.status and error.message. SSE chunks have data: framing without OpenAI's [DONE] marker.

Shared behavior

  • Enabled members and keys, key expiry, pool grants, model policies, account availability, concurrency and member request limits are rechecked through the existing gateway services.
  • For a new conversation or sessionless request, all formats select an eligible Codex account reporting the requested model. Explicit prefixes for disabled providers cannot select an account. Existing conversations bound to a disabled provider fail without switching accounts.
  • Conversation affinity uses Session_id; Claude also uses metadata.user_id when no explicit session is supplied. Identifiers are scoped to the member and pool. Requests without an affinity identifier remain sessionless.
  • Cancellation closes upstream work and releases account/member leases. Missing terminal events and malformed or oversized events are failures, not successful empty responses. Streaming errors are sanitized after headers have been sent.
  • Request history labels operations as messages or gemini, with request IDs, first-output latency and available token usage. Anthropic input totals include cache reads and creation; Gemini input already includes cached tokens, and its output total includes thinking tokens. Trailing Gemini usage chunks are consumed after finishReason.
  • Request bodies default to 128 MiB (configurable with SUBLANE_MAX_REQUEST_BODY_MB); responses and stream events retain separate 8 MiB limits. Ten-minute operation deadlines apply. Streaming generation is not buffered into a complete response for the client.

The consolidated initialization schema includes native protocol operations in request history. No additional persistent service is required.

Scope and verification

The following synthetic compatibility matrix documents the retained implementation. Only the Codex account column is active in the current application runtime:

Client protocolCodex accountClaude accountAntigravity account
OpenAI ResponsesTestedTestedTested
OpenAI Chat CompletionsTestedTestedTested
Claude MessagesTestedTestedTested
Gemini generationTestedTestedTested

TestClientProtocolMatrix exercises all 24 provider/protocol/stream combinations against mocked upstreams; this does not make the paused subscription providers available at runtime. Separate tests cover native signatures, cache usage, permissions, retries, cancellation and malformed streams. These checks establish the protocol behaviors they exercise; they do not certify every tool, thinking mode or provider-specific field. Full real-account client acceptance remains pending for each combination.

Coverage uses synthetic accounts and mocked providers, including native JSON/SSE, tools/thinking, cross-provider translation, authorization, usage, interruption and cancellation. It does not establish complete real-account Claude Code, Gemini CLI or desktop compatibility. Advanced native features depend on the selected provider and SDK conversion path.

In API keys → Setup guide, choose your client or API format. The guide provides Codex configuration, Cline/OpenAI-compatible setup and Chat Completions examples, or native Claude/Gemini request examples, with the correct base URL and authentication header. Select a model from the key's pool. The guide never fetches a key's secret; examples read SUBLANE_API_KEY from the client environment. CC Switch quick import supports Codex, OpenCode, Claude Code, OpenClaw, Hermes, Gemini CLI, and Grok Build.

Gemini CLI 0.37.2 has additionally completed a text-streaming request and a read-only tool round trip against an isolated mock endpoint using a synthetic key and model. This checks the CLI request path and response consumption; it does not validate real-subscription inference, every CLI feature, or the other clients' full workflows.

Token-count endpoints, Message Batches, Files, Gemini Live and Antigravity desktop connection protocols are not implemented by this change. Unsupported actions are rejected rather than treated as generation requests. Native model-discovery resources are not exposed; use the pool-scoped /v1/models catalog.

Protocol references: Claude streaming, Claude errors, and Gemini generation. Routing and compatibility boundaries were compared with sub2api and New API; SubLane retains its own permission, lifecycle and storage implementation.