Add ClinePass across server and VS Code quota paths. Reject malformed windows, preserve valid sibling limits, align credential fallback, and report timeout failures accurately. Validated 30 server quota tests, 129 VS Code quota tests and VS Code type-check. Oxlint findings are confined to pre-existing VS Code code.
17 KiB
Quota Module Documentation
Purpose
This module fetches quota and usage signals for supported providers in the web server runtime.
Node server entrypoints apply ../network-defaults.js before serving requests,
including the daemon launched directly through server/index.js. Connection
attempts get 5 seconds without changing address-family selection. The VS Code
extension applies the same policy in its own process at activation.
Entrypoints and structure
packages/web/server/lib/quota/index.js: public entrypoint imported bypackages/web/server/index.js.packages/web/server/lib/quota/routes.js: Express route registration for quota endpoints.packages/web/server/lib/quota/providers/index.js: provider registry, configured-provider list, and provider dispatcher.packages/web/server/lib/quota/providers/google/: Google-specific auth, API, and transform modules.packages/web/server/lib/quota/providers/claude/: Claude credential discovery, usage transforms, and rate-limit handling.packages/web/server/lib/quota/utils/: shared auth, transform, and formatting helpers.
Supported provider IDs (dispatcher)
These provider IDs are currently dispatchable via fetchQuotaForProvider(providerId) in packages/web/server/lib/quota/providers/index.js.
| Provider ID | Display name | Module | Auth aliases/keys |
|---|---|---|---|
claude |
Claude | providers/claude/ |
Claude Code Keychain entry, Claude Code credentials file, OpenCode auth.json (anthropic, claude), CLAUDE_CODE_OAUTH_TOKEN |
cline-pass |
ClinePass | providers/cline-pass.js |
cline-pass (API key under key or token) |
codex |
Codex | providers/codex.js |
openai, codex, chatgpt |
command-code |
Command Code | providers/command-code.js |
command-code OAuth/API credential in OpenCode auth.json, or COMMAND_CODE_API_KEY |
cursor |
Cursor | providers/cursor.js |
Environment/token files, OpenChamber-managed credentials, or explicit one-time Cursor import |
crof |
CrofAI | providers/crof.js |
crof (API key under key or token) |
deepseek |
DeepSeek | providers/deepseek.js |
deepseek (API key under key or token) |
exe-dev |
exe.dev | providers/exe-dev.js |
Usage API token stored under ~/.config/openchamber/quota/ |
google |
providers/google/index.js |
google, google.oauth, Antigravity accounts file |
|
hyper |
Charm Hyper | providers/hyper.js |
hyper (API key under key or token) |
github-copilot |
GitHub Copilot | providers/copilot.js |
github-copilot, copilot |
github-copilot-addon |
GitHub Copilot Add-on | providers/copilot.js |
github-copilot, copilot |
kimi-for-coding |
Kimi for Coding | providers/kimi.js |
kimi-for-coding, kimi |
nano-gpt |
NanoGPT | providers/nanogpt.js |
nano-gpt, nanogpt, nano_gpt |
openrouter |
OpenRouter | providers/openrouter.js |
openrouter |
zai-coding-plan |
z.ai | providers/zai.js |
zai-coding-plan, zai, z.ai |
zhipuai-coding-plan |
Zhipu AI Coding Plan | providers/zhipuai-coding-plan.js |
zhipuai-coding-plan, zhipuai, zhipu |
minimax-coding-plan |
MiniMax Coding Plan (minimax.io) | providers/minimax-coding-plan.js / providers/minimax-shared.js |
minimax-coding-plan |
minimax-cn-coding-plan |
MiniMax Coding Plan (minimaxi.com) | providers/minimax-cn-coding-plan.js / providers/minimax-shared.js |
minimax-cn-coding-plan |
ollama-cloud |
Ollama Cloud | providers/ollama-cloud.js |
Manual cookie pasted into Settings (aid=...; __Secure-session=... from ollama.com), stored under ~/.config/openchamber/quota/ |
wafer |
Wafer.ai | providers/wafer.js |
wafer, wafer-ai, wafer_ai, wafer.ai |
opencode-go |
OpenCode Go | providers/opencode-go.js |
opencode-go API key from OpenCode auth.json |
neuralwatt |
NeuralWatt | providers/neuralwatt.js |
neuralwatt (API key under key or token) |
xai |
xAI | providers/xai.js |
xai OAuth entry in OpenCode auth.json |
Internal-only provider module
providers/openai.jsexists for logic parity/reuse but is intentionally not registered for dispatcher ID routing.
Response contract
All providers should return results via shared helpers to preserve API shape:
- Required fields:
providerId,providerName,ok,configured,usage,fetchedAt - Optional field:
error - Unsupported provider requests should return
ok: false,configured: false,error: Unsupported provider
Provider modules must export providerId, providerName, aliases, isConfigured(auth?), and fetchQuota().
fetchQuota() should return a quota result with usage.windows keyed by window name (for example 5h, 7d, daily) and optional provider-specific usage.models data.
exe.dev, Ollama Cloud, and Cursor credentials are explicitly managed through Settings. exe.dev usage uses a separately generated HTTPS API token restricted to billing credits usage and aggregates every exe-* model provider into one monthly credit window. Generate the token with ssh exe.dev "ssh-key generate-api-key --label=openchamber --exp=30d --cmds='billing credits usage'". OpenCode Go usage uses GET https://opencode.ai/zen/go/v1/usage with the opencode-go API key from OpenCode auth.json as a bearer token and the stable x-opencode-session: openchamber-usage workload id. The server validates managed credentials before atomic 0600 writes and never returns secrets through its API. OpenChamber never scans browser cookie stores or automatically reads Cursor storage; Cursor import is an explicit one-time user action and never modifies Cursor's database.
Command Code usage resolves account scope through GET /alpha/whoami, then reads server-backed credit balances and five-hour/weekly limits from GET /alpha/billing/credits?orgId=.... Personal accounts return org: null and use /alpha/billing/credits without an orgId; organization accounts include their organization id. Web/Electron and VS Code read the standard command-code OpenCode auth entry (including OAuth access) or COMMAND_CODE_API_KEY; credentials remain in the owning runtime and are never returned to shared UI.
On the first OpenCode Go usage refresh after upgrading, OpenChamber deletes the obsolete quota/opencode-go.json credential file without reading its cookie value.
Claude credential and limit semantics
Claude quota reports the subscription limits Claude Code itself is bound by, read from GET https://api.anthropic.com/api/oauth/usage.
- Credential sources, in priority order: the macOS Keychain entry
Claude Code-credentials, then${CLAUDE_CONFIG_DIR:-~/.claude}/.credentials.json(the Linux/WSL location), then the OpenCodeauth.jsonentry, thenCLAUDE_CODE_OAUTH_TOKEN. The Keychain wins on macOS because the credentials file there is a leftover Claude Code no longer updates. - All sources are read-only. OpenChamber never writes to Claude Code's credential store and never refreshes the OAuth token, because Anthropic does not support two live refresh tokens for one
client_id— refreshing here would sign the user out of Claude Code. Credentials are read fresh per request so a Claude Code refresh is picked up immediately; an expired token yields an explicit "open Claude Code to sign in again" error rather than a bare 401. - Limits come from the
limitsarray, keyed bykind:sessionmaps to the5hwindow,weekly_allto7d, andweekly_scopedto a per-model7dwindow named byscope.model.display_name. The legacyfive_hour/seven_dayfields are only a fallback;seven_day_sonnet/seven_day_opusare no longer populated by Anthropic. Unrecognized limit kinds and Anthropic's rotating internal code names (nimbus_quill,tangelo, ...) are ignored rather than guessed at. - Extra usage is reported as the
extra_usagewindow fromspend, only whilespend.enabledis true, with a moneyvalueLabel. - Rate limiting: Anthropic returns 429 aggressively. The last successful usage payload is cached in memory and reserved during a cooldown (
Retry-After, else five minutes, capped at one hour). The cache is keyed by a hash of the access and refresh tokens, so switching accounts drops it instead of showing the previous account's numbers. - Runtime parity: Web/Electron and VS Code preserve the last successful Claude values during the same bounded 429 cooldown. Quota dispatchers also coalesce concurrent refreshes for the same provider in each runtime, while requests for different providers remain parallel.
Add a new provider (quick steps)
- Choose module shape based on complexity:
- Simple providers: create
packages/web/server/lib/quota/providers/<provider>.js. - Complex providers (multi-source auth, multiple API calls, non-trivial transforms): create
packages/web/server/lib/quota/providers/<provider>/with split modules like Google (index.js,auth.js,api.js,transforms.js).
- Simple providers: create
- Export
providerId,providerName,aliases,isConfigured, andfetchQuota. - Use shared helpers from
packages/web/server/lib/quota/utils/index.js(buildResult,toUsageWindow, auth/conversion helpers) to keep payload shape consistent. - Register the provider in
packages/web/server/lib/quota/providers/index.js. - If needed for direct use, export a named fetcher from
packages/web/server/lib/quota/providers/index.jsandpackages/web/server/lib/quota/index.js. - Update this file with the new provider ID, module path, and alias/auth details.
- Validate with
bun run type-check,bun run lint, andbun run build.
MiniMax M3 / Token Plan migration
In 2025/2026 MiniMax rebranded "Coding Plan" to "Token Plan" alongside the M3 model release. The API underwent breaking changes:
- Endpoint fallback: The provider tries
/v1/token_plan/remains(M3) first, falling back to legacy/v1/api/openplatform/coding_plan/remains. - Field semantics: On the
token_plan/remainsendpoint,current_interval_usage_countreturns remaining quota (not consumed). The provider computesused = total - remainingfor this endpoint. The legacycoding_plan/remainsendpoint retains the old semantics (usage_count = consumed). - Percentage-based plans: Legacy Coding Plan accounts return
current_interval_total_count: 0but includecurrent_interval_remaining_percent. The provider prefers this field when count fields are absent. - model_remains array: Now contains entries for multiple model categories (chat, speech, video, image). The provider selects the chat-model entry by matching
MiniMax-M*, thengeneral/chat/textby name, then any entry with a remaining percent. - Window status: The
current_interval_statusandcurrent_weekly_statusfields indicate whether a window is active. Status3means the window is not applicable for the current plan tier (e.g. legacy plans without weekly limits). The provider omits inactive windows.
ClinePass quota semantics
ClinePass reads data.limits from its usage-limits endpoint. Web/Electron and
VS Code accept only known window types with finite numeric or non-empty numeric
string percentages. Invalid windows are skipped independently; no usable windows
is a failed refresh, not zero usage. Both implementations choose a non-empty
key, then token, and expose auth/fetch dependencies for focused tests.
Saved UI provider-visibility lists remain authoritative; installations without a
saved list include ClinePass through the provider registry.
Charm Hyper balance semantics
GET https://hyper.charm.land/v1/credits returns a team's current Hypercredit balance, not a percentage or reset timestamp. The Hyper FAQ defines one Hypercredit as $0.05. Both runtimes expose credits_balance in dollars and credits as a numeric label under the UI's localized window title. Keep English unit text out of that numeric label.
Web and VS Code accept finite numeric balances and non-empty numeric strings. Missing, blank, or malformed balances remain explicit failures; zero is valid. Credential lookup uses a non-empty string key, then token, so malformed or blank keys cannot mark the provider configured or hide a valid fallback token. Hyper fetchers accept readAuth and fetchImpl dependencies for tests without replacing filesystem or auth modules.
Kimi for Coding field semantics
GET https://api.kimi.com/coding/v1/usages is inconsistent about which field carries consumption:
- The weekly
usageblock returnsused(consumed) with noremainingfield. - Each
limits[].detailrate-limit block returnsremaining(available) with nousedfield.
The provider computes usedPercent from whichever of used/remaining is present (used takes precedence when both exist) rather than assuming one field name. Both packages/web/server/lib/quota/providers/kimi.js and packages/vscode/src/quotaProviders.ts (fetchKimiQuota) must stay in sync — the VS Code extension duplicates this parsing logic rather than importing it.
Ollama Cloud settings-page shapes
Ollama Cloud authentication uses two cookies (aid and __Secure-session) pasted together as one single-line Cookie header value. Ollama serves two different /settings page shapes and parseOllamaSettingsHtml supports both: session/weekly/premium-interaction windows (percent-based plans) and a monthly window derived from "Monthly usage: $X of $Y used" (cost-based plans) with a symmetric $X / $Y money valueLabel that reads correctly in both used/remaining display modes. The "Extra usage" credits block is surfaced as a balance-only credits_balance window with a plain money valueLabel when present, matching the Codex/DeepSeek credits treatment ("Credits Balance" in the UI); a $0 balance is omitted instead of showing an empty credits row. Keep packages/web/server/lib/quota/providers/ollama-cloud.js and packages/vscode/src/quotaProviders.ts (parseOllamaSettingsHtml) in sync — the VS Code extension duplicates this parsing logic rather than importing it.
GitHub Copilot quota semantics
GitHub Copilot usage exposes only the premium_interactions snapshot as the
premium_interactions window. Shared UI labels that window AI Credits and treats it as
the provider's primary usage marker. Legacy chat-request quota and unlimited
completion quota are intentionally omitted. Keep
packages/web/server/lib/quota/providers/copilot.js and
packages/vscode/src/quotaProviders.ts in sync.
The /copilot_internal/user endpoint is undocumented; its quota semantics mirror
what microsoft/vscode-copilot-chat consumes (CopilotUserQuotaInfo). Each
snapshot carries entitlement, remaining, unlimited, and
percent_remaining. Providers must honor these rules:
unlimited: truerenders a percent-less window with an "Unlimited" value label.- Percent math requires a positive
entitlement; entitlements of0,-1, or null are unusable. - When entitlement/remaining are unusable, fall back to
100 - percent_remaining. - Snapshots other than
premium_interactions(legacy annual plans) yield zero windows.
OpenRouter key semantics
OpenRouter quota reads GET https://openrouter.ai/api/v1/key, which is documented as callable with any valid API key. GET /api/v1/credits is documented as "Management key required" and is not used. Calling /credits with a normal inference key has been observed to return HTTP 200 with {total_credits:0, total_usage:0} rather than an error; this behavior is not documented and is why the old implementation silently rendered "$0.00 left · $0.00 spent". A /credits fallback for unlimited keys would render the same zeros, so unlimited keys report usage_monthly instead.
The documented limit, limit_remaining, and limit_reset fields are present and null on unlimited keys; null means unlimited, never missing data. For a limited key, window usage is limit - limit_remaining, not usage: usage is all-time and measures a different axis from the current reset window. Pairing usage with the current limit produces a wrong number. limit_remaining is server-computed and already honors include_byok_in_limit, so byok_* fields are ignored.
Unlimited keys report usage_monthly in a monthly window with no percent. limit_reset is a period string (daily, weekly, monthly, or null), not a timestamp; resetAt is derived from the documented midnight-UTC boundaries, with weeks starting Monday. A set limit with a null limit_reset is a lifetime cap and maps to the credits window with no reset.
Keep packages/web/server/lib/quota/providers/openrouter.js and packages/vscode/src/quotaProviders.ts in sync, as with the Kimi and Copilot providers; the VS Code extension duplicates this parsing logic rather than importing the web provider.
Notes for contributors
- Keep provider IDs stable; clients use them directly.
- Avoid adding alias-based dispatch in
fetchQuotaForProvider; dispatch currently expects exact provider IDs. - Keep Google behavior changes isolated and review
providers/google/*together. - Z.ai Coding Plan exposes separate 5-hour and weekly token/credit limit entries plus a monthly
TIME_LIMITfor MCP tools. The API renamed the limit type fromTOKENS_LIMITtoCREDIT_LIMIT(sameunit/numberwindow semantics);CREDIT_LIMITentries additionally carryusage(total),currentValue(consumed), andremaining, surfaced as a creditvalueLabel, and the payload'sdata.levelbecomesplanLabel. Web and VS Code must preserve these windows and stay in sync.