Import claude-code-switcher from the Windows build
Source of truth so far has been c:\Users\smoido\projects\cli on the Windows box, which has no git history of its own. This is that tree copied verbatim over SSH, minus dist/ - the PowerShell build, the POSIX port under linux/, and the presets both share. Recorded as its own commit so that everything after it is a reviewable diff rather than an undifferentiated first drop.
This commit is contained in:
@@ -0,0 +1,326 @@
|
||||
# claude-mode
|
||||
|
||||
Switch Claude Code system-wide between **Anthropic**, **OpenRouter**, **Z.AI**, and a
|
||||
local **LM Studio** server — with named per-tier model presets.
|
||||
Windows / PowerShell 5.1, no external dependencies.
|
||||
|
||||
```
|
||||
claude-mode # interactive menu
|
||||
claude-mode anthropic # subscription login
|
||||
claude-mode openrouter # remote gateway (preset: default)
|
||||
claude-mode zai # Z.AI GLM coding plan (preset: zai)
|
||||
claude-mode lmstudio # local server (preset: lmstudio)
|
||||
```
|
||||
|
||||
Works for the CLI, the VS Code extension, and the desktop app from a single
|
||||
switch. Restart Claude Code afterwards — nothing else.
|
||||
|
||||
---
|
||||
|
||||
## The interactive menu
|
||||
|
||||
Run `claude-mode` with no arguments. The mode you're already in is omitted —
|
||||
there's nothing to switch to:
|
||||
|
||||
```
|
||||
claude-mode (currently: openrouter / default)
|
||||
switch mode:
|
||||
1) Anthropic - your subscription login, no gateway
|
||||
2) Z.AI - GLM coding plan
|
||||
3) LM Studio - local server, offline, free
|
||||
|
||||
4) show full status
|
||||
5) edit presets
|
||||
6) run doctor
|
||||
0) quit
|
||||
```
|
||||
|
||||
Pick a provider and it lists that provider's presets with the default marked;
|
||||
**Enter** accepts it. Option 5 walks preset → tier → new model ID, and if you
|
||||
edit the preset that's currently live it re-applies immediately.
|
||||
|
||||
When stdin is redirected (scripts, CI) the menu is skipped and `status` prints
|
||||
instead, so `claude-mode` is still safe in a pipeline.
|
||||
|
||||
## Why settings.json and not a profile export
|
||||
|
||||
Three ways to make this persist. I picked the second.
|
||||
|
||||
**1. Export the variables from `$PROFILE`.**
|
||||
The obvious move, and the wrong one here. It only covers processes launched from
|
||||
a PowerShell session that loaded the profile — which is exactly *not* how you use
|
||||
Claude Code. The VS Code extension is spawned by VS Code, not by your shell, so
|
||||
it would never see the exports. Same for the desktop app, `cmd.exe`, and any
|
||||
terminal opened before the switch. Worse, the failure is silent: you switch to
|
||||
`anthropic`, a shell opened five minutes ago still has `ANTHROPIC_BASE_URL` set,
|
||||
and that session quietly keeps billing OpenRouter.
|
||||
|
||||
**2. Rewrite the `env` block in `~/.claude/settings.json`.** ← chosen
|
||||
Claude Code reads this file on every startup, from every launch context. One
|
||||
write, and the next `claude` — CLI, extension, desktop — picks it up. A switch is
|
||||
atomic: one file, one source of truth. `claude-mode anthropic` *deletes* the
|
||||
managed keys rather than blanking them, so nothing can linger and break native
|
||||
auth. The cost is that config is global rather than per-terminal.
|
||||
|
||||
**3. Persistent User-scope environment variables (`setx`).**
|
||||
Also global and reboot-proof, but strictly worse: new processes only, values sit
|
||||
in the registry in plaintext, and a stale entry silently outranks whatever
|
||||
`claude-mode` writes. This tool treats them as a fault condition — `status` and
|
||||
`doctor` flag them and offer removal, backing the old value up first.
|
||||
|
||||
The `claude` wrapper in the profile is a **safety net, not the mechanism**. It
|
||||
strips inherited process-level copies of all thirteen managed variables before
|
||||
launching `claude.exe`. Everything still works without it — including in VS Code,
|
||||
which never loads the profile.
|
||||
|
||||
## Why API keys are not in settings.json
|
||||
|
||||
`settings.json` is a config file you'll hand-edit, diff, and possibly paste into
|
||||
a bug report. A `sk-or-` or Z.AI token does not belong there.
|
||||
|
||||
Keys are stored **DPAPI-encrypted** in `~/.claude-mode/vault/*.cred` — encrypted
|
||||
against your Windows account on this machine, so copying the file elsewhere or
|
||||
reading it as another user yields nothing — with the file ACL restricted to you.
|
||||
Claude Code receives the key at runtime through `apiKeyHelper`, which decrypts
|
||||
and prints it. `settings.json` holds only the base URL and model IDs.
|
||||
|
||||
In `anthropic` mode the helper is removed from settings.json *and* returns
|
||||
nothing when state says `anthropic` — belt and braces. LM Studio's `lmstudio`
|
||||
token is a placeholder, not a secret, so it's written inline and the helper stays
|
||||
out of it.
|
||||
|
||||
## Commands
|
||||
|
||||
```
|
||||
claude-mode interactive menu
|
||||
claude-mode status active mode, preset, model map
|
||||
|
||||
claude-mode anthropic native login (deletes all managed keys)
|
||||
claude-mode openrouter [preset] default preset: default
|
||||
claude-mode zai [preset] default preset: zai (alias: z.ai, z-ai)
|
||||
claude-mode lmstudio [preset] default preset: lmstudio
|
||||
|
||||
claude-mode presets list presets (* = active)
|
||||
claude-mode preset show <name>
|
||||
claude-mode preset new <name> [from] copy an existing preset
|
||||
claude-mode preset set <name> <tier> <model-id>
|
||||
claude-mode preset all <name> <model-id> point every tier at one model
|
||||
claude-mode preset rm <name>
|
||||
|
||||
claude-mode set-key [ref] store a key (hidden prompt, DPAPI)
|
||||
claude-mode models [filter] models available from the active provider
|
||||
claude-mode doctor verify auth, endpoint, model ids, stray env vars
|
||||
```
|
||||
|
||||
Omitting the preset uses a **fixed** per-provider default, not "most recently
|
||||
used" — so `claude-mode openrouter` always means `default`.
|
||||
|
||||
## Presets shipped
|
||||
|
||||
| preset | provider | opus | sonnet | haiku | fable |
|
||||
|---|---|---|---|---|---|
|
||||
| `default` | openrouter | `deepseek/deepseek-v4-flash` | `qwen/qwen3.7-flash` | `openrouter/free` | `z-ai/glm-5.2` |
|
||||
| `cheap` | openrouter | `deepseek/deepseek-v4-pro` | `deepseek/deepseek-v4-flash` | `qwen/qwen3.7-flash` | `openai/gpt-5.6-luna-pro` |
|
||||
| `claude-via-or` | openrouter | `anthropic/claude-opus-5` | `anthropic/claude-sonnet-5` | `anthropic/claude-haiku-4.5` | `anthropic/claude-fable-5` |
|
||||
| `zai` | zai | `glm-5.2` | `glm-5.2` | `glm-4.7` | `glm-5.2` |
|
||||
| `lmstudio` | lmstudio | `kwaipilot_kat-coder-v2.5-dev` (all tiers) | | | |
|
||||
| `lmstudio-qwen` | lmstudio | `qwen3.6-35b-a3b-uncensored-heretic-native-mtp-preserved` (all tiers) | | | |
|
||||
|
||||
Presets are plain JSON in `~/.claude-mode/presets/`. A preset declares its
|
||||
`provider`; `claude-mode lmstudio default` is rejected rather than silently
|
||||
pointing a local URL at remote model IDs.
|
||||
|
||||
## Context windows and early auto-compaction
|
||||
|
||||
**Symptom:** switch to a gateway and the session starts auto-compacting almost
|
||||
immediately, even though every model involved has a huge context window.
|
||||
|
||||
**Cause:** behind a custom `ANTHROPIC_BASE_URL`, Claude Code has no way to
|
||||
resolve a third-party model ID like `deepseek/deepseek-v4-flash` to a context
|
||||
length. It falls back to a conservative default and starts compacting against
|
||||
*that*, not against the model's real 1M window. Z.AI's own docs work around this
|
||||
by setting `CLAUDE_CODE_AUTO_COMPACT_WINDOW=1000000` — they hit the same thing.
|
||||
|
||||
**Fix:** every preset carries a `contextTokens` field, which writes both knobs
|
||||
(confirmed present in CLI 2.1.221):
|
||||
|
||||
```
|
||||
CLAUDE_CODE_MAX_CONTEXT_TOKENS = <contextTokens>
|
||||
CLAUDE_CODE_AUTO_COMPACT_WINDOW = <contextTokens>
|
||||
```
|
||||
|
||||
| preset | contextTokens |
|
||||
|---|---|
|
||||
| `default`, `cheap`, `claude-via-or`, `zai` | 1,000,000 |
|
||||
| `lmstudio`, `lmstudio-qwen` | 262,144 |
|
||||
|
||||
`doctor` cross-checks the declared window against each tier's *actual* model
|
||||
window and names any tier that falls short — `default` maps haiku to
|
||||
`openrouter/free` (200k), which it flags as harmless since haiku only runs short
|
||||
background tasks. Switching without `contextTokens` prints a warning.
|
||||
|
||||
Adjust per preset:
|
||||
|
||||
```powershell
|
||||
# edit ~/.claude-mode/presets/<name>.json -> "contextTokens": 262144
|
||||
claude-mode doctor # re-checks declared vs actual
|
||||
```
|
||||
|
||||
## Z.AI mode
|
||||
|
||||
Replaces `npx @z_ai/coding-helper` — and does something it doesn't: **maps a
|
||||
distinct model to each Anthropic tier** instead of forcing one model everywhere.
|
||||
|
||||
Per [Z.AI's Claude Code docs](https://docs.z.ai/devpack/tool/claude):
|
||||
|
||||
| setting | value |
|
||||
|---|---|
|
||||
| `ANTHROPIC_BASE_URL` | `https://api.z.ai/api/anthropic` |
|
||||
| auth | your Z.AI API key — kept in the DPAPI vault, delivered via `apiKeyHelper` |
|
||||
| `API_TIMEOUT_MS` | `3000000` |
|
||||
| `CLAUDE_CODE_AUTO_COMPACT_WINDOW` | `1000000` |
|
||||
| `CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` | `1` |
|
||||
|
||||
All three extra variables are confirmed present in CLI 2.1.221. Setup:
|
||||
|
||||
```powershell
|
||||
claude-mode set-key zai # paste your key from https://z.ai/manage-apikey/apikey-list
|
||||
claude-mode zai
|
||||
claude-mode doctor # sends a 1-token request to prove the key works
|
||||
```
|
||||
|
||||
## LM Studio mode
|
||||
|
||||
Per [LM Studio's docs](https://lmstudio.ai/docs/integrations/claude-code): base
|
||||
URL `http://127.0.0.1:1234` (**base only** — Claude Code appends `/v1/messages`),
|
||||
token `lmstudio`, plus `CLAUDE_CODE_ATTRIBUTION_HEADER=0`. Gateway discovery
|
||||
stays off; the Anthropic surface is `/v1/messages` only.
|
||||
|
||||
### Model IDs are not what the UI shows
|
||||
|
||||
LM Studio's `/v1/models` lists only **loaded** instances under their display
|
||||
alias. `kat-coder-v2.5` is such an alias — once the model idle-unloads it
|
||||
vanishes, and a request using that name returns `400 No models loaded`. The
|
||||
JIT-loadable ID is the model key, `kwaipilot_kat-coder-v2.5-dev`. `claude-mode`
|
||||
reads `/api/v0/models` instead, which lists every installed model with its load
|
||||
state, so `models` and `doctor` show IDs that actually work.
|
||||
|
||||
### The `[Server Error] ... Unable to generate parser for this template` spam
|
||||
|
||||
Cause: some GGUF chat templates hard-assert message ordering —
|
||||
|
||||
```jinja
|
||||
{%- if message.role == "system" %}
|
||||
{%- if not loop.first %}
|
||||
{{- raise_exception('System message must be at the beginning.') }}
|
||||
```
|
||||
|
||||
Runtimes that auto-generate a tool-call parser probe the template with synthetic
|
||||
message sequences; those probes trip the assertion and the request dies. It's a
|
||||
model-template bug, not a Claude Code or claude-mode bug — it's been reported
|
||||
against several models
|
||||
([LM Studio #1999](https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1999),
|
||||
[llama.cpp #20733](https://github.com/ggml-org/llama.cpp/issues/20733)).
|
||||
|
||||
Scanning your installed models' templates:
|
||||
|
||||
| model | template |
|
||||
|---|---|
|
||||
| `qwen3.6-35b-a3b-uncensored-heretic-native-mtp-preserved` | clean |
|
||||
| `qwen3.6-35b-a3b` | clean |
|
||||
| `qwen2.5-coder-7b-instruct`, `google/gemma-4-12b-qat` | clean |
|
||||
| **`kwaipilot_kat-coder-v2.5-dev`** | **asserts** |
|
||||
| **`qwen/qwen3.5-9b`**, **`prism-ml/bonsai-27b`** | **assert** |
|
||||
|
||||
`doctor` now reports this per model, and `models` flags affected entries with
|
||||
`TEMPLATE RISK`. **The fix is to use a model without the flag** — which is
|
||||
exactly the `lmstudio-qwen` preset, verified end-to-end with a cold JIT load,
|
||||
streaming, and tool calls.
|
||||
|
||||
Honest caveat: KAT-Coder's template *does* contain the assertion, but I could not
|
||||
reproduce the failure against it here — cold JIT, streaming, tools, system
|
||||
blocks, and multi-turn `tool_result` all succeeded. Whether it trips seems to
|
||||
depend on which parser strategy the runtime picks. If it spams, switch:
|
||||
|
||||
```powershell
|
||||
claude-mode lmstudio lmstudio-qwen
|
||||
```
|
||||
|
||||
Other notes: use a model with **>25k context** (`doctor` warns below that), and
|
||||
the model must be **installed** — JIT loading handles "not loaded" fine.
|
||||
|
||||
If you enable authentication in LM Studio, move that preset to a vault key:
|
||||
|
||||
```powershell
|
||||
claude-mode set-key lmstudio
|
||||
# then in ~/.claude-mode/presets/lmstudio.json:
|
||||
# "auth": { "mode": "vault", "keyRef": "lmstudio" }
|
||||
```
|
||||
|
||||
## CLI version
|
||||
|
||||
Verified against `claude.exe` **2.1.221** by scanning the binary — every variable
|
||||
this tool writes is referenced by it:
|
||||
|
||||
`ANTHROPIC_BASE_URL` · `ANTHROPIC_AUTH_TOKEN` · `ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU,FABLE}_MODEL` ·
|
||||
`CLAUDE_CODE_SUBAGENT_MODEL` · `CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY` ·
|
||||
`CLAUDE_CODE_ATTRIBUTION_HEADER` · `CLAUDE_CODE_AUTO_COMPACT_WINDOW` ·
|
||||
`CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC` · `API_TIMEOUT_MS` · `apiKeyHelper`
|
||||
|
||||
(On 2.1.89 the fable and gateway-discovery vars did not exist; the update to
|
||||
2.1.221 added both.)
|
||||
|
||||
## Setup
|
||||
|
||||
```powershell
|
||||
cd c:\Users\smoido\Projects\cli\claude-code-switcher
|
||||
.\install.ps1
|
||||
```
|
||||
|
||||
Installs to `~/.claude-mode/` (ACL: you only), drops `claude-mode.cmd` into
|
||||
`~/.local/bin` (already on your User PATH, next to `claude.exe`), and adds a
|
||||
marked block to `~/Documents/WindowsPowerShell/profile.ps1`.
|
||||
`~/.claude/settings.json` is **not** touched by the installer — only by an actual
|
||||
mode switch, which backs it up to `~/.claude-mode/backups/` first (last 20 kept).
|
||||
|
||||
If the profile doesn't load: `Set-ExecutionPolicy -Scope CurrentUser RemoteSigned`.
|
||||
|
||||
## Restarting sessions
|
||||
|
||||
Claude Code reads all of this **once at startup**. A switch does not affect a
|
||||
running session — that's how the process loads config, not something scripting
|
||||
can change.
|
||||
|
||||
- **CLI** — exit and relaunch `claude`
|
||||
- **VS Code** — `Ctrl+Shift+P` → *Developer: Reload Window*
|
||||
- **Desktop app** — quit and reopen
|
||||
|
||||
`claude-mode status` shows what the *next* launch will use.
|
||||
|
||||
## Layout
|
||||
|
||||
```
|
||||
~/.claude-mode/
|
||||
claude-mode.ps1 main script
|
||||
state.json mode, active preset, and the exact env keys last written
|
||||
presets/*.json provider + model maps
|
||||
vault/*.cred DPAPI-encrypted keys (openrouter, zai, ...)
|
||||
backups/ settings.json snapshots + removed env-var values
|
||||
bin/claude-key-helper.cmd apiKeyHelper shim
|
||||
~/.local/bin/claude-mode.cmd PATH entry point (works from any shell)
|
||||
~/Documents/WindowsPowerShell/profile.ps1
|
||||
claude-mode + claude functions, between markers
|
||||
```
|
||||
|
||||
`state.json` records which env keys the last switch actually wrote, so a custom
|
||||
`extraEnv` key (Z.AI's timeouts, LM Studio's attribution header) is removed when
|
||||
you switch away — even though no other preset knows that key exists.
|
||||
|
||||
## Uninstall
|
||||
|
||||
```powershell
|
||||
claude-mode anthropic # clean settings.json first
|
||||
Remove-Item ~\.claude-mode -Recurse -Force
|
||||
Remove-Item ~\.local\bin\claude-mode.cmd
|
||||
# then delete the block between the >>> claude-mode >>> markers in profile.ps1
|
||||
```
|
||||
Reference in New Issue
Block a user