Ask before switching while sessions are live, and recover the ones already broken

Acts on docs/incident-mode-switch-corrupts-live-sessions.md, which is added here
as the record of why.

The report identifies a consequence that was not modelled. A failed call is
recoverable; a *successful* one may not be. If a running session takes even one
completion from the provider being switched to - which happens when that mode
matches the base URL it already had cached - that provider's message-id format
lands in its transcript. OpenRouter issues `gen-<epoch>-<rand>` where Anthropic
issues `msg_...`, and native Anthropic then refuses to resume the session at
all, with a 400 naming previous_message_id. The only way back is to truncate the
transcript, losing every turn after the cut. That happened here, and was fixed
by hand.

Two changes follow.

Sessions are now settled before the write, not reported after it. A switch with
anything running stops, names the sessions, explains what is about to happen to
them, and offers restart (the only answer that ends with everything on the mode
the bar now claims), close, proceed anyway, or abort - defaulting to abort.
Non-interactively it refuses outright unless given --yes. The old after-the-fact
reporter is deleted rather than left as a second, contradictory account.

`claude-mode repair-session` replaces the hand surgery: it finds a project's
transcripts, reports which are resumable, and on --apply backs the file up and
truncates to the last Anthropic-issued message. Verified against the real
corrupted transcript from the incident - it reproduces the manual cut exactly,
1921 lines to 1813, dropping the two `gen-` completions and the error
placeholders after them, leaving a transcript that ends on a genuine msg_ id. It
refuses a transcript written to in the last 90 seconds, since that one belongs
to a session still running.

The panel passes --yes, having already asked in its own card, and that card now
names the transcript risk rather than only the inconvenient one.

Requirement 4 of the report - documenting the mechanism - landed in 9c301e1;
the README now carries the unrecoverable half as well.
This commit is contained in:
smoido
2026-08-31 02:05:37 +03:00
parent 9ff9b82c25
commit a4afa55580
5 changed files with 490 additions and 27 deletions
+37 -2
View File
@@ -322,8 +322,38 @@ switch reaches into a live session through the one thing that was never cached:
| another preset of the *same* provider, same `keyRef` | same key, same endpoint — this one survives, on the model ids it started with | | another preset of the *same* provider, same `keyRef` | same key, same endpoint — this one survives, on the model ids it started with |
Either of the first two starts failing calls whenever the TTL happens to expire: Either of the first two starts failing calls whenever the TTL happens to expire:
mid-turn as easily as between turns. So restart afterwards — and on Linux, mid-turn as easily as between turns.
`claude-mode sessions --restart` will do it for you:
### The part that is not just an inconvenience
A failed call is recoverable. A *successful* one may not be.
If a running session takes even one completion from the new provider before
anything notices — which happens when the mode it is switched to matches the
base URL it already had cached — that provider's message-id format lands in its
transcript. OpenRouter issues `gen-<epoch>-<rand>` where Anthropic issues
`msg_…`. Native Anthropic then refuses to resume the session at all:
```
API Error: 400 diagnostics.previous_message_id: must be the `id` from a
prior /v1/messages response (starts with `msg_`)
```
There is no supported way back from that. The only fix is to truncate the
transcript to the last message Anthropic issued, losing every turn after it:
```bash
claude-mode repair-session # list transcripts here and their state
claude-mode repair-session <session-id> # show what it would cut
claude-mode repair-session <session-id> --apply
```
It backs the original up first, and refuses to touch a transcript that was
written to in the last 90 seconds, since that belongs to a session still alive.
This is why a switch now asks before it writes rather than reporting afterwards.
So restart affected sessions — and `claude-mode sessions --restart` will do it
for you:
- **CLI** — exit and relaunch `claude` - **CLI** — exit and relaunch `claude`
- **VS Code** — `Ctrl+Shift+P`*Developer: Reload Window* - **VS Code** — `Ctrl+Shift+P`*Developer: Reload Window*
@@ -506,6 +536,11 @@ endpoint is fixed at startup but their credential is re-fetched on a timer, so
it switches under them and the endpoint they are still pointed at refuses it. it switches under them and the endpoint they are still pointed at refuses it.
Which is why this is a decision rather than a notification. Which is why this is a decision rather than a notification.
A switch stops and asks while any are running, offering to restart them (the
only answer that ends with everything on the mode the bar now claims), close
them, proceed anyway, or abort — and abort is the default. Non-interactively it
refuses outright unless given `--yes`.
``` ```
claude-mode sessions claude-mode sessions
running sessions (2) running sessions (2)
@@ -0,0 +1,142 @@
# Incident: switching modes mid-session corrupts and can permanently break that session
**Date:** 2026-08-30/31
**Affected session:** `1fffec9e-e2e3-4255-828c-a9ccfd8f631d` (project `/home/smoido/Work`)
**Severity:** High — silent, mid-task corruption with no warning, and one failure path is not recoverable by normal means (requires manual transcript surgery).
## Summary
The user had a live, running Claude Code session open in `~/Work`. While that
session was mid-task, `claude-mode` was used to switch the active provider
away from `anthropic` and later back to `anthropic`. The running session was
not restarted in between. This corrupted the session's credentials, and — because
the session had picked up a different provider's message-ID format in the
meantime — it became permanently unable to resume under native Anthropic auth,
throwing a hard API error on every resume attempt.
This must not be possible to trigger silently. Whatever ships next needs to
either prevent it, warn loudly before it happens, or make recovery automatic.
## Root cause / mechanism
`claude-mode` writes provider config into `~/.claude/settings.json`
(`env` block: base URL, model IDs, etc., plus `apiKeyHelper` pointing at
`claude-key-helper.sh`). A **running** `claude` process only reads part of
this at startup:
- **Static, cached at startup:** `ANTHROPIC_BASE_URL`, the model ID env vars,
the rest of the `env` block. A running session keeps whatever it started
with here — switching modes does *not* change these for an
already-running process.
- **Not cached — re-fetched on a timer:** the credential. Claude Code
re-invokes `apiKeyHelper` periodically (`CLAUDE_CODE_API_KEY_HELPER_TTL_MS`
is present in the 2.1.251 binary), and `claude-key-helper.sh` answers based
on whatever `~/.claude-mode/state.json` says *at the moment it's called*
not at session startup.
So a mode switch reaches into a live session through the one part that was
never cached:
| switching the global mode to... | what the *running* session gets on next credential refresh |
|---|---|
| `anthropic` | the helper returns nothing (native anthropic auth is expected to need no helper) → session has no credential at all |
| any other provider/preset | the new key, but the session is still pointed at the **old** base URL (cached at startup) → that endpoint rejects the new key |
Either way, the session starts failing API calls from the moment the TTL
next expires — mid-turn as easily as between turns. The one case that does
*not* break: switching between two presets of the *same* provider that share
a `keyRef` (same key, same endpoint survives).
`claude-mode status` and `claude-mode health` only ever report the *global*
config's current state. Neither one has any way to know a specific already-running
session exists, let alone that it's about to be (or has been) knocked over.
## What happened to this specific session (concrete failure chain)
1. Session `1fffec9e` was live in `~/Work`, working on `claude-code-switcher`
itself.
2. User ran a switch to the `openrouter` "default" preset (via `claude-mode`)
to test something, while that session kept running.
3. The running session's next few turns actually succeeded against
openrouter (its cached base URL now matched, since the switch happened to
land on the provider its env pointed at) — those turns got real
completions back, but with **openrouter's message-ID format**
(`gen-<timestamp>-<id>`), not Anthropic's (`msg_<id>`).
4. A later call failed outright and Claude Code recorded a synthetic
client-side placeholder turn (`"model":"<synthetic>"`,
`"isApiErrorMessage":true`, `"error":"unknown"`) as the last message in
the transcript — this is the visible "session disconnected mid-task"
symptom.
5. User switched the global mode back to `anthropic`.
6. Resuming session `1fffec9e` (`claude -r 1fffec9e-... -p "..."`) now fails
unconditionally with:
```
API Error: 400 diagnostics.previous_message_id: must be the `id` from a
prior /v1/messages response (starts with `msg_`)
```
because native Anthropic's API requires `previous_message_id` to be a
real Anthropic-issued ID, and the last message(s) in this session's
transcript are not (`gen-...` or the synthetic placeholder ID).
7. **There is no supported way to resume past this.** The only fix found was
manual surgery on the session's `.jsonl` transcript file: locate the last
message that still has a genuine `msg_...` id (in this case, several
turns earlier, at a clean `end_turn` boundary well before the switch was
even tested), back up the original file, and truncate the transcript to
that point. That rolls the session back to its last-known-good state and
makes it resumable again — at the cost of permanently losing every turn
after that point (the actual code changes from that later work were
separately safe in git, but the chat narrative was not recoverable).
## Impact
- Silent corruption: nothing warns the user before or during the switch that
a live session exists and is about to break.
- Depending on timing, the break can be "just" an auth failure (annoying,
session still resumable once you're back on the mode it started with) **or**
a hard, unrecoverable-by-normal-means failure (if the session round-tripped
through a different provider's ID format before failing) that requires
hand-editing a JSONL transcript to fix.
- This applies to *every* running `claude` process on the machine at switch
time, not just the one in the foreground shell — CLI, VS Code, desktop, any
of them.
## Requirements for the fix
1. **Detect running sessions before switching.** `claude-mode` should scan
for live `claude` processes (and ideally which project/cwd each belongs
to) before performing a switch.
2. **Warn or block, don't silently proceed.** At minimum, print a clear
warning naming the affected session(s)/PIDs/cwds and what will happen to
them (credential will be pulled out from under them on the next TTL
refresh). Consider requiring `--force` (or an explicit confirmation) to
proceed while sessions are running, and defaulting to "abort" otherwise.
3. **Prefer a safe path when sessions are detected:** e.g. offer to let the
user gracefully end/save those sessions first, or clearly instruct them
to restart affected sessions immediately after the switch completes.
4. **Document the real mechanism** (this file's "Root cause" section) in the
tool's own help/README, replacing any prior claim that a switch "does not
affect a running session" or that sessions "keep talking to the old
provider until restarted" — both are wrong; the credential moves under
them regardless.
5. **Make the unrecoverable failure mode recoverable.** Add a `claude-mode
repair` (or similar) command that:
- finds a given session's transcript,
- locates the last message with a valid `msg_...` id,
- backs up the original file,
- truncates to that point,
so this doesn't require manual `jq`/`head`/`grep` surgery next time.
This is the exact procedure used to fix session `1fffec9e` above.
## Repro steps (for verification once fixed)
1. Start a `claude` session in some project directory, mid-task.
2. In a separate shell, run `claude-mode <some-other-provider>` (something
with a different base URL/key format from the session's current mode).
3. Let the running session's next API-key TTL refresh happen (or just issue
another prompt in it) — observe the call fail.
4. Switch back: `claude-mode anthropic`.
5. Try `claude -r <that-session-id> -p "hi"` — currently fails with the
`previous_message_id` 400 error if the session ever got a non-`msg_`
completion in between. The fix should make step 2 impossible (or clearly
confirmed) rather than needing step 5's failure to be caught after the
fact.
+204 -21
View File
@@ -278,7 +278,9 @@ claude-mode - switch Claude Code between Anthropic, OpenRouter, Z.AI, LM Studio
claude-mode preflight <mode> [preset] check a mode can actually serve, without switching claude-mode preflight <mode> [preset] check a mode can actually serve, without switching
claude-mode sessions [--stop|--restart] claude-mode sessions [--stop|--restart]
running sessions; close or reopen them running sessions; close or reopen them
claude-mode repair-session [id] make a session resumable again after a bad switch
claude-mode <mode> --force switch even if preflight says no claude-mode <mode> --force switch even if preflight says no
claude-mode <mode> --yes switch without asking about running sessions
EOF EOF
} }
@@ -818,26 +820,6 @@ EOF_ROWS
return 0 return 0
} }
# Named after the switch, not before it: the switch has already happened, and
# these are the sessions it did not reach.
cm_report_live_sessions() {
local rows n busy
rows="$(cm_session_rows)"
n="$(printf '%s' "$rows" | grep -c . || true)"
[ "${n:-0}" -gt 0 ] || return 0
busy="$(printf '%s' "$rows" | cut -f4 | grep -c '^yes$' || true)"
printf '\n'
warn "$n Claude Code session(s) are running and will start failing their calls"
say 'their key is re-fetched on a timer and now resolves to the new mode, which'
say 'the endpoint they are still pointed at will not accept. Restart them.'
if [ "${busy:-0}" -gt 0 ]; then
warn "$busy of them is mid-request and will break wherever it happens to be"
fi
printf ' %sclaude-mode sessions what is running%s\n' "$C_DIM" "$C_RESET"
printf ' %sclaude-mode sessions --restart close and reopen them on the new provider%s\n' "$C_DIM" "$C_RESET"
}
cm_terminal_cmd() { cm_terminal_cmd() {
local t local t
for t in "${TERMINAL:-}" foot alacritty ghostty kitty; do for t in "${TERMINAL:-}" foot alacritty ghostty kitty; do
@@ -898,6 +880,8 @@ set_mode() {
[ -f "$preset_file" ] || { err "preset '$preset_name' not found"; return 1; } [ -f "$preset_file" ] || { err "preset '$preset_name' not found"; return 1; }
fi fi
cm_confirm_sessions "$mode" || return 1
backup="$(backup_settings)" backup="$(backup_settings)"
if ! "$PY" "$JSON" apply "$CM_SETTINGS" "$CM_STATE" "$mode" "$preset_file" "$CM_HELPER" >/dev/null; then if ! "$PY" "$JSON" apply "$CM_SETTINGS" "$CM_STATE" "$mode" "$preset_file" "$CM_HELPER" >/dev/null; then
@@ -945,7 +929,7 @@ set_mode() {
check_stale_models "$mode" check_stale_models "$mode"
write_health "$mode" "$preset_name" write_health "$mode" "$preset_name"
cm_report_live_sessions cm_apply_session_action
printf '\n %srestart claude (and reload the VS Code window) to pick this up%s\n' "$C_DIM" "$C_RESET" printf '\n %srestart claude (and reload the VS Code window) to pick this up%s\n' "$C_DIM" "$C_RESET"
} }
@@ -2034,6 +2018,203 @@ cmd_setup() {
return 0 return 0
} }
# ---------------------------------------------------------------------------
# Live sessions: asked before the write, not reported after it
#
# The damage a switch does to a running session is not limited to it failing
# calls. If the session takes even one completion from the new provider before
# anything notices, that provider's message-id format lands in its transcript -
# OpenRouter issues `gen-<epoch>-<rand>` where Anthropic issues `msg_...` - and
# native Anthropic then refuses to resume the session at all:
#
# API Error: 400 diagnostics.previous_message_id: must be the `id` from a
# prior /v1/messages response (starts with `msg_`)
#
# There is no supported way back from that. The only fix is to truncate the
# transcript to the last message Anthropic issued, losing everything after it
# (see `claude-mode repair-session`). A confirmation that costs one keystroke is
# cheap against a failure that costs an afternoon of conversation.
# ---------------------------------------------------------------------------
CM_ASSUME_YES=0
CM_SESSION_ACTION=none
CM_SESSION_ROWS=''
cm_confirm_sessions() {
local mode="$1" rows n busy reply
CM_SESSION_ACTION=none
CM_SESSION_ROWS=''
rows="$(cm_session_rows)"
n="$(printf '%s' "$rows" | grep -c . || true)"
[ "${n:-0}" -gt 0 ] || return 0
CM_SESSION_ROWS="$rows"
busy="$(printf '%s' "$rows" | cut -f4 | grep -c '^yes$' || true)"
printf '\n'
warn "$n Claude Code session(s) are running right now"
local pid ppid tty b cwd isself pcmd tag
while IFS=$'\t' read -r pid ppid tty b cwd isself pcmd; do
[ -n "$pid" ] || continue
tag=''
[ "$b" = yes ] && tag=" ${C_YELLOW}working${C_RESET}"
[ "$isself" = yes ] && tag="$tag ${C_DIM}(this one)${C_RESET}"
printf ' %-8s %-8s %s%s\n' "$pid" "$tty" "$cwd" "$tag"
done <<EOF_ROWS
$rows
EOF_ROWS
printf '\n'
say 'Their key is re-fetched on a timer and will resolve to the new mode,'
say 'which the endpoint they are still pointed at will not accept. If one'
say 'of them takes a reply from the new provider first, that provider'"'"'s'
say 'message-id format goes into its transcript and Anthropic will then'
say 'refuse to resume that session at all - recoverable only by truncating'
say 'it (claude-mode repair-session), which loses the turns after the cut.'
[ "${busy:-0}" -gt 0 ] && warn "$busy of them is mid-request and is the most likely to be caught"
if [ "$CM_ASSUME_YES" -eq 1 ]; then
say 'proceeding (--yes)'
return 0
fi
if ! ui_interactive; then
printf '\n'
err 'refusing to switch while sessions are running'
say 'restart or close them first, or pass --yes to switch anyway'
return 1
fi
printf '\n'
say 'r switch, then close and reopen them on the new mode (safest)'
say 'c switch, then close them'
say 's switch and leave them running (risks the above)'
say 'a abort'
printf '\n [r/c/s/A] '
IFS= read -r reply
case "$reply" in
r|R) CM_SESSION_ACTION=restart; return 0 ;;
c|C) CM_SESSION_ACTION=stop; return 0 ;;
s|S) CM_SESSION_ACTION=none; return 0 ;;
*) say 'aborted; nothing was changed'; return 1 ;;
esac
}
# Run after the write, never before: a session reopened first would come back up
# on the mode being left behind.
cm_apply_session_action() {
[ "$CM_SESSION_ACTION" = "none" ] && return 0
[ -n "$CM_SESSION_ROWS" ] || return 0
printf '\n'
cm_session_act "$CM_SESSION_ACTION" "$CM_SESSION_ROWS" 0
}
# ---------------------------------------------------------------------------
# Transcript repair
# ---------------------------------------------------------------------------
# ~/.claude/projects/<cwd with every slash turned into a dash>
cm_project_dir() {
printf '%s/projects/%s' "$CM_SETTINGS_DIR" "$(printf '%s' "${1:-$PWD}" | sed 's|/|-|g')"
}
cmd_repair_session() {
local target='' apply=0 dir='' a file verdict age
for a in "$@"; do
case "$a" in
--apply) apply=1 ;;
--dry-run) apply=0 ;;
--list) target='--list' ;;
-*) err "unknown option '$a'"; return 1 ;;
*) target="$a" ;;
esac
done
# Claude Code keys transcripts by the directory the session was started in,
# which is rarely the one you are standing in when you come to fix it. Walk
# up first, and for a named session fall back to looking through every
# project - the id is unique, so there is nothing ambiguous to resolve.
local probe="$PWD"
while [ -n "$probe" ]; do
[ -d "$(cm_project_dir "$probe")" ] && { dir="$(cm_project_dir "$probe")"; break; }
[ "$probe" = "/" ] && break
probe="$(dirname "$probe")"
done
if [ -n "$target" ] && [ "$target" != "--list" ]; then
if [ -z "$dir" ] || [ ! -f "$dir/${target%.jsonl}.jsonl" ]; then
local hit
hit="$(ls -1 "$CM_SETTINGS_DIR"/projects/*/"${target%.jsonl}".jsonl 2>/dev/null | head -n1)"
[ -n "$hit" ] && dir="$(dirname "$hit")"
fi
fi
if [ -z "$dir" ] || [ ! -d "$dir" ]; then
err 'no session transcripts found for this directory'
say 'run it from the project the session belongs to, or name the session id'
return 1
fi
if [ -z "$target" ] || [ "$target" = "--list" ]; then
head_ 'session transcripts here'
local f v
for f in $(ls -1t "$dir"/*.jsonl 2>/dev/null); do
v="$("$PY" "$JSON" repair-session "$f" 2>/dev/null)" || continue
printf '%s' "$v" | "$PY" -c "
import json,sys,os
d=json.load(sys.stdin)
state = 'ok' if d['healthy'] else ('repairable, would drop %d line(s)' % d['dropLines'] if d['repairable'] else 'no anthropic message found')
print(' %-40s %s' % (os.path.basename(d['path'])[:-6], state))
"
done
printf '\n %sclaude-mode repair-session <session-id> --apply%s\n' "$C_DIM" "$C_RESET"
return 0
fi
file="$dir/${target%.jsonl}.jsonl"
[ -f "$file" ] || { err "no transcript $file"; return 1; }
# A transcript that is still being appended to belongs to a session that is
# still alive; truncating it underneath a running process helps nobody.
age=$(( $(date +%s) - $(stat -c %Y "$file" 2>/dev/null || echo 0) ))
if [ "$age" -lt 90 ] && [ "$apply" -eq 1 ]; then
err "that transcript was written to ${age}s ago - it looks live"
say 'close the session that owns it first'
return 1
fi
verdict="$("$PY" "$JSON" repair-session "$file" $([ "$apply" -eq 1 ] && printf -- '--apply'))" || {
err 'could not read that transcript'; return 1; }
printf '%s' "$verdict" | "$PY" -c "
import json,sys
d=json.load(sys.stdin)
G,Y,R,D,X = '\033[32m','\033[33m','\033[91m','\033[90m','\033[0m'
print()
if d['healthy']:
print(' %sok %s last message is Anthropic-issued; nothing to repair' % (G,X))
raise SystemExit(0)
if not d['repairable']:
print(' %sFAIL%s no Anthropic-issued message anywhere in this transcript' % (R,X))
raise SystemExit(1)
for f in d['foreignIds']:
print(' %swarn%s line %d carries a %s id from %s' % (Y,X,f['line'],f['id'].split('-')[0]+'-',f['model'] or 'another provider'))
n = sum(1 for s in d['syntheticIds'] if s['apiError'] and s['line'] > d['lastGoodLine'])
if n:
print(' %swarn%s %d client-side error placeholder(s) after the last good message' % (Y,X,n))
if d['applied']:
print(' %sok %s truncated to line %d, dropping %d' % (G,X,d['lastGoodLine'],d['dropLines']))
print(' %sok %s original saved as %s' % (G,X,d['backup']))
print()
print(' %sthat session should resume again; the dropped turns are gone%s' % (D,X))
else:
print(' %swarn%s would truncate to line %d, dropping %d line(s)' % (Y,X,d['lastGoodLine'],d['dropLines']))
print()
print(' %sre-run with --apply to do it (the original is backed up first)%s' % (D,X))
"
}
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Dispatch # Dispatch
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@@ -2047,6 +2228,7 @@ _args=()
for _a in "$@"; do for _a in "$@"; do
case "$_a" in case "$_a" in
--force) CM_FORCE=1 ;; --force) CM_FORCE=1 ;;
-y|--yes) CM_ASSUME_YES=1 ;;
*) _args+=("$_a") ;; *) _args+=("$_a") ;;
esac esac
done done
@@ -2109,6 +2291,7 @@ case "$cmd" in
cmd_setup "$_mode" "$_rest" cmd_setup "$_mode" "$_rest"
fi ;; fi ;;
sessions) cmd_sessions "$@" ;; sessions) cmd_sessions "$@" ;;
repair-session) cmd_repair_session "$@" ;;
repair) repair)
scope='' scope=''
for a in "$@"; do [ "$a" = "--all" ] && scope=all; done for a in "$@"; do [ "$a" = "--all" ] && scope=all; done
+97
View File
@@ -603,10 +603,107 @@ def cmd_set_all(argv):
save(path, p) save(path, p)
print(model) print(model)
# ---------------------------------------------------------------------------
# Session transcript repair
#
# Native Anthropic requires previous_message_id to be an id it issued itself -
# one starting `msg_`. A session that took even one completion from a gateway
# while the mode was switched under it has that provider's id format in its
# transcript instead (OpenRouter issues `gen-<epoch>-<rand>`), and every attempt
# to resume it afterwards fails with a 400 naming previous_message_id. Client-
# side error placeholders, written as model `<synthetic>` with a UUID for an id,
# do the same thing when one is last.
#
# The transcript is newline-delimited JSON, one independent object per line, so
# rolling back to the last message Anthropic actually issued is a truncation.
# Everything after it is lost - which is the cost, and why nothing here runs
# without being asked twice.
# ---------------------------------------------------------------------------
def _msg_id(obj):
m = obj.get("message")
if isinstance(m, dict) and m.get("id"):
return str(m["id"])
return None
def cmd_repair_session(argv):
"""repair-session <transcript.jsonl> [--apply]
Prints a JSON verdict. With --apply, backs the file up and truncates it to
the last Anthropic-issued message.
"""
path = argv[0]
apply_it = "--apply" in argv[1:]
with open(path, encoding="utf-8") as fh:
lines = fh.read().splitlines()
last_good = -1 # index of the last line carrying a msg_ id
foreign, synthetic = [], []
for i, line in enumerate(lines):
if not line.strip():
continue
try:
obj = json.loads(line)
except Exception:
continue
mid = _msg_id(obj)
if mid is None:
continue
if mid.startswith("msg_"):
last_good = i
elif mid.startswith("gen-"):
foreign.append({"line": i + 1, "id": mid,
"model": str((obj.get("message") or {}).get("model", ""))})
else:
synthetic.append({"line": i + 1, "id": mid,
"apiError": bool(obj.get("isApiErrorMessage"))})
# Only the *last* id matters for resuming: an error placeholder in the
# middle of a long-finished turn is history, not a blocker.
tail_ids = [i for i in range(len(lines) - 1, last_good, -1)
if lines[i].strip() and _msg_id(_safe(lines[i])) is not None]
healthy = (last_good >= 0 and not tail_ids)
out = {
"path": path,
"lines": len(lines),
"lastGoodLine": last_good + 1 if last_good >= 0 else 0,
"dropLines": 0 if healthy or last_good < 0 else len(lines) - (last_good + 1),
"foreignIds": foreign,
"syntheticIds": synthetic,
"healthy": healthy,
"repairable": (not healthy) and last_good >= 0,
"applied": False,
"backup": "",
}
if apply_it and out["repairable"]:
stamp = __import__("datetime").datetime.now().strftime("%Y%m%d-%H%M%S")
backup = "%s.pre-repair-backup-%s" % (path, stamp)
with open(backup, "w", encoding="utf-8") as fh:
fh.write("\n".join(lines) + ("\n" if lines else ""))
with open(path, "w", encoding="utf-8") as fh:
fh.write("\n".join(lines[:last_good + 1]) + "\n")
out["applied"] = True
out["backup"] = backup
print(json.dumps(out))
def _safe(line):
try:
return json.loads(line)
except Exception:
return {}
COMMANDS = { COMMANDS = {
"health": cmd_health, "health": cmd_health,
"preflight-json": cmd_preflight_json, "preflight-json": cmd_preflight_json,
"sessions-json": cmd_sessions_json, "sessions-json": cmd_sessions_json,
"repair-session": cmd_repair_session,
"stale-models": cmd_stale_models, "stale-models": cmd_stale_models,
"strip-tags": cmd_strip_tags, "strip-tags": cmd_strip_tags,
"or-models": cmd_or_models, "or-models": cmd_or_models,
+9 -3
View File
@@ -220,8 +220,11 @@ Panel {
root.busy = true root.busy = true
root.stage = "list" root.stage = "list"
switchProc.sessionAction = sessionAction switchProc.sessionAction = sessionAction
// --yes because the confirmation already happened, in the card above. The
// CLI now refuses a non-interactive switch while sessions are running, and
// without this the panel's switch would simply stop working.
switchProc.command = root.cli(root.pendingPreset === "" switchProc.command = root.cli(root.pendingPreset === ""
? [root.pendingMode] : [root.pendingMode, root.pendingPreset]) ? [root.pendingMode, "--yes"] : [root.pendingMode, root.pendingPreset, "--yes"])
switchProc.running = true switchProc.running = true
} }
@@ -875,8 +878,9 @@ Panel {
Text { Text {
width: parent.width width: parent.width
text: "They keep pointing at " + Modes.title(root.mode) + ", but their key is " text: "They keep pointing at " + Modes.title(root.mode) + ", but their key is "
+ "re-fetched on a timer and will switch under them. Their calls start " + "re-fetched on a timer and will switch under them. Worse, if one takes a "
+ "failing from that moment, not at a clean stop." + "reply from the new provider first, that provider's message-id format "
+ "goes into its transcript and Anthropic will refuse to resume it at all."
color: Color.muted color: Color.muted
wrapMode: Text.WordWrap wrapMode: Text.WordWrap
lineHeight: 1.2 lineHeight: 1.2
@@ -950,6 +954,8 @@ Panel {
width: parent.width width: parent.width
spacing: Style.space(7) spacing: Style.space(7)
// Restart is the only option that ends with every session on the mode
// the bar is now claiming, so it leads and it is the primary.
PillButton { PillButton {
label: "Switch and restart" label: "Switch and restart"
primary: true primary: true