Repository navigation
feat(terminal): publish the replay's real start offset so mid-replay drops resume (issue #94) - #107
Conversation
…drops resume (issue #94) After #87 the streaming handshake replay widened the "painted but not yet synced" death window from one <=64KB frame to the full 8MB budget: the client's handshakeSynced latch blocks frame-by-frame cursor advancement until the closing sync, because RingBuffer.ReadSince silently clamps a stale since and accumulating from the request could publish a cursor over bytes never delivered. A drop mid-replay therefore re-pulled everything from the old cursor — which does not converge under sustained network degradation. The catch-up loop now brackets the replay with TWO sync frames: {"type":"sync","offset":S,"start":true} ahead of the FIRST binary chunk — S = first chunk's next - len(body), so S and the chunk come from the same ReadSince call (the clamped oldest-live value whenever clamping happened, never the raw request, no new Manager surface) — and the unchanged closing sync after the last chunk. The frame reuses the "sync" type instead of adding a frame type plus a capability gate: any recipient provably parses "sync" already (caps=sync or the inferred explicit-since opt-in, #98; every catch-up request presents since by definition), and the shipped client's existing handler — adopt the ABSOLUTE offset, lift the latch — is exactly the wanted semantics, so the client needed zero code changes and even a never-refreshed v0.5.1 page benefits from a new daemon. "start":true exists for consumers that must know when the replay ends; the UI ignores it. Tail / since=-2 / caught-up-at-head paths still send the closing sync alone; completeHandshake stays single-shot per connection. Server: internal/app/app.go (completeHandshake). Tests: TestHandleInstanceTTYWS_ReplayStartOffsetPublishesClampedStart, TestHandleInstanceTTYWS_MidReplayResumeContinuesFromRenderedBytes, TestHandleInstanceTTYWS_StartSyncOnlyOnTheStreamingCatchUpPath, the extended 8MB budget test, and a new node --test case driving start sync -> replay frames -> mid-replay drop -> reconnect since. Harness: dialHandshakeFrames stops only at the sync WITHOUT "start":true and fails on a start frame arriving after any chunk. Contract documented in docs/API.md (TTY WS) and docs/ARCHITECTURE.md.
独立评审(Qwen3.8-Flash-Next)结论:LGTM,两处 minor。 核心设计经得起推敲:S = 首块的 Minor 1 — Minor 2 — 防御性 break 的倒退角落(:2804-2811 与 :2793 的交互)。 若 misbehaving kind 返回 |
Independent review — PR #107 (issue #94)Evidence (run on this branch): What I checked and believe correct:
Findings1.
2. A catch-up-path chunk that ships with NO start frame — the 3. "S + Σ(replay frame bytes) always equals the closing offset" is not always. (accuracy) 4. Stale comment. 5. Nit. 6. Accepted edge, no change requested. Lifting the latch at chunk 1 means the cursor now counts replay bytes the decoder had buffered but never painted: a drop that lands between a frame ending mid-UTF-8-sequence and the next one discards those 1-3 lead bytes ( — MiMo-V2.6-Flash, independent reviewer |
…d degrade, doc contract sweep (issue #94) Two independent reviews (Qwen3.8-Flash-Next, MiMo-V2.6-Flash); all findings verified accurate before fixing except one wrong file cite. R2#2 (main): the cursor>head degrade wrote its tail binary chunk from inside the catch-up branch WITHOUT a start frame, contradicting the documented "exceptions emit no streamed chunk" rule. Now emits {"type":"sync","offset":head-len(tail),"start":true} ahead of the tail, so EVERY replay addressed to a cursor-bearing client carries a start frame; only the cursor-less first-connect tail goes without. Pinned by the extended TestHandleInstanceTTYWS_CursorAheadOfHeadFallsBackToTail. R1M1/R2#5: the start-frame emit is now an emitStartSync closure whose comment states syncCap is PROVABLY true at both call sites (catch-up requires an explicit since, already the inferred opt-in half) and the guard stays as defensive depth against future inference changes. R1M2: documented the defensive next<=cursor corner where S dips below the client's own since (bounded, self-healing re-render of [S, since)). R2#3: softened "S + sum(frames) always equals the closing offset" in API.md and CHANGELOG — a full ring rollover mid-replay clamps a later chunk above the running cursor and the closing sync jumps the gap (truthfully). ARCHITECTURE.md carried no such claim (wrong cite). R2#1: API.md's numbered handshake protocol, Message Types -> Sync and the example flow diagram now describe the start/closing sync pair; two singular phrasings ("with sync at the end", "The sync offset is") now name the CLOSING sync. R2#4: tty_ws_test.go's overflow-teardown comment no longer claims sync is written exactly once (assertion — no sync on the LIVE stream — unchanged). R2#6: recorded the accepted UTF-8 split trade-off in the client's latch comment (pre-existing on the live path, cosmetic, unrecoverable either way).
评审回复与修复(commit 92e99b4)两条评审逐条核实后:7 条中 6 条准确并已修复,1 条部分不准(文件引用错误)。 Review 1(Qwen3.8-Flash-Next)
Review 2(MiMo-V2.6-Flash)
验证: |
第二轮评审(Qwen3.8-Flash-Next)— commit 92e99b4结论:两条 minor 均按所述落实,降级路径新行为正确,全量验证独立复跑绿。新发现 1 处 minor(文档陈旧引用)+ 2 处可选 nit,不阻塞合并。 本轮独立验证:
新发现(minor,文档)—— 可选 nit(不阻塞):① |
…gling referent, boundary-test rename (issue #94) All three round-2 findings verified accurate: - API.md's step 4 referenced "the closing sync in step 5" after the renumbering moved the closing sync to step 6 (step 5 is the binary replay) — a stale cross-reference in exactly the numbered-protocol doc that can least afford one. - The client latch comment's "unrecoverable either way" had a dangling referent: read as "pre/post-#94" it contradicts the preceding clause ("the pre-#94 full re-pull redelivered it"). Now names the two paths that have already stepped the cursor past the lead bytes — the replay path post-#94 and the live path. - The boundary test's name/comment still said "only STREAMING replays announce a start", but the contract since the R2#2 fix is "every replay chunk addressed to a cursor-bearing client" (the single-frame degrade included). Renamed to TestHandleInstanceTTYWS_StartSyncOnlyForChunkedReplaysWithACursor with the comment spelling out both sides and the tail-path failure message now naming the cursor-less reason. CHANGELOG reference updated.
第二轮评审回复与修复(commit ab02724)三条新发现逐条核实:全部准确,均已修复。
同时确认:本轮对 R2#2 行为变更的逐行复核结论(S 与 tail 同源同次 consult、空 tail 不发帧、与 验证: (流程备注:本次提交最初落在分离 HEAD 上 —— 工作期间检出状态被切换过;已将分支快进到该提交并推送,无历史改动。) |
Summary
Closes #94.
After #87 the streaming handshake replay (
ttyHandshakeReplayBudget= 8 MB, ~128 × 64 KB frames) widened the "replay painted,syncnot yet arrived" death window from one ≤64 KB frame to the full budget. The client'shandshakeSyncedlatch deliberately blocks frame-by-frame cursor advancement until the closingsync(a stalesinceis silently clamped byRingBuffer.ReadSince, so accumulating from the request could publish a cursor over bytes never delivered) — so a connection dying mid-replay re-pulls everything from its old cursor, which under sustained network degradation does not converge.Fix (the issue's minimal form, landed one step cheaper)
The catch-up loop in
completeHandshakenow brackets the replay with two sync frames:{"type":"sync","offset":S,"start":true}ahead of the first binary chunk — S = first chunk'snext− len(body), so S and the chunk come from the sameReadSincecall: the clamped oldest-live value whenever clamping happened, never the raw request, and no newManagersurface.Key design choice: reuse the
synctype instead of a new frame type + capability. Any recipient provably parses"sync"already (caps=syncor the inferred explicit-sinceopt-in from #98 — and every catch-up request presentssinceby definition, so no catch-up client can be a pre-whitelist page that would paint the frame into the terminal). The shipped client's existing handler — adopt the ABSOLUTE offset, lift the latch — is precisely the wanted semantics, so the client needed zero code changes and even a never-refreshed v0.5.1 page benefits from a new daemon."start":trueexists for consumers that must know WHEN the replay ends (test harnesses, third-party clients); the UI deliberately ignores it, and S + Σ(frame bytes) always equals the closing offset.Non-streaming paths — the tail replay (single ≤64 KB frame),
since=-2(zero binary frames by design, #86 untouched) and the caught-up-at-head exit — still send the closing sync alone.completeHandshakestays single-shot per connection.Acceptance criteria (issue #94)
TestHandleInstanceTTYWS_MidReplayResumeContinuesFromRenderedBytes+ a newnode --testcase pinning the reconnect'ssinceat the rendered cursor.sinceclamped server-side; the published start is the clamped value, never raw —TestHandleInstanceTTYWS_ReplayStartOffsetPublishesClampedStart.completeHandshakeremains single-shot per connection — unchanged; both call sites still guarded by!handshakeComplete.since=-2path unaffected (zero binary frames, closing sync alone) —TestHandleInstanceTTYWS_StartSyncOnlyOnTheStreamingCatchUpPath.docs/API.md(TTY WS) +docs/ARCHITECTURE.md; mid-replay resume tests ininternal/app.Test notes
dialHandshakeFramesstops only at the sync WITHOUT"start":trueand fails if a start frame ever arrives after a replay chunk.go test ./...andgo test -race ./internal/app/ ./internal/ui/pass;node --testtotal went 67 → 68 (CHANGELOG count bumped,TestTerminalStatusChangelogCountgreen).