Skip to content

SSE stream tail is not delivered under Playwright's WebKit, blocking the WebKit CI matrix entry (plus a cross-chunk parseSSE bug) #2132

Description

@cliffhall

Corrected 2026-08-26. This issue was originally filed as "Web client SSE stream hangs in WebKit/Safari" and asserted that a Safari user opening an MCP App gets a permanently "loading" widget. That was not verified and appears to be wrong — a maintainer opened an MCP App in real Safari and it worked normally. The failure is real under Playwright's WebKit and blocks gating that engine in CI; it is not, on current evidence, a bug users hit. Rewritten below to say only what was actually observed. The parseSSE defect in the second section is engine-independent and stands on its own.

What actually fails

smoke:web:app and smoke:web:elicit pass in Chromium and Firefox and fail under Playwright's WebKit:

smoke:web:app [webkit] FAILED — app never reached data-app-status="ready" (last: "loading")

This is why webkit is absent from the Sandbox smokes matrix added by #2086 (PR #2133). Adding it back is a one-word diff once this is resolved.

What does NOT fail

Real Safari. Opening an MCP App in Safari renders it normally — no loading hang. So the caveat that Playwright's WebKit "is a WebKit build, not Safari" applies in both directions, and the original framing of this issue only applied it in one. Playwright ships its own WebKit build whose networking stack is not Safari's, which is precisely the layer this failure lives in.

Anyone with a reproduction in real Safari should say so here — that would change both the severity and the fix.

Diagnosis under Playwright's WebKit (verified, not inferred)

Instrumenting the page put the failure at the very end of the chain:

  1. POST /api/mcp/send for resources/read → 200 {"ok":true}. ✅
  2. The backend issues the upstream POST http://…/mcp and gets 200 text/event-stream. ✅ (visible in the fetch_request telemetry event, with its request body and response headers)
  3. The response never reaches the browser. ❌ No message event with that id, no fetch_request_body_update, no error, forever.

Teeing the /api/mcp/events response body inside the page shows the engine stops delivering bytes: it holds the tail of the stream until more data arrives. Chromium and Firefox deliver each write promptly, so ids 0–4 arrive normally in every engine; only the last write before the stream goes quiet is stranded — which, in an interactive session, is the response being waited on.

Two experiments pin the mechanism:

Change Result under Playwright's WebKit
Periodic :\n\n keep-alive on the stream (tried at 300 ms and 5 s) data-app-status="ready" ✅
One-time 2 KB padding on the existing priming comment still hangs ❌

So it is not the familiar "Safari buffers the first ~1 KB" threshold — a one-time pad does not fix it. It behaves like a per-delivery tail buffer, flushed only by more bytes.

A second, independent defect found in the same area

This one is engine-independent and worth fixing regardless of what happens above. parseSSE in core/mcp/remote/remoteClientTransport.ts resets its in-progress event inside the read loop:

while (true) {
  const { done, value } = await reader.read();
  …
  let currentEvent = "message";     // ← reset on every chunk
  let currentData: string[] = [];
  for (const line of lines) { … }
}

buffer correctly carries a partial line across read()s, but a partial frame is discarded: if a frame's data: line lands in one chunk and the blank line terminating it arrives in the next, the accumulated payload is dropped and the frame is never yielded — a silent hang, from a cause that has nothing to do with any particular browser.

It has not bitten us because Chromium and Firefox happen to deliver whole frames at the payload sizes this transport sees. Nothing guarantees that: chunk boundaries are a function of payload size, network conditions, and proxying, none of which we control.

Suggested scope

Split by confidence:

  1. Fix parseSSE (clear win, no open questions). Hoist currentEvent / currentData out of the read loop and flush a trailing frame when the stream ends without a final blank line. Unit-test with a chunk split placed deliberately between a data: line and its terminator — the current code fails that test.
  2. Decide what, if anything, to do about the WebKit tail buffer. A keep-alive on primeAndHoldSseStream fixes it and would also protect against idle-connection timeouts in proxies, which is independently reasonable. But the interval bounds worst-case message latency, and spending that on an engine no user runs is a real trade-off — worth confirming first whether the behavior exists in any shipping browser. Investigating Playwright's WebKit network stack, or reproducing under webkit on another platform, would settle it.

Notes

  • Reproduce with SMOKE_BROWSER=webkit npm run smoke:web:app (needs npx playwright install --with-deps webkit).
  • smoke:web:browser passes under WebKit — only the two App smokes fail.

Activity

  1. added this to the v2.5.0 milestone on Aug 26, 2026
  2. added
    bugSomething isn't working
    v2Issues and PRs for v2
    on Aug 26, 2026
  3. changed the title [-]Web client SSE stream hangs in WebKit/Safari: the last message on /api/mcp/events is never delivered[/-] [+]SSE stream tail is not delivered under Playwright's WebKit, blocking the WebKit CI matrix entry (plus a cross-chunk parseSSE bug)[/+] on Aug 26, 2026
  4. cliffhall commented on Aug 26, 2026

    @cliffhall
    MemberAuthor

    Correction — the Safari claim was wrong, and it was mine.

    @cliffhall opened an MCP App in real Safari and it rendered normally. No loading hang.

    I filed this asserting that "a Safari user opening an MCP App gets a permanently loading widget." I never tested Safari. I tested Playwright's WebKit, observed a real failure, and then wrote it up as a Safari bug — which is exactly the inference scripts/lib/headless-browser.mjs warns against in its own header, in the sentence I wrote:

    Playwright's WebKit is a WebKit build, not Safari … a green run here is not a Safari guarantee.

    I applied that caveat in one direction only. A red run is no more a Safari indictment than a green one is a Safari guarantee — and the divergence between the two is widest in the networking layer, which is where this failure lives.

    Retitled and rewritten to claim only what was observed. What changes:

    • Severity drops. This is a test-infrastructure blocker (it keeps webkit out of the CI matrix), not a user-facing bug. Lowering the board priority accordingly.
    • The fix is no longer obviously worth its cost. A keep-alive resolves it, but the interval bounds worst-case message latency, and spending that for an engine no user runs needs justifying rather than assuming. Now framed as a decision to make, not a fix to apply.
    • The parseSSE half is unaffected and still worth doing. A frame split across chunk boundaries is silently dropped; that is engine-independent, and nothing guarantees Chromium and Firefox keep delivering whole frames — chunk boundaries depend on payload size, network conditions and proxying.

    If anyone does reproduce this in a shipping browser, say so here — that would reverse both points above.

  5. cliffhall commented on Aug 26, 2026

    @cliffhall
    MemberAuthor

    Re-triage: Priority Medium (total 7), down from High.

    The original High rested on "actively hurting users on a published version right now", which is not supported by the evidence.

  6. cliffhall commented on Aug 26, 2026

    @cliffhall
    MemberAuthor

    Second correction: the mechanism described above is not supported by the evidence. Retracting it.

    I built an isolated repro of the mechanism I claimed — a plain HTTP server that writes three SSE frames (200B, 4KB, 4KB) and then goes silent with the connection held open, and a page that reads it with fetch + getReader(). If the stream's tail were being held until more bytes arrived, frame 3 would never be delivered.

    It is delivered. In all three engines, including Playwright's WebKit:

    chromium  clean  NOT REPRODUCED — all 3 frames arrived, including the tail.
    firefox   clean  NOT REPRODUCED — all 3 frames arrived, including the tail.
    webkit    clean  NOT REPRODUCED — all 3 frames arrived, including the tail.
    

    So "a per-delivery tail buffer, flushed only by more bytes" is wrong. The keep-alive experiment that appeared to confirm it was a single observation, and I read a mechanism into it that a controlled test does not support.

    Two further observations that also point away from a clean buffering story — the failure is not stable in shape. Two consecutive runs of SMOKE_BROWSER=webkit npm run smoke:web:app:

    run 1: FAILED — app never reached data-app-status="ready" (last: "loading")
    run 2: FAILED — locator.waitFor: Timeout 45000ms exceeded.
    

    Different stages: run 1 got as far as the App render, run 2 failed earlier, at the connect wait. A deterministic buffering bug would fail the same way every time.

    What still stands

    • The failure is real and reproducible. SMOKE_BROWSER=webkit npm run smoke:web:app fails consistently; smoke:web:browser passes. That is why webkit is out of the CI matrix, and that reason is unaffected.
    • The parseSSE cross-chunk defect stands, and is the one thing here I am confident about — it is read directly off the code, not inferred from an experiment. currentEvent / currentData are reset inside the read loop, so a frame split across a chunk boundary is dropped. Engine-independent, and worth fixing on its own.

    What does not

    • The mechanism. Unknown. Not a tail buffer.
    • The keep-alive as a fix. It was proposed on the strength of the retracted mechanism; there is now no reason to think an interval-based keep-alive is the right answer, and its latency cost should not be paid for a theory that does not hold.

    Next step for whoever picks this up

    Do not start from the diagnosis section — start from the varying failure stage. Worth ruling out first: connection/socket limits (the Inspector holds a long-lived SSE stream plus the sandbox-proxy and app-origin ports open simultaneously), and whether the earlier-stage failure in run 2 reproduces independently of the App path at all.

    The isolated probe above is worth keeping as a negative control — it establishes that plain SSE streaming over fetch is fine in all three engines, so whatever this is, it is not that.

  7. added a commit that references this issue on Aug 26, 2026
  8. added a commit that references this issue on Aug 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingv2Issues and PRs for v2

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions