Skip to content

feat: add annotation image attachments and harden feedback delivery - #4

Closed
Julian-Dasilva wants to merge 24 commits into
mainfrom
fm/lavish-229-v2
Closed

feat: add annotation image attachments and harden feedback delivery#4
Julian-Dasilva wants to merge 24 commits into
mainfrom
fm/lavish-229-v2

Conversation

@Julian-Dasilva

Copy link
Copy Markdown
Owner

What Changed

  • Image attachments on annotations. New src/attachment-store.js and src/async-mutex.js add content-addressed attachment storage with upload admission, disk/object caps, and a reference-aware sweep; src/server.js gains the upload/delete/serve routes, src/chrome-client.js mediates uploads (rate, in-flight, and cumulative-byte limits), and src/artifact-sdk.js renders the annotation card's chip UI. Every SessionStore mutation now serializes through one AsyncMutex, and queuePrompts re-derives attachment metadata from disk all-or-nothing.
  • More accurate annotation and diagram targets. New src/table-cell.js attaches semantic row/column names to table-cell annotations and stays silent when merged cells make a name unprovable; src/whiteboard-core.js restores Mermaid label line breaks (with bound-text recentering) and routes whiteboard init through resolveWhiteboardInitAction so an unmodified scene re-converts instead of prompting; new src/self-paint.js returns a fail-open self_paint_warning from open/export/share.
  • Server hardening and feedback-loss fixes. Mutating routes now 403 a foreign Origin/Referer, the artifact asset route is confined by realpath, and a poll that disconnects mid-take re-queues its batch via restoreClosedFeedback instead of dropping it. Docs follow: README, AGENTS.md, a new author-approved VISION.md, and removal of the committed .no-mistakes/evidence tree (now gitignored).

Risk Assessment

⚠️ Medium: The branch is a large upstream sync (~9k lines) whose bulk is already-reviewed upstream work, but it introduces one crash on a documented reverse-proxy deployment (attachment routes calling isSameOriginRequest with a stale arity) and leaves the exact feedback-loss failure its head commit claims to close still reachable through the long-poll respond() path.

Testing

Ran the three node:test files that own this behavior (server, chrome-client queue, session store) — all green — and then proved the intent at the product level rather than resting on unit tests. I drove the real CLI, a real local Lavish server, and real Chrome through the reviewer flow on this build and on a scratch checkout of the pre-fix commit. With the agent in the "Working…" state, the fixed build keeps Send to Agent and Send & End enabled and delivers the message queued during that window on the next poll; the pre-fix build greys both out and swallows the click, leaving the session open with nothing queued. A 10-trial interrupted-poll probe lost every batch before the fix and none after. The final send-and-end batch also arrives with session_ended/ended_by and releases SSE presence back to waiting. Visual evidence is a side-by-side screenshot with a zoom on the composer buttons; no test failures or flakiness surfaced.

  • Evidence: Before/after: reviewer send controls while the agent is working (zoomed on composer actions) (local file: /var/folders/f0/ts_nzbk17j7d63nm7wbjkqjh0000gn/T/no-mistakes-evidence/01M0ESZ9GD1CGAZKWEYZCDMVXQ/11-send-availability-before-after.png)
  • Evidence: AFTER fix — Chrome shows "Working…" and both send controls are live (local file: /var/folders/f0/ts_nzbk17j7d63nm7wbjkqjh0000gn/T/no-mistakes-evidence/01M0ESZ9GD1CGAZKWEYZCDMVXQ/04-presence-working-send-enabled.png)
  • Evidence: AFTER fix — second message accepted and shown in the conversation while the agent works (local file: /var/folders/f0/ts_nzbk17j7d63nm7wbjkqjh0000gn/T/no-mistakes-evidence/01M0ESZ9GD1CGAZKWEYZCDMVXQ/05-second-send-accepted-while-working.png)
  • Evidence: BEFORE fix (ec50b1e) — Send to Agent and Send & End greyed out in the same state (local file: /var/folders/f0/ts_nzbk17j7d63nm7wbjkqjh0000gn/T/no-mistakes-evidence/01M0ESZ9GD1CGAZKWEYZCDMVXQ/10-BEFORE-fix-sends-blocked-while-working.png)
Evidence: Rendered before/after comparison page (source for the PNG above)
<!doctype html><html><head><meta charset="utf-8"><title>Lavish #229 - sends during agent work</title>
<style>
body{margin:0;background:#101319;color:#e8ecf2;font:14px/1.5 -apple-system,system-ui,sans-serif}
h1{font-size:19px;margin:20px 24px 4px} p.sub{margin:0 24px 18px;color:#98a2b3;max-width:1300px}
.row{display:grid;grid-template-columns:1fr 1fr;gap:20px;padding:0 24px 28px}
figure{margin:0} figcaption{margin:0 0 8px;font-weight:600;font-size:13px}
.bad figcaption{color:#f87171} .good figcaption{color:#4ade80}
img{width:100%;border:1px solid #2a3040;border-radius:8px;display:block}
.zoom{margin-top:12px;height:130px;border:1px solid #2a3040;border-radius:8px;
      background-repeat:no-repeat;background-size:3200px auto;background-position:-2530px -2000px}
.zoomlabel{margin-top:6px;font-size:12px;color:#98a2b3}
</style></head><body>
<h1>Lavish issue #229 &mdash; reviewer sends while the agent is working</h1>
<p class="sub">Same artifact, same moment: the reviewer sent one message, the agent&rsquo;s poll took it, and presence is <em>Working&hellip;</em>. The strip under each shot zooms on the composer actions.</p>
<div class="row">
<figure class="bad"><figcaption>BEFORE (ec50b1e) &mdash; Send to Agent and Send &amp; End are disabled</figcaption>
<img src="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAABQAAAANcCAIAAACPGpUWAAAQAElEQVR4AeydBVxUTRfGpUtBBRFbUbFFLOzu7u7u7njtLuzu7u5uBRU7UDGxBZXu91lGr9eldul43t/5xjNnzpyZ+d/49tzZvWimz5iDQgIkQAIkQAIkQAIkQAIkQAIkQAJJnoBmimT9HxdPAiRAAiRAAiRAAiRAAiRAAiSQXAgwAU4uRzqsddJGAiRAAiRAAiRAAiRAAiRAAsmIABPgZHSwudR/CbBGAiRAAiRAAiRAAiRAAiSQvAgwAU5ex5urJYE/BPgvCZAACZAACZAACZAACSQ7AkyAk90h54JJgARSpCADEiABEiABEiABEiCB5EiACXByPOpcMwmQQPImwNWTAAmQAAmQAAmQQDIlwAQ4mR54LpsESIAEkisBrpsESIAESIAESCD5EmACnHyPPVdOAiRAAiSQ/AhwxSRAAiRAAiSQrAkwAU7Wh5+LJwESIAESIIHkRIBrJQESIAESSO4EmAAn9zOA6ycBEiABEiABEkgeBLhKEiABEiCBFEyAeRKQAAmQAAmQAAmQAAkkeQJcIAmQQDwT0NXVT5kqdZq05qbpMkRB0BHdESSay2ACHE2A7E4CJEACJEACJEACJEACCZwAp0cC8UlAW1vHOLWpgaFRQIDfr5+u379+jIKgI7ojCEIhYJTXwwQ4yujYkQRIgARIgARIgARIgARIIOET4AzjkwD2bE3SmPn5ev/88d3H2yswMCBqs0FHdEcQhEJAhI1aHCbAUePGXiRAAiRAAiRAAiRAAiRAAiSQ8AnE5wyxVZvKJA02b5G7xtQ8EAoBERbBoxCTCXAUoLELCZAACZAACZAACZAACZAACZBAJAQMUxp7evz09/ONxE/NZgREWARXoZ+yCxNgZSKskwAJkAAJkAAJkAAJkAAJkAAJRJOArq6+RooU2LCNZpwwuyMsgmOIMFsjMCazBDgCEmwiARIgARIgARIgARIgARIgARKIIQK6evq+vt4xFCyMMAiOIcJoiNDEBDhCPEmskcshARIgARIgARIgARIgARJQh0DTKumXjcx3dU2Jl/vLqSXogo7ors5oScpXR0fX388v9paE4BhC3fgJOgFOb24W6XqyZ8tcrUr5wQN61K1VNVJnOiRrAlw8CZAACZAACZAACZAACahMALkrktjZ/XPXLGVqYaqncr/fjuiCjuiOIAj125qc/tHU0gqM6jufVeGE4BhCFU+5T8JNgJs0rH357P4+PTvKpxta79S+5eL5Uwf27eobm08XQo9LCwkkMgKcLgmQAAmQAAmQAAmQgMoEJnS3RO6KJFblHinuPPk+foljlW4nc9bZC4GCKowIglAIqHooesYegSgmwP+NGTx25IAhA3sMH9x70vhh5uki36pVdw15rHLp6upWLFcq4o4Tp857/NQJPo73HqKkkAAJkEAYBGgiARIgARIgARIgAZUJIFntUCejyu4KR+S6zYde2HbM+c0HD0U9RQooqMKIJlgQEGGhUMIjYGxiAgmvNabsUUyADxw6Ubyo9YA+XevUrnr46Olv311jakJSnHkLV86Ys7hbn+GSJTwlc6aM710+fv/uFp4D7SRAAiSQrAlw8SRAAiRAAiRAAioTaFolPZJVld0Vjp3GXUGuq9DC+h+a4IAWhEVwKJTQBAoVtpm/cCUESujWGLREMQG+//DJvQePMY9Va7bcunMvKCgIesyKn5/fyjVb3N1/P0EJL3jKlEbpzc3EZMLzoZ0ESIAESCD5EuDKSYAESIAESEAdAkPaZFXHPQU2eC/f+RxxFzjADT4qBtfV1bWyygP/5COWOXM52F+HQInVVUcxAcac8ljlRInsF2U8Sp7cimk8fPQsHufAoUmABEiABEgggRLgtEiABEiABNQhgB1aC3Xed3XnyXds8KoyAtzgjOAYIlL/eg0arV63KVduq0g91XKwzJmze88+EXepVLlKzVp1IvaJjdYL58+YmppBoMRGfClm1BPgvHly+fj4vnR+I8UKT8mWNXOhAnlzZM/asH5N4aOtrV27ZuVUqVKKqijr1qqaKaOF0C1zZKtauZyGhoaoijK1iXGVSmXLl7PFrq+woLTKbYny4eOnKPX19apXrWBinAo6hQRIgARIgARIILkT4PpJgARIQE0CVUukVavH/rNvVfcXzqoMUat23devX8V4Impunh75bcQTzpuvQOHC1hH7xEbrzx9u06eMh0CJjfhSzCgmwOnNzdKkNnnm9CLiLz/36tb+4J71C2ZP3L9r7fmTu4sWKYSB9fR01yyfs3zRzFFD+6IqpHOHlksXTm/XuommpmbvHh1PHdm2dsW8kiVsRCvKpo3rXji9d9rEkZvXLjp1eJuRoQGMEJEAP3rshOT5wK51q5fNwUBIsNFEIQESIAESIAESIIHkS4ArJwESUJ+Ade5/tugiDXD17pdIfSQH4RzpENmz5zAzM/tv3Kjq1WtqaWmJ7rmt8ixbsaZhoya79h5cs25T7br1hH3s+ImNGjebv3DJoaMnR40eb26eXtitrPLMmbfwyPEzS5avKle+AowVK1UeNea/DBky7tpzsGy58rAUti6ycvX6E6cvrN+4rXSZsrD07jugQcPGVavXhE+aNGlgwbgizrwFixETltiTajVqQ2IvvogcxQQ4b57c6P/k2QuU4cnIoX0qlC/VpkOfJq26jRw3HW5Pnj5HWalCmROnLmzbeaBG9YqoCvny5RsUh9v3kNB6enoePnYa1QzpzVFCsOs7Z/q49p37l67UYOioyXcfPPbx9YUdkjdPro+fvpimTb1907KtO/YfPXE2bdo0RYsURBOFBEiABEiABEiABEggmRLgskkgSgQs1Pn+M0Z48+edz9AjFeEc6RA1a9c5d/bMm9evv3z9UtL299/E0dHRyZsvv4lJ6s4d227csG7osJFGRkYYMXXqNE2bNV9sN79b5w6mZqYNGzeFMU3atLPn2l2+dKFls4abNqwfPfY/9L186eK8OTM/ffzYoV3L69euYt+xe4/eu3Zua9yg9vZtmydNmYEh1qxafvTIoQvnzsLnx48fCD5j5txzZ08jzpkzJ2fMmmdsbIz4sSFIfW1LlYVAiY34UswoJsD58uRCCKfnzijDlO5d2jZvUr/vwDGeXt5wyJTBAuXjJ4q/V3Ty9IUduw/aOzimMzM1+fN15Tx5cj5/8er8xWtPn73YtHXP16+u8H/15h1KCI4Kym+ubij37j/aZ8DowMDfr93Klze3r6/fsoUzBgwdv3nbHvFjYGwywzMC0Vf8Z6CvrxBdXV3JU1/xH4z6kkVSdHR0IVJVP6SvVOJ0kZqgaGhoZMmSVU9PD7pc9BX/Ib4BHOR26iRAAiRAAiRAAiRAAiQQYwQYKNESQF5ao0btM6dPYgVnT5+Sfwva399/y+YNnh4ely9deP/uXbHiJeEDOXni2KtXzl++fD5y+GCZMuVgKVasxNu3bw4d3I+dRfub10+dOI7t36CgIH9/v+AUwT4+PtAh/fv2PHvmtLe3NwIic8mcJau/v3+g4r8A+AQHBxcrXvzdu7cnjh/18vI6cewoEvKCheLh29FYUQxKFBPgvHkVCfDzl6/CnIqxcarBA7rv3HPI7cdP4ZArZ3akrE4v/ibMHz59RlO2rJlRamtp4aHCvIUrQRlVCPz9/PxEwozqS+fXOBXmz5qgr/9PSpnBwjy1ibGWlmb3PsPu3nsEz0wZFZv+b966QI9ANm3duWLV2inTZkBat2kveQr71Omz5y9cgiclGTNmkpqGjxw9asw4UUXOjI6Qw8dOiTi1atcVTUiS8ShlzbpNPXr1mW+3ZPrMuRYZMogmlIi/dMXq/yZOWbp89djxEw0NDWGkkAAJkAAJkAAJkAAJkAAJxBiBaAT69P3390xVjJEtoxpfmRbOEQ9RtFhxs3TpGjRsPGLU2OIlSlauUi1lyt9vOPr58weyVjExV9fvBga/fxP63fX7b+N3V4OQH4pmy54diaswonz3/l22bDmgKEnzFq02bNq298DRVWs3oklT458XMMGSPYcl5rNzzwEh+fMXSGf++yu6aI1ZOXPq+M0bVyFQYjayUrSoJsAhX4F+/vxvQiuP26p5A0MDg70HjknGvHlyvXr91sfn7/n08aMiAU5vbgaf5k3rv3f5cOLUeegQDQ2N4kULP3j0FE8gUIWsWrv1yjX70rbF1q9agMiwCEFYKHPtVrx+8x4KpFCBvMi6376LJAGG5549u4YPHQTZuGEtqpLAPmzIgKGD+uvrGyBTlexyBck5OkLc3FzhD+XwoQPCoWq16g0aNYFl/NhR/fr00NDU6DdgsGgS5f59e8aMGjZkUH+czcVL2AojSxIgARIgARIgARIgARIggXgncO+5h1pzKFtEjYRQOEc8BPbVPnxwuet4B3L61AlMpnKVqijVkBQpPri4pE+v+Aau6AXdxeV3uiQsKPMXKNi5aw/kLE0b1e3YrpWPjzeMSoKZXLt2pUXThkIqlrM9uH+vkk8MVs+cOg6JwYBhhopKAqytrZ3LMpunp9enz1/DDFq+bKmAgIDXf77AjP1SS8ts2MWVO3//7oaqkZERNnUH9O06edoCVIXkyW2JPWT7W3dFFSUy4R59Rly+qsiB166cp6f3+0vL+fMp3gz+7M9PkTGxfHlz33/wGF0ildKly3bo2AViXeTvq7akXtiLPnf2VA7LnNh5loyqKLlyWz17+gSPZOCMIDeuX7PKHcaf8CpatJiOjg5OKbhRSIAESIAESIAESIAESIAEEgKBsw6KX2KqPpPGVbOq6xzBENjULV+h0qgRQ06dPC5k3pyZNWur/UeJ7ty5lS9//pK2pTE37AbXrlP3xrWr0L99/Wpunj516jTQ9fX1AwMCfvxQJGWNGjfV1zeAEfLt6xds/CJVge5453bBAoWKFSsB3dTUbPbcBblyKV4FhWqMi0nqNGPGT4FAiVbwyDprRuYQRnsuy2xINZ1fK7/yO0e2LDmyK86ArFkyfvnyTWzQw3P+7AnaWlrSJq2I6OOr2A1OmdKoc/uWV6/Z35NlrSWKKzLSO3fuC09Renl7d+8z7Nade9gHrl61ojAiAUam/cL5959iyp83N5Ltu/dVSoDTmppmyZoVYmKSWkRTKr29fXDgtbS0lewRVw0NDeWPT3y8vXEey7sMHTby7IWrM2bP27Z104vnih9Fy1upkwAJkAAJkAAJkAAJkAAJxBeBvec+R/wVZaWJFc1n2qaO4s+yKtlDV+EGZwTHEKFbhaVS5arv37978/q1qKK8cOFc3rz5M2bKDF11+fzp08T/xvYfMOjwsdPz5i9eu2bV7dsO6P7ixfNrV69s27m3UuUq9+463rx5fe+Bo/sPHUuZMtXPnz/gAMGImlpaR46fTpvW9NPHj1MmTxg0dPihoyc3b9354MF9RIBPbEilytW+f/8GgRIb8aWYUUmAixYtjP7Of9JO6EKGDOyZJXNG6GlSm2hpK/JGAwP9tSvmBgYEwvjl79RhWQAAEABJREFU6zckw5Y5skGHID328/PLlMmiXZumM+cthUWSggUUW6aPnz7Pni0zMmfswYqvPfv4+NotXgM3AwM9lBAkwMh+kQNDh6CK0vnVG+TV4o8MoxqeHD1yaNqUiZBLF39/9VrJs0CBgm/fvvH391OyR1zFOZEzV27MWbhZ5ckLi9BFuWrF0s4d2yJJfvvmd94u7CxJgARIgARIgARIgARIgATincD8bcr7fBFPaUo/m/JFFe8hisANDnCDQ8TBjx870rVTO7hJ8uvnz2qVy31wef/40cMWTRtK9iGD+p88cQzV4UMHHj96BArkwYN7rZo3hgJxsL/Zvm3LVs0bNWtSX/695QnjR9evU/3ihfOBgYFTJ09oULdmq+aNN21cB+XlS8Wf+Pn29WvvHl3q1qomvtNqf/N6+zYt2rZqVq9O9c0b1yNyLInzyxclSpaGQImlIURYtRNgLS3Npo0Uu/BBwUEF8udBjpota2abIgXHjhxQu1YV+1uOiPvsuXN6c7NxowaeOLT1yPEzm7bugbFCOduNa+zyWP19QOLt49u+ddONW3Z9+/bPNw3SplVsyjdrXHf+rInIY8ePHrRr60rzdGbIn2vVqPTu/YdTpy8iIJqyZ838QvYiLlNTRUfrQvmXLZyezswUPhFI2jRps2TJCjEzU/wOWfKEPUcOy4aNm7Zt3+HEsd8nk9QaqXLzxrU0adL26NnH2MSkTNnyDRs1vnD+rLyXu4fH2zevDx862L5DZylPljvEoM5QJEACJEACJEACJEACJEACahHADu2mYx/U6rJhajls8IbXBU1wQCvCIjiUOBNPT8/QYyH1DQ4OFnZsy/mGfDNXVKVS2mIUFnd3d+xfCj2Wygf3HbGfCoESS0OIsOolwMWLWh/YvT5blszvXT7aWBdCnrll/RJkp+tWzm/XuqmT00ts0iLu/kPH/f39q1WpMOa/mbv3Hnn4+Kmr249iNoVXrd16/OTf7VZA/PL127qNO9FFLrdv30O1ds0qvQeMcvvxE2MVyG917eIhh6vH8ljl6tBlwM9f7nAoXDCfhoaG9KZoWBzvPkRZt3a12fOXYR8YegTStXvPLdt3QwYNHSF3g339pm2NmzTbtmXzju1b5U2q6C7v30+ZNL5KteqHj54aP2Hy/r17Dh7YF7rjju1bzNOnr1a9ZugmWmKKAOOQAAmQAAmQAAmQAAmQQBQITFrtjGRVrY5T+tnsnlcJuW62P++FhoIqjGhCKAREWCiU8AhguxsSXmtM2dVLgG/duVe/ScfiZWuXq9KoUo2mFas3hWJbvq6NbY18RSrWafT77wlt27G/VMX6lWs2u3pd8V1zX1+/CtWawO3i5evyeY8YPaV5255

... [378651 bytes truncated] ...

IECCQqQLz585u2Lh5ps7evAkQqFEjkX4vuOCCuJfjji4vSToE4PLOWXsCBAgQIECAAAECyQgsWbJo4YK5zVq0TuZi1xAgUN0Chen3gYeejHs57ujyzkgALq9YpbfXIQECBAgQIECAQNUJzJo+tXnLNnX9TeCqIzcSgcoRKEy/d9x576QpM+JeTqJfATgJNJdUooCuCBAgQIAAAQJVKhCPjKZNmdCqTcd8Pw66SuENRqBCAoXp97b/3jl1xtx6+YuXLFmURI8CcBJoLiFQWQL6IUCAAAECBKpBYO7sGfPmzm7TrqvnwNWgb0gC5RcoTL/x7HfG7IWN6uefffZZ5e+m4AoBuEDBLwIEqkPAmAQIECBAoNoEZs2YMnvW9LYde/r7wNW2BwYmUGaBRQvnX3DBBQ889OSCpbUnjhtx9lm/veXW28t89c8aCsA/4/CGAAECVSVgHAIECBCoZoF4DjxhzPA69ep36LrhWmuv26BhE98UXc1bYngCxQTirox7s3Xbzv974Y1xEyfHPZufXzvS7/AR3xdrW6YKAbhMTBoRIECAQKUK6IwAAQJpIbBkyaKpk8dPHj9q+fIVTZuvvW7Hnp17bKoQIJA+AnFXxr2ZX6vOoPfeibs17tma+bWSTr/xuSMAB4JCgAABAgSqUsBYBAikl0D8kXrWjCmTJ44eN2ro6BFfKwQIpI9A3JVxb8YdGvdppXxwCMCVwqgTAgQIECBAoIwCmhEgQIAAgWoTEICrjd7ABAgQIECAQO4JWDEBAgQIVKeAAFyd+sYmQIAAAQIECOSSgLUSIECgmgUE4GreAMMTIECAAAECBAjkhoBVEiBQ/QICcPXvgRkQIECAAAECBAgQyHYB6yOQFgICcFpsg0kQIECAAAECBAgQIJC9AlaWLgICcLrshHkQIECAAAECBAgQIEAgGwXSaE0CcBpthqkQIECAAAECBAgQIJDFAkfv0/W+K3b86ulDpr97nFJ2gRALt9Cr+O+N6gjAFZ+1HggQIECAAAECBAgQIJA5Avvv3DFS3E2XbLvfjh3art0gcyaeFjMNsXALvTAMyYrMSQCuiF5S17qIAAECBAgQIECAAIFcEjjvuA3vvXyHSHG5tOiUrDUMQ/LPp22WdO8CcNJ0LkxKwEUECBAgQIAAAQIEckkg0lqUXFpxytcaX1BImlQATvn2GIDATwKOCBAgQIAAAQIEcklg1z5tI63l0oqraK2hmtz3QgvAVbRDhiFAoAYCAgQIECBAgECOCdxw0TY5tuKqW+4/ztoyicEE4CTQXEKAAIHyC7iCAAECBAgQyDGBo/fp2tbPu0rZpodtCJe3ewG4vGLaEyBAgED5BVxBgAABAgRyT2DPvu1yb9FVuuIkhAXgKt0hgxEgQIBALgpYMwECBAjkpMAWvVrm5LqrbtFJCAvAVbc9RiJAgAABArkoYM0ECBDIVYG2vv85xVufhLAAnOI90T0BAgQIECCQywLWToAAAQLpJCAAp9NumAsBAgQIECBAIJsErCXtBbZuudYj220d5dz1uqf9ZE2QQCUICMCVgKgLAgQIECBAgAABAqsLpP37CL3nrNd90NRpUT6YOr0i862Xn79B0ybN69SpSCdFr9277Tp/WL9n2/r1i1YWHresW7dOzVQFmVO7d7mgV4/Csar4YPzkOUNHTl+2fEWljDtr7qJ/3Pb+o/2+rZTesqOTVP2+yQ4dqyBAgAABAgQIECCQrQJ9Wq41aOq064d8F+WDqdOSW2bd/Jr/2mzjr/bd44Wdt/tkn92e26lvx4YNCrqq2K9d11n7zB5d29Svt1o3PZo0fm3XHQbvvWuM+McN1lvtbOLtdVtsMvLAfYqW47t0Spwqy+uvO3c8vXvXsrSs3Dbvfzp+84Pu2eLge3c89qGuu91280OfVLz/2XMW3fjAR/97dehqXb05aMw1d384YsyM1erX/HbWnEXtd7plnb7/+dP1b6+5ZXJn73/mq5jV4iXLkru8jFcJwGWE0owAAQIECBAgQIBATggkvi+6jEuNx8iHd2z3+qQfzvv4s1uGjVi/aZPb+mxRxmuTaHbR+j2b1qm91xvv3DpsRDyqjeFK66T/xMn3jhiVKN/Oml1as1TVl7Pf2XMXH37es/MWLLns7O1vuHS3Hp1aXH7LwBcHjChnN2Vt/uag0dfcNei70eULwC+8NXzJynT67OvfVdYz6qIzLgjAdw1auEgALqrimAABAgQIECBAgACBVArE0+B4MvzIdluXZZC+rQr+pZ8rvvr2mbETrv5m6O8/+fyFcRPq5+fHtc3q1L52i03e33OXt3bf6bKNN0h80/L2a7d8Y7cdz+jR9d9bbBJPjG/davMNmjaJxlH2arvOUztsM2ivXePRbl5eXtQUL/l5eUuXrxg1b97EhQvj7LIVpX6r8EMjx1z25TeJ8uG0gm/wvmvrLWPonVu3emXXHWJKv+vZLXpIlLN6dosHy1G5f7u2iZoqfv3wiwmRLffo2/m0IzY9Yp9e91y5zx9P26Z+vdqJaTzdf9j+ZzzZc8/bDzvnmXc+GpuoPO7C5/se+cDr74+KJ8bbHHH/9fcNTtRHij7nitc2O+iefU554sthUxKVRV//dvPAx1Z+U/TF17z128v6x6lQjDy86/GPrLf3HUdf8NxHX02KyuLlqVeG1q9X6/cnbTVl+vzCaUSza+8ZvP0xD8Uc/vfqsAPOfGqX4x+pEbU1asycvfCsy1+NmcSpS64dkHi0+9aHY2La/3ng47P//ur6+9x58iUvJiZ54G+fHjaqYJv2OOnRu578fGUHKXnxBAqZhJEAAA9GSURBVDglrDolQIAAAQIECBAgkEECtWvWrFukxPPVj6ZNf2z7X87AQ2bPiWXeufWWp3XvskWL5s+Nm3jzsBELli2LpHrvNr33XbfNI6PHvjRh4nFdOl6z+cbRsmGtWp0bNfxdj24rVqyIayP0nrfyL9yu16TxDVtuumGzpu9NmbZVyxZ7tVknGhcvV349pEmd2hGS/77JhtcP+W7oytGLN4uaeDjct9VaiRKDRs26DerH0Of36hFLa1K71gW9esRwUX90pw5RGQ+WB02dHget6tWNyiouPTq3yK+Z93T/oedd+Xo8X61TO/+cX2+5c58OMY2X3/n+d5f3z6uRd+EpfSJSHnfhC4nEOG7SnBFjZ/7zjg/6bNJ25pxF/7z9gy+HFsTdC/75xmMvftuscd1uHZtH7IweViubr9+6S4dmUdl74zZ9t2gXB5fd9O41d3/YvGm9Ew7a6NNvJx91/rOJLBqnCsvEKXM/+HzCzn06HrR7wV+QjkyeOBVh9ao7P5i/YMlOfTpGCP/0m8nfj50Zp+IR8VEXPPfM698du/8Gv9qp2z1Pf3H231+L+sjnMe1oWaNGXq+ua/UbMOLqOwfVqFFjj76dmjYukN9nx649O68VNSkqVRGAUzR13RIgQIAAAQIECBAgUCkCA/fcecj+exUt8YB0q7VaPNi3z5r7/9sX37w0YVLnhg0v3mC9Jwue3+5ySIeCTNWjSeNNmjd7d8q0Z8aOf2TU2CGzZu/Xrm39/IInw9HhG5N/+P0nXxw38MPZS5Zs16plrZp5u66zdjwivnHId+d9/Nlh77w/eeUD3mi5WokHyA3y85vWrh3R9z9Dh5+zXveRB+7TuHat1ZrF24s26BmTT5RORf5a8rkff37pZ1/dMqzgu4t3at0qWu7dtiBsnz7ok4s+/eLYgYNqp+zHa8VYpZUObZrcdtlebVs3fuSFb077y8ubHHD30Rc8N2NWwVPuSLPLl6/461nb7bZtp/NP3GrhoqUvvDm8sJ9b/2/Pq/6wc6TlqHnt/VERO196e0RE2edvO/TGP+12yenbRv1qZd+du/XeqE1UHrJHz6P3XT8e/z7S75u1mtV/6Jr9Lz516z+dse2ceYtfeLPAJ9oUlni6G9PYa/suXTs0j2j94oARMZM4+/zKlvdfte+V5+/42PUHLlm66huYh3w/LcLwDlu2P3TPnscdsMH6XVs+89qwBQuXxiVRYi0xvcdvOLBpo7oDBo9ZsnT5mUdv3nqthnHq3ON7b7cylsdxKooAnArVn/XpDQECBAgQIECAAIE0F9julTfXe+7louWmocM/nDY9AuGaZz536dIzP/yk98uvnfLBx4+NHtu8Tp140tuuQf14ohsX7ty61YDdd4rSa+X3OXdv0igqo4ydNz9el61YMX7+grr5NWvn1ezUqCD8fDqj4OHh0uUrvpw5KxqsVk7q2vnSDXvFQ+CzBn8a/f9zs426Nmo4ddGiOUtWxaqi7SPl7vXGO4kyfO7cwlOJocesnEDjWrWiPoZeUaPGZyuHHjd/wZSFi6Ky6st+u3Qb/OTxr9595N/O3r5bh+ZvfDD6rze9G9P4Zvi0eN37lMf7HHb/iX/sF8dffTc1XhOlQ9smcdCxbdN4jeA68Ye5ixYv69GpRaMGBT+Re4sNC7J9nFpDiUe7s+Ys6tqhWd06+dEskmq8fjPipyHibZSnXin4SVqvvTfq4mveGj56xtz5i18dOCrqx0yYlZdXo2fnFnG8TsuGLZqu+sHd3wwv6OH190fFtKN8vfLt0JEFa4mWiWnHQ+926zSOCS9dujwqq6YIwFXjnLujWDkBAgQIECBAgED6CyxevnxRkXJGj65brtXiiHc+WPPMa+bl/bZH18ilMxcveW3S5Is//XLA5ILvwl23Qf1R8+bFtbcOG9H5mRcLyxczSoi10SxKJM947dpoVULu2njVQVQWlt3arL1k+fK7ho98YfzESz//6rAO7eKp8js/FAStwjaFB9FhPCVOlEXL1pSvxs2fn1ejRpeVCbxZndpr1S2IjoX9VM3BoM8n/PueD78cNmWjnq1OPWLTm/6ye4w7duLseO3SvmkExe9fP33SwLMS5aFr9ov6Ekvrlg1r16o5evyseKYaDSKpxmtpZdnKf2wpLmlQr/bYSXPiAW+0HJMYtF3B90jH20T5btT0RIIdMWbG4C8nJiqf6l8QidfrslY8Q37no3FRGfOfNnNBHETp0r6gh7OO2yIx58Trpr1ax6k1l+XL17RZa762LGcF4LIoaUMgSQGXESBAgAABAgQyTmDrlmv1abnWUe/+QvqNdS1fsSJy8p836vXPzTbav13b3/Xstk2rteYtXfrp9Jlfz5w9bv6CYzt3PLj9unu0af3yLtu/vcdO9fILnjHGhcXLwClTo7ffr9/j3PW639R7s15NGhdvM3TWnNo1a0abjZo1nffjt9rmx/PH4k1r1Di6c/s/bdQrUXqvVfB8sqRWBXWJCH3zVptHmL93m94Rhgtqq/ZXndr5V9856MQ/9rv7qS+eeHnIZTcNjPF3Wvl3gPfZsWsk1XOveP2tD8f87eaBPfb87+2PfRZnSyyRfvtu3m7S1HnHX/TCtfcM/uO/3yqxWed1C54Y3/XE5/0Hjsyvmbfn9p3j0fH5/3z9sRe/veLW92rl19xju85FL3z61WHx9vJzdnj9vqOiTHz3rLXXavD6B6Pj0fEph28aPZxw8Qv7nf7kCRf3a1h/1Q/u2qhHq/Ztmtz3vy8ff2nIS29/v/OvH97q0PsS3zUdXZVYOrUrmNUVt72f+MvMJbapeKUAXHFDPRAgULKAWgIECBAgQCATBT6YOq0s6TextAs//SKe+h7esf0NW256Qa8eExcsPPLdQfE8OcoxAwd9N2fOVZtv/N8+W0RMPe+jzxcuW/UXRBPXFn39aNqMy7/6dsnyFees171+fn6/8aseMxZtc+2QYS9PmHRKty7P7dT3X5tt9MrESf8bOz6C90Ub9CzaLHG8Z5t1Tu7aOVHWX/kN2In64q+3fff9k2PGtWtQ/7xePd6YPCWWULxNqms2W7/1jX/abc68xZdcO+Csy1/96KuJpx2x6ZlHbx7jHnfAhhedsvX7n40/8rxn73ji80P3XO+kQwt+nFicKrH858+799m4baTl2x//7HfHlvxPUh28R8/d+3b6+rup/320IEtff8luB+7W4/k3hp9zxWs1a+bd8tc9ttzwZ987nfiRV3ttvyoVx9cc9tyuy5Ily55/c/jOfTo8ceNBR++3QTwKfvDq/Zo0rlszTteoEZH+yRsP7NGpxflXvh7BfumyFTf/3x716hZ8z3mJ047KM47aLGYeIbz/eyPjbYqKAJwiWN0SIJDrAtZPgAABAgTSXOCGId/Fk9544pooyc12ysJFJ7w/eMPnX/nVm+9u/uJru7424Ksf//rumHnzD3n7/Y1f6B/1u7/+9sfTZ8QQkWA7P/PiVd8UfPdsvN3nzXfj7YKVwfjeEaN6v/Tahi/0P/mDj87+6LOoT1wSzRJlzpKlZ3z4yfrPv7J9/zej2emDPjn/48+j2b++XtVbotl5KyujvrDc933B31bd6413oiaSeTTrP3FyHF/59ZA4jifPf/jki037vRo93zjku+36v9n9uZeivorL4Xv3GvLSqR89dcLbDx0z4rXTLzt7+3gSm5jDeSf0/vL5kz9/7qSRr5/+j/N3TNS/ef/RkwaeFTkz2uy9Q5c4/stv+8ZxqxYNnr31kO/6n/ZNv9+cdMjGUf/Y9QdGfdHStHHdB67ab8jLpz71n4Oivm6d/Nsu23NYXPLibz54/Nf779I9KouWQU/8OvqJJ7qFlVdfuHPUHLv/Bs+98V2/t0YctlfPqJm3YEk8Sd6019qJZh3bNn3hv4cN63/qty+e8s5DxyR+8tavduwaF17640/niufJ8bZ+vYJgvMUG68TMR7955gUnbpXoIRWvAnAqVPVJgACBXBewfgIECBBIf4F40jto6rQ+K7/hOV4rMuH5y5Z9M2v2jMWLi3cSp0qsL94yalbUqDFvaQk/0SpOFZZIsOPmL4jUWlhTKQfRbZRK6SrpTuLRabt1GsdT00TEXa2f1ms1LLF+tWaJtw3r145nuYnj0l5juKKn8mvmFf4Iq6L1az7u0r5ZZOB9Tnmi1z537HvaExG//3jaNkUvaVCvdvOm9YrWrPl4tVmtuXESZwXgJNBcQoAAAQIE1iTgHAECBDJF4Poh3x317geJkilzNs+0Etiwe6t3Hzn2/qv2veg3Wz9y7QFxnHjSm1aTLDoZAbiohmMCBAgQIECgogKuJ0CAAIGEwIQfCv61p8RxFr82bVR3j76dTzh4o537dIjjqlxpEsICcFVukLEIECBAgACBLBewPAIECBQKfPxtyf9EU2EDBxUUSEJYAK6gucsJECBAgAABAgRWCfgPAQJFBV4ZOK7oW8eVLpCEsABc6bugQwIECBAgQIAAgVwUsGYCqwk8/OKIJL5Hd7VOvC1NIGxDuLSzpdULwKXJqCdAgAABAgQIECBAoKwC2pUocMl/PiqxXmXFBZKzFYArLq8HAgQIECBAgAABAgRyWqC0xT/35ujrHviqtLPqkxYI1bBN4nIBOAk0lxAgQIAAAQIECBAgQKBMApf/99NIa2VqmrGNqnji4RmqyQ0qACfn5ioCBAgQIECAAAECBAiUSSDS2gl/fntCbvyrSGUSSbZRGIZkeCbbQY1UBOCkJ+NCAgQIECBAgAABAgQIZKHAc2+O3vDgp373j/eeHzAmUlwWrjCVSwqxcAu9MAzJigwlAFdEr8RrVRIgQIAAAQIECBAgQKAEgYdfHHH8pQMixbXY7gGl7AIhFm6hV4JpOasE4HKCaf4LAk4TIECAAAECBAgQIEAgTQUE4DTdGNPKTAGzJkCAAAECBAgQIEAgfQUE4PTdGzMjkGkC5kuAAAECBAgQIEAgrQUE4LTeHpMjQCBzBMyUAAECBAgQIEAg3QUE4HTfIfMjQIBAJgiYIwECBAgQIEAgAwQE4AzYJFMkQIAAgfQWMDsCBAgQIEAgMwQE4MzYJ7MkQIAAAQLpKmBeBAgQIEAgYwQE4IzZKhMlQIAAAQIE0k/AjAgQIEAgkwQE4EzaLXMlQIAAAQIECKSTgLkQIEAgwwQE4AzbMNMlQIAAAQIECBBIDwGzIEAg8wQE4MzbMzMmQIAAAQIECBAgUN0CxieQkQL/DwAA///J4mZcAAAABklEQVQDAMLkeeHtPSXOAAAAAElFTkSuQmCC)"></div>
<div class="zoomlabel">send.disabled = false, sendAndEnd.disabled = false</div></figure>
</div></body></html>
Evidence: Agent poll receives the message the reviewer sent while presence was "working"

session: file: .../review-artifact.html status: feedback prompts[1]{uid,prompt,selector,tag,text}: "",Also drop the footnote line while you are in there.,"",message,Freeform message

[lavish-axi] Long-polling for user feedback on /private/var/folders/f0/ts_nzbk17j7d63nm7wbjkqjh0000gn/T/no-mistakes-evidence/01M0ESZ9GD1CGAZKWEYZCDMVXQ/review-artifact.html. This stays silent until the user sends feedback or ends the session - leave it running. Detected layout issues do NOT return this poll: they wait in the user's Layout issues inbox until the user queues them as ordinary feedback. If it gets killed or times out, re-run `lavish-axi poll /private/var/folders/f0/ts_nzbk17j7d63nm7wbjkqjh0000gn/T/no-mistakes-evidence/01M0ESZ9GD1CGAZKWEYZCDMVXQ/review-artifact.html` - queued feedback is never lost.
session:
  file: /private/var/folders/f0/ts_nzbk17j7d63nm7wbjkqjh0000gn/T/no-mistakes-evidence/01M0ESZ9GD1CGAZKWEYZCDMVXQ/review-artifact.html
  status: feedback
dom_snapshot: "uid=1 body \"Pricing page draft Review the two tiers below and tell me what to change. Starte\"\n  uid=2 main \"Pricing page draft Review the two tiers below and tell me what to change. Starte\"\n    uid=3 h1 \"Pricing page draft\"\n    uid=4 p \"Review the two tiers below and tell me what to change.\"\n    uid=5 div \"Starter $0 One project Community support Team $29 / seat Unlimited projects Prio\"\n      uid=6 section \"Starter $0 One project Community support\"\n        uid=7 h2 \"Starter\"\n        uid=8 div \"$0\"\n        uid=9 ul \"One project Community support\"\n          uid=10 li \"One project\"\n          uid=11 li \"Community support\"\n      uid=12 section \"Team $29 / seat Unlimited projects Priority support\"\n        uid=13 h2 \"Team\"\n        uid=14 div \"$29 / seat\"\n        uid=15 ul \"Unlimited projects Priority support\"\n          uid=16 li \"Unlimited projects\"\n          uid=17 li \"Priority support\"\n    uid=18 footer \"Draft 1 · prices not final\"\n  uid=19 script"
prompts[1]{uid,prompt,selector,tag,text}:
  "",Also drop the footnote line while you are in there.,"",message,Freeform message
Evidence: Final send-and-end batch plus SSE presence release

session: status: feedback session_ended: true ended_by: user prompts[1]{uid,prompt,selector,tag,text}: "",Looks good - ship it.,"",message,Freeform message --- presence on the SSE stream right after the final batch --- event: agent-presence data: {"state":"waiting"}

session:
  file: /private/var/folders/f0/ts_nzbk17j7d63nm7wbjkqjh0000gn/T/no-mistakes-evidence/01M0ESZ9GD1CGAZKWEYZCDMVXQ/review-artifact.html
  status: feedback
  session_ended: true
  ended_by: user
prompts[1]{uid,prompt,selector,tag,text}:
  "",Looks good - ship it.,"",message,Freeform message
event: agent-presence
data: {"state":"waiting"}
Evidence: Interrupted-poll durability, BEFORE fix — 10/10 feedback batches lost

{ "build": "before-fix (ec50b1e)", "trials": 10, "feedback_lost": 10, "results": [ { "trial": 0, "sent_by_reviewer": "trial-0: tighten the Starter card copy", "rerun_status": "waiting", "rerun_prompts": [], "feedback_kept": false }, ... ] }

{
  "build": "before-fix (ec50b1e)",
  "trials": 10,
  "feedback_lost": 10,
  "results": [
    {
      "trial": 0,
      "sent_by_reviewer": "trial-0: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 1,
      "sent_by_reviewer": "trial-1: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 2,
      "sent_by_reviewer": "trial-2: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 3,
      "sent_by_reviewer": "trial-3: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 4,
      "sent_by_reviewer": "trial-4: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 5,
      "sent_by_reviewer": "trial-5: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 6,
      "sent_by_reviewer": "trial-6: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 7,
      "sent_by_reviewer": "trial-7: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 8,
      "sent_by_reviewer": "trial-8: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    },
    {
      "trial": 9,
      "sent_by_reviewer": "trial-9: tighten the Starter card copy",
      "rerun_status": "waiting",
      "rerun_prompts": [],
      "feedback_kept": false
    }
  ]
}
Evidence: Interrupted-poll durability, AFTER fix — 0/10 lost

{ "build": "after-fix (9ea0104)", "trials": 10, "feedback_lost": 0, "results": [ { "trial": 0, "sent_by_reviewer": "trial-0: tighten the Starter card copy", "rerun_status": "feedback", "rerun_prompts": ["trial-0: tighten the Starter card copy"], "feedback_kept": true }, ... ] }

{
  "build": "after-fix (9ea0104)",
  "trials": 10,
  "feedback_lost": 0,
  "results": [
    {
      "trial": 0,
      "sent_by_reviewer": "trial-0: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-0: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 1,
      "sent_by_reviewer": "trial-1: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-1: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 2,
      "sent_by_reviewer": "trial-2: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-2: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 3,
      "sent_by_reviewer": "trial-3: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-3: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 4,
      "sent_by_reviewer": "trial-4: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-4: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 5,
      "sent_by_reviewer": "trial-5: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-5: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 6,
      "sent_by_reviewer": "trial-6: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-6: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 7,
      "sent_by_reviewer": "trial-7: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-7: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 8,
      "sent_by_reviewer": "trial-8: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-8: tighten the Starter card copy"
      ],
      "feedback_kept": true
    },
    {
      "trial": 9,
      "sent_by_reviewer": "trial-9: tighten the Starter card copy",
      "rerun_status": "feedback",
      "rerun_prompts": [
        "trial-9: tighten the Starter card copy"
      ],
      "feedback_kept": true
    }
  ]
}
Evidence: Reproducible probe script used for both builds
// Product-level probe for the contract Lavish prints on every poll:
// "If the poll gets killed or times out anyway, just re-run it - queued feedback is never lost."
//
// Each trial: the reviewer queues feedback in the browser chrome, an agent poll connects and its
// process/connection dies before the server has written the response, then a normal poll re-runs.
// The re-run must return that exact feedback. Nothing here reads source code; it drives the
// real HTTP surface of a real `lavish-axi server`.
import { connect } from "node:net";
import { setTimeout as delay } from "node:timers/promises";

const [, , base, key, file, label] = process.argv;
const port = Number(new URL(base).port);

async function queueFeedbackFromBrowser(text) {
  const res = await fetch(`${base}/api/${key}/prompts`, {
    method: "POST",
    headers: { "content-type": "application/json", origin: base },
    body: JSON.stringify({ domSnapshot: `uid=1 body "${text}"`, prompts: [{ prompt: text, tag: "message" }] }),
  });
  if (res.status !== 200) throw new Error(`queue failed: ${res.status}`);
}

// The agent's poll connection dies before the server can write its response.
async function pollThatDiesBeforeTheResponse() {
  const socket = await new Promise((resolve, reject) => {
    const client = connect(port, "127.0.0.1", () => {
      client.write(`GET /api/poll?file=${encodeURIComponent(file)} HTTP/1.1\r\nHost: 127.0.0.1:${port}\r\n\r\n`, () =>
        resolve(client),
      );
    });
    client.on("error", reject);
  });
  socket.on("error", () => {});
  socket.destroy();
  await delay(120);
}

async function rerunPoll() {
  const res = await fetch(`${base}/api/poll?file=${encodeURIComponent(file)}&timeoutMs=0`);
  return res.json();
}

const results = [];
let lost = 0;
for (let trial = 0; trial < 10; trial += 1) {
  const text = `trial-${trial}: tighten the Starter card copy`;
  await queueFeedbackFromBrowser(text);
  await pollThatDiesBeforeTheResponse();
  const rerun = await rerunPoll();
  const prompts = Array.isArray(rerun.prompts) ? rerun.prompts.map((p) => p.prompt) : [];
  const ok = rerun.status === "feedback" && prompts.length === 1 && prompts[0] === text;
  if (!ok) {
    lost += 1;
    // Drain whatever is left so the next trial starts clean.
    await rerunPoll();
  }
  results.push({ trial, sent_by_reviewer: text, rerun_status: rerun.status, rerun_prompts: prompts, feedback_kept: ok });
}

console.log(JSON.stringify({ build: label, trials: results.length, feedback_lost: lost, results }, null, 2));
process.exit(lost === 0 ? 0 : 1);
Evidence: Before/after transcript of the reviewer's Send & End during agent work
Reviewer action while the agent is working ("Working..." bubble on screen), same artifact,
same moment, driven through the real browser chrome.

BEFORE the fix (ec50b1e, port 4418, session 33427e5f20b7f5e5)
  chrome DOM state ........ send.disabled = true, sendAndEnd.disabled = true
  reviewer types "Looks good - ship it." and clicks Send & End
  server-side result ...... state.json session stays status="open", ended_by=null, pending prompts = []
                            -> the click is a no-op; the reviewer can neither send nor end
                               until some later poll happens to attach.

AFTER the fix (9ea0104, port 4417, session e0cd74d7439829f1)
  chrome DOM state ........ send.disabled = false, sendAndEnd.disabled = false
  reviewer types "Looks good - ship it." and clicks Send & End
  server-side result ...... state.json session becomes status="ended", ended_by="user",
                            pending prompts = ["Looks good - ship it."]
  agent's next poll ....... status: feedback / session_ended: true / ended_by: user
                            prompts: "Looks good - ship it."
  SSE agent-presence ...... {"state":"waiting"}   (the final batch releases "working";
                            no later poll or reply could have)

Pipeline

Updates from git push no-mistakes

⏭️ **intent** - skipped

✅ No issues found.

⚠️ **Rebase** - 1 warning
  • ⚠️ AGENTS.md - merge conflict rebasing onto origin/main
⚠️ **Review** - 4 issues (2 errors, 1 warning, 1 info)
  • 🚨 src/server.js:489 - The durable fix in 9ea0104 ("restore feedback from closed immediate polls") only covers the immediate-take branch; the identical loss is still reachable through the long-poll respond() path. Concrete sequence: an agent runs a no-timeout lavish-axi poll, the poll attaches and waits; the user hits Send, events.emit(&#34;feedback&#34;) fires onFeedback -> respond(); respond() sets responding = true and awaits store.takeFeedback(key) (mutex + two fs ops, a real multi-tick window); the agent CLI is SIGINT/SIGTERM'd in that window, the socket closes, onRequestClose -> cleanup() runs. takeFeedback then resolves having already cleared session.prompts/artifact_failures, finishFeedbackDelivery marks delivery, and res.end(JSON.stringify(result)) writes to a destroyed socket. Note res.writableEnded stays false on client abort (only res.destroyed flips), so neither the respond() guard nor onFeedback's guard catches it. The batch is gone: the next poll returns waiting, and presence is left stuck on "working" with no active poll. This violates the invariant AGENTS.md states for the interrupted-poll path ("queued feedback persists, so re-running the same poll is safe"). Fix at the shared boundary rather than duplicating the immediate-branch patch: have respond() (and any future delivery site) route through one helper that re-checks requestClosed || req.destroyed after takeFeedback returns and calls restoreClosedFeedback instead of finishFeedbackDelivery + res.end.
  • 🚨 src/server.js:1214 - POST /api/:key/attachments (line 1214) and DELETE /api/:key/attachments/:id (line 1276) call isSameOriginRequest(req) without the allowedHostnames, allowAnyHostname arguments that every other call site passes (lines 347, 534, 745, 838, 1133, 1158, 1185). With allowedHostnames === undefined and allowAnyHostname defaulting to false, any request carrying an X-Forwarded-Host header reaches allowedHostnames.has(host.hostname) at src/server.js:1616 and throws TypeError: Cannot read properties of undefined (reading &#39;has&#39;), which the route's catch forwards to next(error) -> 500. Failure scenario: Lavish behind a reverse proxy (a configuration AGENTS.md and README document as supported), user drops an image into an annotation card -> chrome POSTs to /api/:key/attachments -> 500 -> the chip never leaves its error state and the card blocks queuing. Secondarily, these two routes also skip the forwarded-host-vs-allowlist validation the sibling guarded routes perform. Pass allowedHostnames, allowAnyHostname at both call sites.
  • ⚠️ src/session-store.js:227 - restoreClosedFeedback re-queues through queuePrompts with restore: true, which unconditionally overwrites session.artifact_failures (line 227) and session.dom_snapshot (line 232) with the values captured before the take. takeFeedback releases the store mutex before restoreClosedFeedback re-acquires it, so any mutation that wins the lock in between is silently discarded. Failure scenario: the SDK posts a fatal artifact-asset-unavailable to /api/:key/artifact-failures in that window; the restore then writes artifact_failures: [] back over it and the fatal signal never reaches the agent. Same window for a concurrent /prompts POST: its fresher dom_snapshot is replaced by the stale restored one, and its prompts end up ordered before the older restored batch. Merge (append restored failures to whatever is currently stored, and keep the newer snapshot) instead of overwriting.
  • ℹ️ src/server.js:254 - finishFeedbackDelivery reads result.chat, deletes it, and emits chat-sync on events, but neither half is live: SessionStore.takeFeedback (src/session-store.js:535-541) never puts a chat key on its result, and the SSE handler at src/server.js:936 registers listeners for reload, agent-reply, agent-presence, and layout-warnings only - there is no events.on(&#34;chat-sync&#34;, ...), so the emit has no subscriber. The block is inert; either wire the listener (if browser chat sync on delivery was the intent) or drop the three lines so the helper reads as what it does.
✅ **Test** - passed

✅ No issues found.

  • npx -y pnpm@10 install --frozen-lockfile (node_modules was absent in the worktree)
  • node --test test/server.test.js — 189 pass, includes the new a disconnect during immediate feedback take requeues the batch without working presence, a poll dropped before it arms never leaves presence listening, immediate poll delivery leaves presence working and preserves the next send, immediate send-and-end delivery clears working presence without an active poll
  • node --test test/chrome-client-queue.test.js test/session-store.test.js — 142 pass, includes send controls stay enabled while the agent works and lock only once the session ends, chrome client sends queued prompts while the agent is working, warning fixes stay queueable while the agent is working
  • Manual end-to-end on the fixed build: node bin/lavish-axi.js &lt;artifact&gt; --no-open --no-gate -> real Chrome via chrome-devtools-axi -> reviewer sends a message -> node bin/lavish-axi.js poll &lt;artifact&gt; delivers it and presence flips to Working -> reviewer sends a second message while Working -> node bin/lavish-axi.js poll &lt;artifact&gt; --agent-reply &#34;...&#34; returns that second message
  • Manual end-to-end on the pre-fix build (git archive ec50b1e into a scratch tree, second server on port 4418): same flow shows send.disabled = true / sendAndEnd.disabled = true while Working, and a Send & End click leaves state.json at status=open, ended_by=null, zero pending prompts
  • node interrupted-poll-probe.mjs (10 trials per build) — reviewer queues feedback, the agent poll connection dies before the server writes its response, then a normal poll re-runs: 10/10 lost before the fix, 0/10 lost after
  • Reviewer Send & End on the fixed build -> poll returns status: feedback, session_ended: true, ended_by: user; curl -sN /events/&lt;key&gt; then reports agent-presence {&#34;state&#34;:&#34;waiting&#34;}
  • Cleanup check: git status --porcelain shows only the pre-existing untracked .planning/; both test servers stopped and scratch state dirs removed
✅ **Document** - passed

✅ No issues found.

⚠️ **Lint** - 1 info
  • ℹ️ .prettierignore - The untracked .planning/ directory (agent-harness scratch: skill-dedup JSON and telemetry/hook-metrics.json) fails prettier --check ., which is part of npm run check. It is neither gitignored nor prettierignored. It is not part of this change, so I did not format it or modify the ignore files (out of scope). The executor should exclude it from any commit; if it recurs, adding .planning/ to .gitignore is the fix.
✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 24 commits August 20, 2026 01:55
* test: tolerate coalesced heartbeat chunks in the long-poll heartbeat test

Under full-suite load two 10ms heartbeat writes can arrive in one TCP chunk,
failing the one-byte-per-read assertion. Collect bytes until two heartbeats
have streamed and assert they are all whitespace, which is the actual contract.

* feat: steer agents away from unpainted, invisible artifacts

An agent-built artifact that styles light text while never painting its own
page background renders invisible over Lavish's light surface, because Lavish
deliberately injects no design system. Make the authoring guidance and the CLI
itself steer agents away from that failure:

- SELF_PAINT_RULE (single-sourced in src/design-reference.js): every artifact
  must set an explicit page background plus matching high-contrast text color,
  theme-aware via prefers-color-scheme. Leads home visual_guidance, so it
  reaches the no-args home, top-level --help, the SessionStart hook context,
  and the generated skill; also surfaced in `lavish-axi design` output and help.
- RENDER_VERIFY_RULE (src/cli.js): render-verify the composed page by loading
  the saved HTML file itself in a browser before presenting it - verifying the
  inputs is not verifying the page, and the served session URL must never be
  the verification target because a chrome load supersedes the user's reviewer
  handoff. In home help and as an explicit skill workflow step.
- src/self-paint.js: render-free, fail-open check wired into open/export/share
  that returns a one-line self_paint_warning plus fix-first next_step guidance
  when an artifact has no background signal on html/body/:root and no
  stylesheet that could provide one. Any stylesheet link, @import, Tailwind
  runtime script, or color-scheme suppresses it, so false positives stay near
  zero; it never blocks the open.

* no-mistakes(review): Correct unsupported share warning guidance

* no-mistakes(document): Verify self-paint documentation and lint
* fix: slim invisible artifact guidance

* no-mistakes(review): Fix guidance coverage and remove dead references

* no-mistakes(document): Refresh invisible-artifact guidance and pass lint
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
… hardening) (kunchenguid#194)

* fix(server): confine the artifact asset route by realpath to block symlink-escape (Phase-B sweep)

resolveArtifactAsset() confined asset paths lexically only (path.resolve/relative — a
string check that never touches disk), so a symlink whose NAME sits inside the artifact
dir but whose TARGET points outside was served by res.sendFile, which follows it. The
export path (guardedRead) already defends this exact threat with realpath resolution.
Live-verified: GET /artifact/<key>/<symlink> returned 200 + leaked an outside file
pre-fix, 403 post-fix. The route is reachable by script inside the artifact (its HTML
embeds the session key). Fix: make resolveArtifactAsset async, realpath the resolved
path + root, reject anything escaping (mirrors guardedRead); both call sites await it.
node --test 592 pass / 0 fail (+2 new symlink-escape tests), eslint + tsc clean.

* no-mistakes(review): fail closed on non-ENOENT realpath errors in resolveArtifactAsset

* no-mistakes(document): document realpath symlink confinement on artifact asset routes

* fix(server): return the resolved asset path and cover both routes with regression tests

Harden the realpath confinement added in this branch and prove it:

- resolveArtifactAsset now hands back the symlink-resolved path instead of the
  requested one. A real path contains no symlinks, so sendFile re-opening it
  cannot be redirected by a link swapped in between the check and the read.
- Regression tests for both routes it guards (/artifact and /whiteboard-assets):
  a leaf symlink escape, an escape through an intermediate directory symlink,
  and an in-directory symlink that must still resolve (no over-blocking).
- Guard the pre-existing lexical ".." rejection at the HTTP level via raw
  requests, since fetch collapses ".." in a URL before it reaches the server.

All seven escape tests fail against the pre-fix resolver (the /artifact case
serves the outside file with 200); the two lexical-traversal guards pass in
both states, which is what makes them preservation checks.

---------

Co-authored-by: Nic Nogueira <tibernero@proton.me>
Co-authored-by: kunchenguid <kun@kunchenguid.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(server,chrome): bind whiteboard channels to their session and guard prompt submission

The session key is derived from the artifact path, not a secret, and the
whiteboard frame page is framable by any origin, so neither key nor channel
token possession can stand in for authorization:

- Whiteboard channel tokens are now signed over the session key, making a
  token a capability for exactly one session; /whiteboard-frame requires
  ?key= and both call sites pass their own.
- The chrome accepts inline whiteboard messages only from windows that
  descend from the current artifact frame, mirroring the guard the ordinary
  artifact-annotation handler already applies.
- POST /api/:key/prompts is same-origin guarded like /share and the
  whiteboard write routes.
- The session chrome page answers X-Frame-Options: DENY and
  frame-ancestors 'none'. Scoped to that route: /artifact/* is framed by the
  chrome, and /whiteboard-frame is framed by the artifact document, whose
  sandbox gives it an opaque origin no frame-ancestors expression can name.

Addresses GHSA-w887-pf37-frrv.

* no-mistakes(review): Preserve secure proxied same-origin feedback

* style: apply prettier to the proxied same-origin fix

* no-mistakes(review): Harden host authority validation against origin bypasses

* no-mistakes(review): Reject authorities without valid canonical origins

* no-mistakes(review): Preserve wildcard proxy routing with strict authority validation

* no-mistakes(document): Document proxy headers and fix lint formatting
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
… them to the agent (kunchenguid#188)

* feat(attachments): content-addressed image storage in the state dir

Add src/attachment-store.js: magic-byte image validation (PNG/JPEG/WebP),
content-hash ids, atomic dedup writes under <state-dir>/attachments/<key>/,
and a resolveAttachment trust boundary that re-derives every field (absolute
path, mime, bytes, dimensions) from disk rather than from the caller. Limits
are LAVISH_AXI_* env-configurable.

* feat(attachments): upload, fetch, and remove endpoints

Add POST/GET/DELETE /api/:key/attachments routes. Upload takes raw image
bytes via express.raw (per-image byte cap -> 413) and is same-origin guarded
like the whiteboard writes; fetch serves the stored bytes with the resolved
mime; delete removes the file idempotently. Non-images 415, unknown sessions
404.

* feat(attachments): resolve prompt attachments at the trust boundary

Prompts carry only a client id + display name; queuePrompts re-resolves each
id against the on-disk store, injecting the authoritative path/mime/bytes/dims
and enforcing per-prompt count and total-byte caps. Unknown ids are dropped and
a missing resolver drops attachments rather than trusting client metadata, so a
crafted /prompts POST cannot aim an attachment at an arbitrary file.

* feat(attachments): reference-aware TTL sweeper and disk-cap backstop

Add listAttachments + sweepAttachments to the store and referencedAttachmentIds
to SessionStore. A file is reaped only when past its TTL AND not referenced by a
pending prompt (so send-and-end batches survive); the optional disk cap then
evicts oldest UNREFERENCED files, never referenced ones. The server sweeps at
startup and hourly, skipping entirely when neither TTL nor disk cap is set.

* feat(attachments): capture images in the annotation card (SDK)

The annotation card gains paste, drag-drop, and an Attach-image picker for
PNG/JPEG/WebP. Captured files render as chips with a thumbnail and upload/
ready/error state; the SDK reads bytes and hands them to the chrome (which owns
the server round trip), applies upload results, and supports remove/retry. On
queue it rides the ready attachments' server ids along with the prompt.

* feat(attachments): chrome upload orchestration and pill thumbnails

The chrome performs the same-origin attachment upload/delete on behalf of the
sandboxed iframe and reports the server-vetted id back to the card; queued-prompt
pills render image thumbnails from the attachment endpoint and label image-only
annotations.

* feat(attachments): poll delivery guidance, docs, and kunchenguid#123 non-regression

Poll next_step tells the agent when prompts carry image attachments and to open
their local paths. Document the feature and its LAVISH_AXI_* limits in README
(user-facing contract) and the storage/trust/sweeper invariants in AGENTS.md.
Add a non-regression test that export coexists with the raw-upload route and
never leaks attachment ids/paths/SDK into an exported bundle.

* fix(attachments): deliver a clean 413 on over-cap upload instead of hanging

VERIFIED: over-cap uploads DID get an HTTP 413 (curl/undici recovered it), but
express.raw aborts on Content-Length WITHOUT draining the body, so a browser
mid-upload is reset and the chip hangs on "uploading" with no recovery.

Fix both paths: the route now reads the stream via readAttachmentUploadBody,
which buffers up to the cap but always drains to end-of-body before sending a
clean 413 the browser reliably receives. Belt-and-suspenders, the chrome pre-
checks byte length against attachmentMaxBytes (surfaced in the session JSON) and
fails the chip locally before uploading. Both reach the error+retry state; the
errored chip never blocks queueing.

* fix(attachments): intercept non-image drops and show a visible remove X

Drops over the annotation card now preventDefault unconditionally, so a dropped
PDF can't navigate the frame; non-images surface a dismissible UNSUPPORTED_TYPE
error chip (no thumbnail, no retry). The chip's remove control is a larger,
higher-contrast, titled X. Trim the card helper text.

* polish(attachments): use Lavish's native close icon for the chip remove control

Match the app's own close buttons (.pill-close / .share-close): same path,
viewBox, 1.6 stroke weight, and round caps, in a restrained circular button
(subtle white, hover fill) instead of the heavier foreign-looking glyph. Remove
function, title, and aria-label are unchanged.

* fix(attachments): render the chip remove-X (padding override) + clean inline glyph

Root cause of the empty/squished remove-X: the generic '.lavish-annotation-card
button' rule (padding:8px 10px, specificity 0,1,1) outweighs '.lavish-attachment-
remove' (padding:0, specificity 0,1,0), so with box-sizing:border-box the 22px
button's content box collapsed to ~2x6px and squished the glyph into a speck. Add
padding:0!important (the sibling .lavish-attach button already uses !important for
this same conflict) so the icon gets its full box. Also inline a clean, self-
contained X path (no sprite/symbol/use/mask reference). Verified end-to-end by
driving the real annotation card in a sandboxed iframe (real drop -> real chip):
the X now paints crisp and centered. Remove function, title, aria-label unchanged.

* harden(attachments): seal v2.5 concurrency, lifecycle, and security gaps

Round-3 conformance to the sealed design v2.5 for the kunchenguid#103 attachment
feature (no-mistakes flagged Risk HIGH). Four groups:

Group 1 — lifecycle & concurrency
- D5: one shared AsyncMutex serializes upload-finalize, /prompts
  resolve+persist, delete, and the reference-aware sweep, closing the
  window where a reference is acquired between the sweeper's snapshot and
  its delete (src/async-mutex.js).
- refcount: DELETE is reference-counted under the lock — a content-
  addressed file still referenced by a queued prompt survives a chip
  removal (status "referenced") instead of breaking the queued thumbnail.
- B3: dedup re-upload refreshes the file mtime so a re-referenced aged
  file isn't reaped by the next sweep.

Group 2 — batch & queue integrity
- C4: attachment resolution is all-or-nothing — an unknown id or a
  count/byte-cap breach rejects the whole batch (400, persist nothing)
  with { rejected, caps }; the chrome keeps its queue and surfaces the
  reason instead of silently dropping images.
- R2.4: the SDK card gates queuing on hasPending() so an in-flight upload
  can't be dropped by collectReady/closeCard, and only deletes a removed
  chip's file when no sibling shares its deduped id.

Group 3 — security
- The chrome mediates the same-origin confused-deputy: rate limit + per-
  session cumulative-byte quota before any upload hits the loopback
  server, plus a bounded default disk quota (512 MiB, off/0 disables).

Group 4 — perf
- D6: persist vetted dims in an <id>.meta sidecar at upload;
  resolveAttachment reads it instead of re-parsing the image, and the
  thumbnail GET serves via a lightweight stat (statAttachmentForServe),
  removing the ~2x full-image read per render.

* no-mistakes(review): Harden attachment payload validation and JPEG TEM parsing

* no-mistakes(document): Document attachment guards and fix lint

* fix(attachments): close state race and harden the annotation card (round 4)

Round-4 conformance fixes for the kunchenguid#103 image-attachment feature, on top of
c2db2df. Addresses the four defects no-mistakes flagged (1 error + 3 warnings).

E1 (state race, DRIVER): queuePrompts held its pre-resolve state snapshot across
the attachment-resolution await while takeFeedback/recordLayoutWarnings mutated
state.json without any shared lock, so a poll landing mid-resolution clobbered the
write (resurrecting delivered prompts or dropping queued ones). SessionStore now
owns ONE AsyncMutex and serializes EVERY state.json read-modify-write under it
(queuePrompts, takeFeedback, recordLayoutWarnings, upsertSession, endSession,
addAgentReply). The server routes its attachment disk-lifecycle sections through
store.runExclusive so the D5 file lock and state consistency share one lock;
referencedAttachmentIds stays lock-free (called from inside runExclusive).
Covered by a regression test that fails without the unified lock.

W1 (hardcoded cap): the card's per-prompt count cap was a literal 4. createSdkJs
now threads the server's maxPerPrompt (LAVISH_AXI_MAX_ATTACHMENTS_PER_PROMPT) into
the SDK, and an over-cap pick is surfaced instantly instead of silently swallowed.

W2 (error attachment lost): queuing gated only on in-flight uploads, so a text send
closed the card and discarded an errored/rejected chip with its retry/remove UI.
tryQueue now also gates on hasErrors(), keeping the card open with a notice.

W3 (card leaves viewport): the card top was clamped once with the initial height, so
attachment rows grew it past the frame and hid Queue/Cancel. The card re-clamps after
every attachment-row render, plus a max-height/scroll backstop on the chip list.

* docs: reconcile SessionStore locking invariant with the E1 mutex

The E1 change made every SessionStore mutation acquire a shared AsyncMutex, but
the 'Things to know' bullet still claimed SessionStore has 'no locking'. no-mistakes
flagged the contradiction; update the bullet to point at the store lock and the
Image attachments (E1) section that owns the invariant.

* fix(attachments): count hidden preview images and defer racy deletes (round 5)

W-A: the queued-prompt pill always sliced its thumbnails to four, but
LAVISH_AXI_MAX_ATTACHMENTS_PER_PROMPT is configurable, so a prompt could
legitimately carry more and the preview silently hid the surplus - the queue
looked like it had lost accepted attachments. The overflow now collapses into a
+N badge instead of disappearing.

W-B: removing a ready chip deleted its stored file whenever no OTHER chip
already held that content-addressed id. A twin uploading identical bytes dedups
onto that very id but carries none until its upload returns, so the check missed
it and the delete pulled the file out from under the twin: its upload finalized
onto a deleted id and the send then failed as not-found. classifyAttachmentDelete
now parks an id whose fate an in-flight upload could still change and re-decides
it as uploads settle; the server's reference count and the TTL sweeper stay the
backstop, so parking can only leak a file, never break a live chip.

The card's count-cap notice shares its line with the neutral keyboard hint and
inherited its passive gray, which read as help text rather than a rejection. It
now renders in the error color and is restored once the condition clears.

* no-mistakes(review): Fix attachment reference, notice, and dedup handling

* no-mistakes(review): Preserve cap notice across transient attachment states

* no-mistakes(review): Derive attachment notices from current controller state

* no-mistakes(document): Document attachment preview invariant and fix formatting

* fix(attachments): make image files private and close store accounting gaps

Closes four findings from the review-only gate on 0c02f37:

- E4 (security): images, dims sidecars, and their dirs were created with the
  process umask, landing 0644 in 0755 dirs and exposing screenshots to other
  local users. Modes are now set explicitly at creation (0600/0700), and the
  dir modes are re-asserted so installs that uploaded before this hardening
  do not keep their existing images exposed.
- W2: a dedup upload swallowed a failed mtime refresh and still reported
  success, leaving the file TTL-expired and sweepable before the prompt was
  queued. It now falls back to an atomic rewrite of the identical bytes, and
  propagates when that fails too, instead of reporting a false success.
- W3: an expired orphan whose delete failed was dropped from `survivors`, so
  its still-present bytes vanished from disk-cap accounting and the quota
  could stay exceeded while nothing was evicted. It is kept as an
  unreferenced survivor.
- W5: a fractional limit such as `0.5` passed the positivity check and then
  floored to 0, disabling uploads server-side while the SDK kept advertising
  its own default. Both env resolvers now floor first and require >= 1.

`touchFile` is injected into writeAttachment so the W2 refresh failure is
testable without depending on a filesystem that rejects utimes. The mode and
permission-failure tests are POSIX-only; Windows has no equivalent.

* fix(attachments): close the untrusted-input, race, and DoS findings

Closes the remaining seven findings from the review-only gate on 0c02f37,
each with its own failing-then-passing test.

- E1 result-correlation: the SDK's message listener accepted any sender and
  correlated upload results by chip id alone. Chip ids restart at att-1 on
  every document load, so a result in flight across an iframe reload marked a
  new chip ready with the previous document's image, and a same-window forged
  message could hand a chip any id. Uploads now carry a per-document nonce
  that results must echo, and the listener requires event.source === parent.
- E2 delete confused-deputy: the chrome honored deletes driven by the
  untrusted iframe, while its reference checks could not see a ready-but-
  unqueued chip in another tab, so one tab could destroy bytes another live
  card still needed. The eager delete is gone; the reference-aware sweeper
  owns reclamation. This removes the classifyAttachmentDelete/defer machinery
  it existed to support.
- E3 lock-DoS: attachment resolution ran an uncapped sequential stat per ref
  while holding the store's single global mutex, because the cap counted
  RESOLVED refs and unknown ids never advance it. Raw per-prompt and
  request-wide ref counts are now rejected before the resolver is called.
- E5 preview-poison: a queued prompt's attachments were dereferenced
  unvalidated, and the queue persists before it renders, so attachments:[null]
  threw out of render and re-threw on every reload. Refs are filtered to
  well-formed {id} objects at the enqueue and restore boundaries.
- W1 malformed-drop: malformed attachment fields were silently normalized
  away and the POST then succeeded, so the chrome cleared a queue whose images
  were never delivered. Malformed input now fails the whole batch with 400.
- W4-a mixed-drop: a drop of images plus unsupported files accepted the images
  and ignored the rest, because the error branch only ran when no image was
  found. Both halves are now reported: the images attach and each unsupported
  file raises a visible error chip.
- attachment-post-poll-retention: the sweeper's reference set covered only
  pending prompts, which takeFeedback clears on delivery, so an attachment
  could be reaped while the agent was still reading the path it had just been
  handed. Delivered ids are retained for a bounded read grace.

Two existing tests asserted the old contracts that two of these findings
identify as defects (the eager delete, and "after the agent takes the feedback
the reference clears"); both are updated to the corrected behavior rather than
deleted.

* docs: reconcile the attachment invariants with the round-6 fixes

AGENTS.md documented the behaviors these findings removed - the eager
`lavish:removeAttachment` delete and its classifyAttachmentDelete parking, and
"takeFeedback clears delivered prompts, which is what makes their attachments
sweep-eligible" (the post-poll-retention defect, written down as intent).

Records what neither the code nor the tests show at a glance: why the raw ref
bound must stay ahead of the first filesystem await, why no eager delete may be
reintroduced, why upload results are nonce-scoped, and that the delivery grace
is a bounded read window rather than a second lifetime.

* docs(server): note the delete route's grace-aware guard and its unused-by-chrome status

* fix(attachments): close three defects the after-gate found in the round-6 fixes

The after-review on e841887 confirmed all eleven findings closed and then found
three defects in the fixes themselves. Adjacent bugs bred by a fix are the exact
failure this round exists to correct (E5 was born in round-5's own W-A fix), so
these are part of closing E5 and post-poll-retention, not new scope.

- The E5 sanitizer filtered entries but returned the artifact's own objects.
  postMessage delivers a structured clone, which preserves BigInt values and
  cycles that JSON.stringify then refuses, so the junk rode along into
  sessionStorage and the POST body and made the queue unsendable - the same
  wedge one step later. Each ref is now projected onto a fresh {id, name}.
- `upsertSession` rebuilds a session from an explicit field list and did not
  carry `delivered_attachments`, so re-opening the artifact inside the grace
  hour erased the retention and handed the next sweep a path the agent was
  still reading. The field is carried, with a note that this constructor
  silently drops anything it omits.
- Delivery-grace entries were appended per delivery, so one reused
  content-addressed image consumed many slots and could evict distinct
  attachments still inside their own grace. Retention is now keyed by id: a
  re-delivery refreshes the window instead of taking another slot.

Two further after-gate findings (an object-count/sidecar disk-quota bypass, and
the SDK materializing a file buffer before checking its size) are real but
pre-existing and outside the ruled scope of the eleven; they are surfaced to
firstmate rather than fixed here.

* fix(attachments): derive the delivery retention bound from the request bound

The final after-review caught a contradiction between two constants introduced in
this same round: /prompts accepts up to MAX_REQUEST_ATTACHMENT_REFS (256) images
per batch (the E3 pre-resolver bound), but delivery retained only the newest
MAX_DELIVERED_ATTACHMENTS (200). A max-size batch therefore left 56 delivered
paths immediately sweepable while the agent was reading the very response that
handed them over - the post-poll-retention hole reopening for the largest batch
the E3 bound permits.

Retention is now derived from the request bound rather than hand-picked
separately, so whatever a single batch may queue, a single delivery can always
protect, and the two cannot drift apart again.

* fix(attachments): retain the whole delivery, not a constant's worth of it

Third defect found in this round's own retention work, and the same root cause
each time: the bound was a NUMBER when the invariant is structural.

Prompts accumulate across an unbounded number of accepted POSTs until a poll
drains them, so a single takeFeedback can legitimately deliver far more than one
request may queue. Retention sized to the per-request bound therefore sliced the
delivery itself: two legal 256-ref batches queued before one poll produced 512
delivered paths and left the first 256 immediately sweepable - while the agent
was reading exactly those paths.

`takeFeedback` now retains every id in the current delivery in full, whatever its
size, and MAX_DELIVERED_ATTACHMENTS bounds only the history of EARLIER deliveries
that rides along. The current delivery is never trimmed, so no constant has to be
correct for the invariant to hold.

* feat(attachments): round 7 - client size gate + real disk accounting

Round 7 = the two deferred items César scoped IN (a+b); the (c) cluster
(DEFERRED-3/4/5/6) stays tracked and deferred.

(a) The SDK now rejects an over-limit file in add() BEFORE createObjectURL or
    arrayBuffer, so a multi-gigabyte drop is never read into a buffer and
    structured-cloned into the chrome only to be rejected afterwards. The server
    byte limit is threaded into the SDK via createSdkJs (mirroring the existing
    maxAttachmentCount wiring); attachmentSizeError is a pure exported helper
    serialized into the bundle, and the oversized file surfaces as a dismissible
    error chip. The server still re-checks authoritatively.

(b) The disk cap now charges REAL allocation instead of logical image bytes:
    listAttachments reports chargedBytes = the image rounded up to whole 4096-byte
    blocks plus the sidecar's own block, and sweepAttachments enforces the cap on
    that. This closes the measured ~683x undercount, where a flood of 12-byte
    magic-prefix uploads sat far under the reported total while consuming real disk
    and inodes. A derived object-count bound (disk budget / min per-object charge,
    never a separate knob) backstops the inode dimension.

Each closed with its own forced RED->GREEN test. Two existing disk-cap tests that
expressed the cap in logical .bytes are rewritten in terms of chargedBytes, since
that is the accounting basis the fix intentionally changed.

* docs: record the round-7 disk-accounting and client size-gate invariants

* perf(attachments): make the disk-cap sweep O(n log n), not O(n^2) under the mutex

The round-7 gate caught a regression in round 7's own (b) code: the eviction loop
recomputed the charged-byte and object totals by filtering and reducing over ALL
survivors on every candidate, making a large over-cap sweep O(n^2) - while holding
the store's single global mutex. That is an E3-class lock hazard (the very thing
the lock-DoS fix closed), reintroduced by the disk-accounting change.

The totals are now computed once and decremented per successful removal. Behavior
is identical - a new test pins that a multi-file over-cap sweep still evicts the
exact minimum, oldest-first - so this is a complexity fix, not a behavior change.

* docs: note the sweep keeps running totals to avoid an under-mutex O(n^2)

* feat(attachments): round 8 - upload concurrency bound, temp-file reap, batched render

The tight bundle César scoped for round 8 (D8 + D6 + D7), each with its own
forced RED->GREEN test. D3/D4/D5/D9 stay tracked as architectural/ask-user.

- D8 (chrome-client.js): the confused-deputy guards bounded upload rate and
  cumulative bytes but not how many ran AT ONCE, so a hostile artifact could post
  ~30 large bodies in one tick and hold hundreds of MiB of clones + server buffers
  concurrently before the cumulative quota tripped. An in-flight ceiling
  (UPLOAD_MAX_IN_FLIGHT) now refuses over-cap uploads with a retry hint, and a
  settled upload frees a slot. RED: eight held-open uploads all hit the network at
  once (got 8).
- D6 (attachment-store.js): writeFileAtomically writes `<name>.<pid>.<n>.tmp` then
  renames; a crash in between left temp files that ID_RE excludes, so no TTL, disk,
  or object cap ever counted or removed them - a permanent leak. sweepAttachments
  now reaps temp files matching the store's own pattern older than a 5-minute grace
  (a live write renames in ms). RED: a stale orphan temp survived the sweep.
- D7 (artifact-sdk.js): add()/rejectUnsupported each rebuilt the whole chip DOM, so
  a multi-file drop was O(N^2). A new pure classifyAttachmentBatch decides the whole
  batch in one pass; addFiles and a batched rejectUnsupportedBatch now render once.
  RED: seam-first (the classifier did not exist). The round-7 (a) size gate now
  lives in the classifier and still precedes the single createObjectURL, so an
  oversized file is decided "error" before any read - re-pinned accordingly.

Two existing tests updated for the new structure: the confused-deputy rate-cap test
now settles each upload so the new in-flight bound does not mask the rate cap, and
the round-7 (a) bundle pin follows the size gate into the classifier. Neither is
weakened - the invariants they assert are unchanged.

* docs: record round-8 upload-concurrency, temp-reap, and batched-render invariants

* fix(attachments): reap orphan .meta sidecars and delete them first (ATTACH-002)

The round-8 gate found a leak in the same orphan-file class D6 addressed:
removeAttachment/removeFile deleted the image first, then the sidecar with a
suppressed failure, so a crash between the two left an orphan `.meta` that ID_RE
hides from every cap - a permanent leak. Completes D6's orphan-file hardening.

- The sweep's orphan reap now also removes `<id>.meta` sidecars whose image is
  gone. No grace: a live upload writes the image before its sidecar, so a
  sidecar-without-image is unambiguous debris (unlike a .tmp, which may be a live
  write). RED: an orphan sidecar survived the sweep.
- removeAttachment/removeFile now delete the sidecar FIRST, so a crash after the
  first removal leaves only the counted image (TTL/disk/object reclaimable), never
  the uncounted sidecar.

The other round-8 gate finding (ATTACH-001: the fixed 256 request-wide ref bound
vs a configurable per-prompt count) is a config-coherence ask-user decision, left
tracked as DEFERRED-10 for César alongside D3/D4/D5/D9 - not patched.

* docs: record the orphan-sidecar reap and sidecar-first delete invariant (ATTACH-002)

* fix(attachments): enforce the disk cap at the upload admission chokepoint and shield ready cards from eviction

Item 1 (root cause B): uploads never checked current usage, so several chrome
pages could each write past `maxDiskBytes` before the next periodic sweep,
defeating the documented disk-cap backstop. Route every byte-adding write path
through one admission chokepoint (`admitAttachmentCharge`) inside `writeAttachment`:
before a NEW image+sidecar (or a dedup sidecar repair) touches disk, reclaim
unreferenced/expired bytes toward (cap - newCharge), then measure the true
committed allocation on disk and refuse with 507 when it plus the new charge still
exceeds the cap. The whole reference-snapshot + reclaim + write runs under the
server's lifecycle lock. Committed bytes are measured by an independent tree walk
so `sweepAttachments`' public return shape is unchanged.

Item 2 (attachment-store.js eviction filter): cap eviction removed the oldest
UNREFERENCED file with no minimum-age floor, so it could delete a freshly uploaded
"ready card" (an image dropped into the composer but not yet queued on a prompt,
hence unreferenced) out from under the imminent Send -> unrecoverable not-found,
user-visible data loss. Add a bounded upload grace (`ATTACHMENT_EVICTION_GRACE_MS`,
1h, symmetric with the delivery grace): files younger than the grace are never
cap-evicted. The grace is opt-in via a `sweepAttachments` option (default 0 = off)
so the pure mechanism and its existing tests are unchanged; the server passes it on
both the periodic sweep and the upload admission. The grace never lets the total
exceed the cap (admission still 507s when only referenced/fresh bytes remain); it
only makes the cap prefer refusing a new write over destroying a ready card.

Tests: test/attachment-disk-admission.test.js reproduces both vectors (RED on the
base tip, GREEN here) - a new over-cap upload is refused, and a fresh ready card
survives disk pressure so Send still resolves it.

* test(chrome-client): stop reserving 300 MiB in the quota test (item 3)

The "oversized upload is refused" case allocated a real
`new ArrayBuffer(300 * 1024 * 1024)` (300 MiB) purely so the upload size check
would read an over-quota `byteLength` - 300 MiB of real memory per run and a
needless OOM risk in CI. The size check only reads `byteLength`, so spoof a real
view (`ArrayBuffer.isView` still true) whose `byteLength` own property REPORTS the
over-quota length without reserving the bytes. The test still exercises the exact
same session-quota refusal path.

* test(attachments): cover attachment routes under the Host-allowlist (DNS-rebinding) guard

The attachment upload/fetch/delete routes must sit behind Kun's Host-header
allowlist, not only the same-origin guard. A DNS-rebound page carries its hostile
domain in both Origin and Host, so isSameOriginRequest still matches; only the
Host allowlist rejects it. Assert a rebound request (Origin == forged Host) is a
clean 403 forbidden host on all three attachment routes while a legitimate
loopback same-origin request still succeeds.

* no-mistakes(document): document attachment disk-cap admission and eviction grace

* style(server): format rebased SDK response

* fix(attachments): route SDK uploads through postArtifactMessage so the chrome mediates them

The annotation card sent lavish:uploadAttachment via a raw parent.postMessage
with no artifact_load_token, and the chrome drops every artifact message whose
token is not the current load's before dispatch - so real uploads were silently
discarded while every mocked harness stayed green (the chrome harness patched
the missing token into test messages). Send the upload through
postArtifactMessage like every other SDK message, make the chrome harness send
messages verbatim, pin the token gate for uploads, and add a real-browser e2e
that round-trips an actual image through the gate into the attachment store
and back to the agent poll.

---------

Co-authored-by: “420tombombadil” <“dijongui@gmail.com”>
Co-authored-by: kunchenguid <kun@kunchenguid.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
kunchenguid#246)

* fix(whiteboard): preserve mermaid label line breaks in Excalidraw

The converter left <br> and \n as literal characters in skeleton labels, so adjacent words fused and the box sized as one line. Convert those markers to real newlines and size the bound text from the resulting line set.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(whiteboard): recenter bound text when linebreak pass grows the box

Growing the container from its center left the independently stored bound-text x/y at the old coords, so a multi-line label sat off-center on first paint. Reposition the bound text into the resized box.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Skip labelled-arrow resize in linebreak restore

* no-mistakes(document): Document labelled-arrow skip for linebreak restore

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
…nchenguid#249)

* chore(agents): use @AGENTS.md import instead of CLAUDE.md symlink

* chore(agents): exempt the CLAUDE.md pointer from prettier

* test(agents): assert CLAUDE.md is a real @AGENTS.md pointer, not a symlink

---------

Co-authored-by: Kun Chen <kun-1@kunchenguid.com>
* fix(whiteboard): prompt on source change only for real scene edits

Conversion autosaves on view, so a live Mermaid rewrite was treated as
stale user edits whenever a sidecar existed. Compare against the
conversion baseline and silently re-convert unmodified scenes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Preserve empty edited scenes without conversion baselines

* no-mistakes(review): Detect all meaningful whiteboard edits before reconversion

* no-mistakes(review): Compare geometry jitter using symmetric raw deltas

* no-mistakes(review): Correct whiteboard preservable-edit invariant

* no-mistakes(review): Normalize whiteboard baselines through Excalidraw restore

* no-mistakes(review): Cover mounted autosave conflicts with real Excalidraw regression

* no-mistakes(document): Format whiteboard conflict regression files

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* docs: add author-approved VISION.md

Adds the project's first VISION.md at the repo root, co-authored with the
repository owner over six rounds on a Lavish review board.

Every principle traces to real merged history (80 merged PRs mined, 15 bodies
read) or to a recorded author verdict. Twenty-four hypotheticals were
stress-tested; three reversed positions the initial evidence-only draft had
asserted, and four principles came from the author rather than the history.

Also registers VISION.md in the AGENTS.md documentation-ownership list as the
owner of the acceptance policy.

* no-mistakes: apply CI fixes
* fix(whiteboard): preserve Mermaid label line breaks

* fix(design): retain Mermaid label markup

* no-mistakes(document): Verify Mermaid label preservation documentation and lint
* fix(sdk): give table-cell annotations semantic row and column names

Port the accepted work from kunchenguid#242 onto current main so filtered or sorted
tables keep semantic row and column names while the clicked element's own
selector, tag, and text stay the locator.

Co-authored-by: Adelin <adelin-b@users.noreply.github.com>

* no-mistakes(review): Remove source-token SDK serialization test

* no-mistakes(document): Clarify semantic table target documentation

* no-mistakes: apply CI fixes

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Adelin <adelin-b@users.noreply.github.com>
* fix(server): reject cross-origin mutating requests on the loopback server

Add a global Origin/Referer CSRF guard after the Host allowlist so a
foreign page that can reach 127.0.0.1 cannot CSRF previously unguarded
mutating routes. Header-less CLI requests still pass; per-route
isSameOriginRequest checks stay in place.

Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

* no-mistakes(document): Clarify CSRF guard documentation

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
Keep feedback actions durable while the agent is working, close poll cleanup races, and clear delivered state for final feedback on ended sessions.

Fixes kunchenguid#229
@Julian-Dasilva

Copy link
Copy Markdown
Owner Author

Re-raised upstream as kunchenguid#263; closing this intermediate fork PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants