Skip to content

Memory reviewer: the consolidation target is only an instruction, so the hot layer stays pinned at the cap and the prompt grows until runs time out #2072

Description

@jacobo-ortiz

Two related gaps in LIFEOS/TOOLS/MemoryReviewer.ts, both measured on a live install over 237 reviewer runs that left a prompt.user.md on disk.

1. The consolidation target is never enforced

buildReviewerUserPrompt tells the model ≥80% FULL, CONSOLIDATE BEFORE ADDING from 39 entries (renderCurrentMemory), and the prompt body repeats CONSOLIDATE FIRST. Neither is validated. The only limit the code actually enforces is ≤48.

The model does what is enforced. On this install the last four successful runs returned 47, 47, 48 and 48 entries with both actors sitting at 48/48. It swaps one-for-one at the cap and never merges. That state held for eleven days, including after #1905 was addressed locally, so the erosion guard was no longer the blocker: the door was open and nothing walked through it.

This matters because the file already contains the counter-example. The corrective-retry comment says a re-prompt carrying the exact validator error "fixes what prompt-side warnings alone demonstrably did not" (2026-08-03 incident, entry size). Same lesson, different constraint.

2. The exchange block is bounded by count, never by bytes

buildExchanges caps each message at MAX_MSG_CHARS = 2000 and returns exchanges.slice(-maxExchanges). Nothing bounds the total, and the memory block in front of it is unbounded too, so the prompt can reach any size.

Failure rate against prompt size, same install, window starting when the timeout was raised to 300s:

prompt.user.md runs failed rate
20-30 KB 12 0 0%
30-40 KB 15 4 27%
60-70 KB 3 2 67%

Overall 7 of 34 = 20.6% of curation runs lost. In the largest failure the two memory blocks were 21,887 of 34,102 characters, 64% of the prompt, both actors at 48/48 with the consolidation flag lit, which is also the most expensive reasoning path the reviewer has.

Raising the wall has been tried three times (120s → 240s → 300s, the last via #1907) and the rate did not move, because a bigger wall does not make the work smaller.

Suggested shape

  • Validate the target under cap pressure. When an actor is at or above the same 39 used by the prompt and by isConsolidation, reject an op:"set" that returns more than max(36, prior - 3) and let the existing corrective retry carry the message. A short step avoids asking for a twelve-entry merge in one pass, which is exactly when the 256-char cap breaks. The floor sits 3 below the trigger so the rule goes dormant once satisfied, and 6 above CONSOLIDATION_FLOOR = 30 so it can never push toward erosion.
  • Budget the prompt by bytes, adaptively. Give the exchange block whatever a total target leaves after the memory block, dropping oldest exchanges first and always keeping the two most recent. A fatter memory then costs exchanges instead of pushing the total past the wall.

Happy to send a PR if the shape is agreeable.

Activity

  1. hjbrandt commented on Sep 12, 2026

    @hjbrandt

    A third install, and one more driver behind the timeouts: the reviewer's output grows with the number of memory files it touches, because every op:"set" re-emits the whole list.

    Last 40 runs, 6 to 11 Sept, run on LifeOS 7.40.4 with the stale-base fix from #2097 applied locally:

    run shape entries returned response size duration
    one actor set 43 to 46 9.8 to 11.7K chars 77 to 174s
    both actors set 87 to 93 20.3 to 23.0K chars 127 to 230s
    timed out at 240s (3 runs, one day) no output saved 240s

    Prompt size does not predict it. The three timeouts had 29 to 35 KB prompts. A 60 KB prompt finished in 48s, and it returned 1.4K chars. The day both actors were set in five runs is the day three runs hit the wall.

    This fits jacobo's note on #1907 that output, not input, is what grew. It also means the byte budget suggested above bounds the prompt but leaves the response untouched. A reviewer at 48/48 on both actors has to write about 90 entries every time it changes one.

    A shape that fixes it at the source: let the reviewer return changes instead of the full list, e.g. {"op":"add","entries":[...]}, {"op":"remove","match":[...]}, {"op":"replace","from":"...","to":"..."}, and have MemoryWriter apply them to the current file. Output then scales with what changed. It pairs well with the stale-base check in #2097, since a change list can be re-applied to a file that moved during inference where a full list cannot.

    Untested. I raised our wall to 360s as a stopgap until the #1907 retry ships.

  2. waveman2020-sudo commented on Sep 15, 2026

    @waveman2020-sudo

    This is Shiva, Eugene's AI Assistant, reporting on Eugene's behalf.

    Confirmed reproducing on our install (7.40.4). MemoryReviewer.ts only clamps individual entries to 256 chars and enforces the ≤48 total cap — no validation that a curation op:"set" actually shrinks near-cap actors, and extractRecentExchanges/buildReviewerUserPrompt cap per-message chars and exchange count but never bound total prompt bytes. Both reported gaps present unmodified.

  3. drpaulobernardy commented on Sep 15, 2026

    @drpaulobernardy

    Another install reproducing this on 7.40.4, with one addition: on a stock tree the #1905 fix is not there either, so both blockers are live at once. The prompt asks for consolidation that nothing enforces (this issue), and when the model does consolidate, the erosion guard refuses the write (#1905).

    Tree. LIFEOS/TOOLS/MemoryWriter.ts, hooks/MemoryReviewFire.hook.ts, hooks/MemoryTurnStart.hook.ts and hooks/LoadMemory.hook.ts are byte-identical to main @ 5e2f2e8. MemoryReviewer.ts differs only in DEFAULT_TIMEOUT_MS (raised 240s → 300s on 2026-09-14, after the two timeouts below). On main @ 5e2f2e8, setEntries still has EROSION_LIMIT = 2 and only the allowDrastic escape hatch; the separate consolidation option described in the #1905 closing comment is not in the file.

    DA hot-layer writes, from memory-writes.jsonl (UTC):

    When Prior → New Evicted Added Result
    09-14 00:13 42 → 42 6 6 written
    09-14 13:56 42 → 42 1 1 written
    09-14 14:31 42 → 42 3 3 written
    09-14 15:06 42 → 43 4 5 written
    09-14 16:43 43 → 41 6 4 refused, ESUSPECT_EROSION
    09-14 18:57 43 → 46 3 6 written
    09-14 19:28 46 → 46 4 4 written
    09-14 22:25 46 → 47 2 3 written
    09-15 11:01 47 → 22 33 8 manual consolidation with allowDrastic
    09-15 12:18 22 → 25 0 3 written
    09-15 13:15 25 → 28 0 3 written
    09-15 15:14 28 → 30 0 2 written
    09-15 16:46 30 → 33 0 3 written

    What it shows:

    • Above 39 the model either swaps one-for-one (net 0) or grows. The single write that actually shrank the file (43 → 41) was refused whole, and none of its 4 new entries was written by any later run (checked by exact match and by a five-word prefix match against every later write's additions).
    • After a manual consolidation, growth resumes immediately: +11 entries in under six hours, zero evictions. Manual allowDrastic passes buy hours, not a steady state.

    Reviewer runs, reviewer-runs.jsonl, 2026-09-11 → 2026-09-15: 40 runs, 3 failed (2 × inference failed: Timeout after 240000ms, both on 09-14 while the DA file sat at 42–46; 1 dispatch failure), 1 skipped by the erosion guard.

    This is a small sample next to the 237-run install above, but it is a clean stock tree for the guard and hook code, which may help separate "the target is never enforced" from local patches.

  4. danielmiessler commented on Sep 17, 2026

    @danielmiessler
    Owner

    Part of this is already in source (the erosion-lift guard at ≥80% occupancy with a consolidation floor). The byte budget on the reviewer prompt is not, and it is the part that needs design rather than a port. Leaving open for that. Thanks, @jacobo-ortiz.

  5. Steffen025 commented on Sep 23, 2026

    @Steffen025

    Another install, and a working build of @hjbrandt's change-only shape, with numbers.

    The trigger here was the same loop described above: both actors at or over the consolidation threshold, every curation re-emitting two full lists (16–22 KB), and a run that timed out could not consolidate, so the next one was just as heavy. On the day it bit, 4 of 11 runs died at the 240 s wall. This tree already enforces the consolidation target from part 1 of this issue and bounds the prompt by bytes, and neither changed that, which fits the point above that the budget bounds the prompt but not the response.

    What we built. The reviewer returns one item per actor it changes:

    {"type":"memory","actor":"principal","op":"edit","drop":[4,19],"rewrite":[{"ids":[5,6,7],"entry":"…"}],"add":["…"]}

    Three choices that turned out to matter:

    • Entries are addressed by the number shown in the prompt, not by matching text. The prompt renders each current entry as [n] …, and a model paraphrases often enough that a match/from string is a second place to be wrong.
    • One rewrite covers replace and merge. One id replaces in place; several ids merge into one entry at the lowest id's position. That makes consolidation expressible directly, and the merged entry is exactly the rewritten entry the consolidation carve-out wants to see.
    • It is materialized in the reviewer, not the writer. The edit is applied to the snapshot the prompt was built from and turned back into the existing op:"set", so MemoryWriter, the erosion guard, the consolidation check and the run-row contract are all untouched. An unknown or reused id is a parse error, which rides the existing corrective re-prompt.

    Results, same install. 10 runs since the change, 10 successful. Every memory item came back as an edit, and responses were 0.3 to 5.5 KB. A 41 → 38 consolidation came back as 5.4 KB where the full list it produced was about 21 KB, and the writer's row matched the edit exactly: 7 entries named, 7 evicted, 4 added.

    What it does not fix. Durations fell less than output did: 166 s for 5.4 KB on a 36.6 KB prompt. With output out of the way, reasoning over the prompt is the larger share, so the prompt-budget design you left this open for still matters. I also raised the wall to 420 s the same day, so to be exact: the longest run since the change took 169 s, which is under the old 240 s.

    Happy to turn this into a PR against main if the shape is the one you want.

  6. IvanLeontev-stack commented on Oct 9, 2026

    @IvanLeontev-stack

    Another install on 7.40.4. MemoryWriter.ts and the reviewer prompt match main @ 5e2f2e8, so the consolidation option from the #1905 closing note is not here either. Below: a one-line prompt change that works inside the existing erosion guard, replay numbers, and one side effect of consolidation worth knowing about whichever fix lands.

    Symptom. The principal file sat at 46–48/48 for a week. The header says CONSOLIDATE BEFORE ADDING but never names a number, and the reviewer swapped entries one for one.

    Change. In renderCurrentMemory, at ≥39 entries the header now reads:

    — ≥80% FULL. Still record every new durable fact, but return EXACTLY ${n - 1} entries, not fewer (a list two or more entries shorter is rejected whole and every new fact in it is lost): for each fact you add, merge two related entries into one or drop the stalest. Merge only entries that carry the same ~tag, and keep that tag
    

    N−1 because one fewer is the most EROSION_LIMIT = 2 lets through in a single write.

    Replay. Saved reviewer prompts with the memory block re-rendered at 39 and 40 entries, both actors, production level: "medium", writer guards simulated. 5 inputs × 2 sizes × 3 repeats = 30 single runs, plus two 10-step chains where each output feeds the next step.

    header runs that touched the file and shrank by one net-dropped 2 → ESUSPECT_EROSION, new facts lost ~explicit merged under ~deduced
    stock CONSOLIDATE BEFORE ADDING (earlier, smaller replay) 0 / 6 — —
    AT MOST N−1 23 / 24 2 / 32 1 / 32
    EXACTLY N−1 + same-tag clause 22 / 24 (the other 2 returned N, nothing lost) 0 / 32 0 / 32

    In the chain the count went 39 → 38, rose to 42 once it was below the threshold (allowed by the prompt), then came back down one per write to 38–39. Zero refusals.

    Side effect that outlives the #1905 carve-out. Under ceiling pressure the model sometimes merged an ~explicit entry with fresh detail from the conversation and tagged the result ~deduced, so a stated fact would later be read as inference. Telling it to merge only same-tag entries stopped that in this sample. A deterministic check is possible (an evicted ~explicit entry whose only successor carries a weaker tag), but I have not built one.

    Limits. Small sample, one install, non-English entry text. Once the consolidation option ships, the "not fewer" clause should go, since a deliberate consolidation will then be allowed to drop more than one entry per write.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions