Skip to content

Pi-on-Herdr: fm-control relaunch permanently refuses after the Pi process dies, even with an fm-recovery idle event applied #2908

Description

@atibus

Reproduced three times in a production home (2026-08-21 twice, 2026-08-23 once). When a Pi agent on the herdr backend exits without its extension closing the turn (fm-control-delivered exit, or an ENOSPC crash printing 'pi exiting due to uncaughtException'), the pane shows a bare shell and pgrep finds no pi process, yet fm-control <id> relaunch fails every attempt with: 'failed while stopping the old agent, which is still running; its original instructions were restored'.

Today's new datapoint: applying the documented recovery reset (fm-busy-event.sh apply <state> <id> idle --current-gen --source fm-recovery --event ...) succeeds and the busy record reads idle, but relaunch STILL refuses identically - so the stop-phase liveness check is not (only) reading the busy record. The same shell-in-pane situation relaunches fine for harness=claude, so the defect appears specific to the pi agent-state classification on herdr. The stop phase also types '/quit' into the bash shell ('bash: quit: command not found'), suggesting it believes a Pi composer is present.

Operator impact: the only recovery is captain-authorized manual WIP commit/push plus forced teardown and a fresh spawn - heavy for a routine crash. Happy to provide the exact state files from the incident on request.

Metadata

Metadata

Assignees

No one assigned

    Labels

    ready-for-prTriage: real bug or VISION-aligned feature, open for a PR

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions