Repository navigation
[Feature]: Periodic agent nudges for long running PTY #51
Description
Activity
Would it be nice, instead of having a nudge parameter, would it be nice if the PTY system actually detects that the current command needs input or needs something, and that's when it actually nudges the agent itself about the current command needing an input or something, where the agent now can read it and respond to it?
So yep, what do you think about this idea? If this idea makes sense to you and if you are willing to make a pull request, it will be nicer.
Well that sounds good in practice. But it doesn't solve the time blindness factor that my original solution presents to fix. For example if a download or an upload is taking unusually long, or a program thats been starting up and the agent has left waiting for it to quit has thrown an error and is not quitting (which isn't that uncommon and can happen in applications that don't do proper error handling, which AI is prone to make) then you only solve one half of the battle.
Not only that, I haven't looked at the technical aspect, but how do you detect if stdin is expected? Can't you enter input to stdin all the time? I don't know, maybe I'm wrong, maybe it is possible. but even then, even if it can be detected reliably, that leaves the other half of the problem completely hanging.
I've thought about this for a long time because its an issue I've frequently encountered with opencode and any other background plugin similar to (and including) opencode-pty, and I cannot think of any other architectural solution. I believe the better direction to discuss would be how the nudge would be done, in a token efficient way, what controls the agent has over it, like; dynamically increasing nudge intervals, active agent control over nudge etc etc.But that's just my personal opinion, if the problem I mentioned in the foregoing can be solved by a different change, then I could evaluate that as well.
Also, I wouldn't mind making a pull request, but as mentioned I want to discuss the details that I'm also not sure about before I send code that'll be rejected, henceforth why I opened this issue.One more thing, I didn't completely invent this idea. As far as I'm aware Windsurf does something similar, which Antigravity also does it (AFAIK its because theyre made by the same team of people or something). Its clearly not a new idea its an existing solution thats already in place, this is something opencode should've had in the first place but I'm not gonna bother fix it underneath there because its much harder to make changes to there than here
Ok, that sounds interesting. Would you be interested in making a PR for this?
I would be interested in making a proof of concept for sure, but there is a lot of things we need to establish first,
- Should input detection also added alongside the nudge mechanism?
- Should nudge be optional or mandatory? Should it be controllable as a parameter within pty_spawn? should it be editable later?
- Should the agent be able to decide to make the next nudge happen sooner or later? And wouldn't this require me to add a new tool?
- Should the agent have control over input detection? There is an inherent question here that a command the agent doesn't think might require inputs, could require inputs, leaving that up to the agent's decision is basically gambling the LLM to predict his own mistake, which if he knew in the beginning wouldve applied a different method, which is why I'm sort of against the "input detection" not even mentioning the platform based issues we could encounter with stdin detection and its reliability etc. In a way this applies to nudge too, but from existing philosophies it seems it would be likely preferred the nudge is kept optional (because notifyOnExit is also optional but the agent mostly needs it)
- Should there be an option for a nudge to occur on substring or regex detection? What are the implications since there are no tools to edit the properties of running PTY instances as far as I'm aware? (could require a new tool), and what about the problem of that substring/regex having to get detected again, what about multiple of them? maybe only detect after a certain offset?
- Should nudges be staticly periodic or dynamic? as in for example increasing nudge intervals, while it sounds like the solution there is the cost benefit problem of KV cache invalidation, especially for something that will be running frequently, the current notifyOnExit does not concern itself about this but I can't leave that question open because this is something periodic. But there's also that static nudges provide more info.
- Should nudges still happen if the agent is actively working?
- Should calling pty_read delay the nudge?
- should the nudge give info on time? should it be saying how much elapsed since last nudge or since idle start?
I could decide these myself, but i'd rather you have a say in them if you prefer.
Problem Statement
It's frustrating when the agent runs long commands and the command gets stuck, or takes longer than expected, with this plugin, the agent isn't notified of anything until the pty finishes, meaning it can hang for a long time by waiting for a non responsive command or a command that is waiting for input.
(Example of this issue can be found in additional context)
Proposed Solution
Similar to the existing notification system that works when a command exits, I suggest a similar approach could be taken for new a nudge parametre within pty_spawn. Something like
nudgeIntervalSecondsto nudge the agent in an interval, or anudgeStartso if the process takes longer than this then the agent is told and it is notified again and again by the interval mentioned beforehand.I haven't established any specifics so it should be up to debate, but the nudge system could be made to only run when the agent is idle (to save tokens)
Alternatives Considered
I've instructed agents to be more strict with timeouts, but timeouts can kill processes that you don't actually want to kill (say for example you're transferring files)
I've also tried instructing the agent via system prompt to be more mindful of the running pty instances, like tell it to manually check by running bash sleep in between pty_read operations. However, this is unstable, the agent occasionally forgets to check and/or checks too often bloating context for no reason. And on top of this, it's very hard to instruct cheaper models with lower prompt adherence to do this, and it's an absolute nightmare getting them to do this, because that is essentially trying to get the agent to do the plugin's job.
Use Case
Monitoring long running tasks: Say for example a heavy database migration or a long install is running, if for example, input is required mid run, or an error happens, if this suggested nudge feature were available, the agent could be woken up to check the status, if there are any errors, it could restart the process, or if an input is required, it could allow the process to continue. Opposed to the only alternative today which is a hard timeout, which could lead to an install/database migration getting corrupted when it couldve been saved by a nudge.
Additional Context
The agent would have waited for a long time unless I intervened.