Repository navigation
feat(process): preauthorize inner host integrations from trusted config - #55
Merged
Merged
Conversation
Trusted global configuration now preauthorizes the inner action sandbox: `[permissions] ssh_agent/dbus/gpg_agent = true` puts the matching integration into the ordinary action baseline once its endpoint validates, and `readonly_paths`/`readwrite_paths` were already directly inner-visible. `review_paths`, `deny_paths`, network, and product-private paths keep their per-action gates, project-local configuration still never grants host-wide access, and endpoint validation keeps running per action. Why: requiring a second request for a capability the user had already declared in their own trusted config made the sandbox impractical for agent work without adding any authority the configuration had not granted. Behavior changes in this commit: - Path policy keeps deciding reachability, not announcement. SSH_AUTH_SOCK, DBUS_SESSION_BUS_ADDRESS, and GNUPGHOME stay named even when a deny or review rule masks the mount, so the failure appears at connect time instead of the capability silently disappearing. - One sandbox plan resolves the enabled host-integration bindings once and shares them between mounts and the client environment, replacing the unnamed positional tuple with a typed binding. - Drop the unused PathAccessRuleSource::TrustedGlobalConfigWritableCeiling variant; configured read-write paths are preauthorized directly. - Share the sandboxed re-entry test scaffolding and fail closed when an `--exact` child filter matches no test, which previously passed as an empty run. - Split the permission admission tests by the policy each group owns and document the simplified model in README and examples/config.toml. Verification: cargo fmt --all --check; cargo clippy --all-targets --all-features -- -D warnings; cargo test --all (2083 passed, 0 failed across 58 suites); git diff --check.
Why: the model only requested network after a sandboxed command already failed, and it reached for `grep -r` even when `rg` and `fd` were installed. The prompt named the capability mechanism but never told the model that network is never implicit, so `gh`-style commands were tried first and recovered second. Behavior: - `CODING_AGENT_POLICY_PROMPT` now states that an action starts with no network and decides per command instead of per failure, and that a withheld capability is the first explanation for a credentials, authentication, or connectivity error. - The network schema, the `permissions` object description, the `run_process` description, and the failed-action recovery guidance repeat the same contract where the model reads it. - The coding profile probes the action PATH once per composition and appends `Action PATH search tools:` to the workspace capability summary, so the prompt asks for `rg`/`fd` only when they exist and names the bounded `grep -r`/`find` fallback otherwise. - `action_process_path()` becomes the single owner of the PATH a sandboxed action inherits; the probe and the bubblewrap environment both resolve it there. - The long capability summary moves to `concat!` so each bullet is one reviewable line. Verification: cargo fmt --all --check, cargo clippy --all-targets --all-features -- -D warnings, and cargo test --all (58 suites, 2089 passed, 0 failed) all pass. Six new probe tests cover rg+fd, the `fdfind` distribution name, each missing-tool wording, candidate construction from named PATH directories, and the composed profile summary; contract tests pin the model-visible text. Refs: #55
Why: `test_close_cancels_an_unfinished_run_and_keeps_terminal_result` started a run with an instantly completing provider and closed it immediately. Whether `close()` observed cancellation or an already finished run therefore depended on thread scheduling, so the CI job failed with `COMPLETED` instead of `CANCELLED` on a loaded runner. Evidence: under 64 busy loops on a 32-core host the old formulation reported `completed` in 273 of 300 iterations (first at iteration 2), while the same engine build with a provider that cannot finish on its own reported `cancelled` in 300 of 300. Behavior: - The cancellation test now closes a pending-provider run, which reaches the runtime's cancellation path deterministically and still proves the durable terminal result survives for `result()`. - A new test covers the other half of the contract: a run drained to EOF keeps its completed result after `close()`. - `AgentRun.cancel`, `AgentRun.close`, and the SDK README now say that cancellation is cooperative and that a run which reached its terminal state first keeps that result instead of being rewritten. Verification: uv sync --locked, ruff format --check ., ruff check ., ty check, and pytest tests -q (53 passed) all pass; tests/test_run.py passes three times in a row under the same CPU load that reproduced the failure. uv build succeeds. Refs: #55
10 of 12 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Trusted global configuration now preauthorizes the inner action sandbox.
[permissions] ssh_agent/dbus/gpg_agent = trueputs the matching integration into the ordinary action baseline once its endpoint validates;readonly_paths/readwrite_pathswere already directly inner-visible.review_paths,deny_paths, network, and product-private paths keep their per-action gates.Why: requiring a second request for a capability the user had already declared in their own trusted config made the sandbox impractical for agent work without adding any authority the configuration had not granted.
What changes
MerryConfig::host_integrations()toaction_process_backend_optionstoProcessBackendOptions.host_integrationstoBwrapProcessEnvironment, withprepared_action_process_backend_optionsas the single production entry pointSSH_AUTH_SOCK/DBUS_SESSION_BUS_ADDRESS/GNUPGHOMEstay named even whendeny_pathsorreview_pathsmasks the mount, so a masked endpoint fails at connect time instead of the capability silently disappearingHostIntegrationBindingPathAccessRuleSource::TrustedGlobalConfigWritableCeilingvariant; configured read-write paths are preauthorized directly--exactchild filter matches no test; permission admission tests split by the policy each group ownsexamples/config.toml, model-visible capability text, and permission schema descriptionsBoundaries preserved
review_pathsremain the per-action gate;deny_pathsalways mask and are never reopened by an approval..gitmetadata stays read-only by default with a per-action grant path.Verification
cargo fmt --all --check- cleancargo clippy --all-targets --all-features -- -D warnings- exit 0cargo test --all- 2083 passed, 0 failed across 58 suitesgit diff --check- exit 001ce81eis GPG-signed (%G? = G, signerLocez <locez@locez.com>)Notes and follow-ups