Skip to content

Add authenticated MCP Events webhook support - #17

Open
quinnj wants to merge 22 commits into
mainfrom
feat/mcp-events
Open

quinnj wants to merge 22 commits into
mainfrom
feat/mcp-events

Conversation

@quinnj

@quinnj quinnj commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

MCP clients can now discover event definitions, create authenticated webhook subscriptions, and receive signed occurrences when a Julia application emits an event. This implements the webhook profile described by OpenAI's MCP Events guide and the experimental MCP Events proposal on MCP 2026-07-28.

The server requires explicit identity, authorization, storage, and filtering hooks. Subscription identity includes the authenticated owner, exact callback URL, event name, and canonical arguments. Refreshes preserve identity and rotate signing keys; unsubscribe waits for a delivery already in flight and stops later retries. File-backed subscriptions survive process restarts through private files, atomic replacement, and fsync.

Callbacks must pass a signed challenge before activation; a live subscription from the same owner to the same URL counts as consent for refreshes and new filters. Delivery rechecks access and expiry, validates every DNS answer, pins the connection address while preserving TLS hostname verification, and disables redirects, proxies, and netrc credentials. Standard Webhooks HMAC signatures cover the exact body bytes, including fresh timestamps on bounded retries. Signing secrets are redacted from verbose logs, and callback errors expose only server-generated categories. Fuzzing also drove fixes in shared JSON-RPC envelope validation and malformed-parameter handling.

Scope and decisions

  • Event APIs remain namespaced. Existing stateful MCP support and the static tools server are preserved.
  • Webhook delivery is emit-only and synchronous: emit_event! validates every body first, then delivers each owner's subscriptions in order and different owners concurrently (16 at a time), and returns when all deliveries finish. The application owns its reliable queue or outbox.
  • Default subscriptions last 30 minutes, capped at one day. ttlMs: null gets the cap; other integers are clamped, never rejected. maxAgeMs is ignored because replay is not offered.
  • Callback verification and its cooldown are scoped per owner, so users who share a hosted receiver cannot block each other. Each owner verifies one callback at a time.
  • Polling, push streams, replay, dynamic event catalogs, no-expiry grants, and gap/termination envelopes are not advertised.
  • JSONSchema.jl validates drafts 4, 6, and 7. Omitted dialects are explicitly published as draft 7; JSON Schema 2020-12 validation and external schema references are unsupported.
  • The file store has one process owner. Replicas need a shared transactional backend and distributed coordination. System DNS resolution can outlast the request timeout; the deadline is checked again before connecting.

Review round

A Claude review, followed by four Codex cross-review rounds (the last returned VERDICT: CLEAN), found and fixed these. Each fix has a regression test that fails on the earlier code:

  • Security: webhook requests sent the server's ~/.netrc credentials to subscriber-chosen hosts; every delivery leaked an idle TLS connection for 30 seconds.
  • Cross-tenant: one user's failed verification blocked every user's subscriptions to a shared receiver host; one user could fill the verification cache and lock others out; the global events lock was held while a slow delivery finished, stalling other users' requests; one dead endpoint delayed every other subscriber's delivery by about 42 seconds per emission.
  • Races: retries could reach a subscription re-created after unsubscribe; a stale revocation check could delete a refreshed subscription, even an identical one; unsubscribe could wait on a replacement subscription's deliveries; concurrent subscribes could send redundant or over-quota challenges; a subscribe that waited for a turn or a slow store could save after access was revoked.
  • Spec: TTL values are clamped rather than rejected; maxAgeMs is ignored; events no longer appear in the 2025-11-25 discovery manifest.
  • Robustness: sub-millisecond deadlines no longer disable libcurl's timeout; the cooldown applies however the host is spelled; the store syncs the file and directory and checks its directory at startup; a transform error sends nothing instead of a partial emission; a Windows-only test leaked a dead proxy into the trim test.

Validation

  • CI passes all 11 checks on 61134e4: Julia 1.10, current, and nightly on Linux, macOS, and Windows, plus the documentation build and preview.
  • Full Pkg.test() passes locally on Julia 1.12.6 with HTTP.jl 2.8 and with HTTP.jl 1.11, including the JuliaC --trim=safe check.
  • End-to-end interop against the official Python standardwebhooks library, over real HTTPS with address pinning: challenge echo, verified deliveries, key rotation with dual signatures, filtering, and unsubscribe.
  • A multithreaded stress run (4 threads, about 1,700 deliveries per run with retries, concurrent refresh, unsubscribe, and emit) found no delivery after unsubscribe returned and no leaked in-flight or verification state.
  • The seeded request, identity, signature, and lifecycle fuzzers from the original PR still pass. Codex's review probes (180+ concurrency, lifecycle, persistence, and transport checks) pass on Julia 1.12 and 1.10.
  • Documenter build and doctests pass.
  • Live ChatGPT plugin delivery has not been exercised.

Co-authored by Codex

🤖 Generated with Claude Code

quinnj and others added 22 commits September 30, 2026 00:10
Every delivery built a fresh Downloader with the default 30-second grace
period, so each POST left its TLS connection open for 30 seconds. A busy
server would run out of file descriptors. A zero grace period releases the
connection as soon as the request finishes; the pinned connection is never
reused anyway.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Four cross-tenant problems, each reproduced before the change:

- A failed callback verification put the whole callback host on cooldown.
  Hosted clients such as ChatGPT use one receiver host for every user, so
  any authenticated user could block everyone's new subscriptions by
  subscribing a bad URL every few seconds.
- The verification cache rejected new verifications once it held
  max_subscriptions entries. One user could fill it by verifying URLs they
  control, then unsubscribing, locking out every other user for minutes.
- Refreshing a live subscription re-sent the challenge once the 5-minute
  verification cache expired, so refreshes depended on the endpoint
  answering again and on no cooldown being active.
- The global events lock was held while waiting for a subscription's
  delivery lock, which a delivery holds across its network request. One
  slow endpoint plus its owner's unsubscribe stalled everyone's
  events/list, events/subscribe, and emit_event! calls.

Verification is now rate limited per (principal, host), runs outside any
lock, and treats a live subscription to the same URL as consent, as the
draft allows for refreshes and new arguments. Subscription records are
plain data, so custom stores no longer have to preserve a process-local
lock. One condition-backed lock guards definitions, store writes, and
in-flight counts, and is never held across network I/O or application
hooks. Unsubscribe deletes the record, then waits only for an attempt
that already started, so no delivery reaches the endpoint after it
returns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
emit_event! delivered one subscription at a time, sorted by ID. A dead
callback costs up to four attempts at the request timeout plus backoff,
about 42 seconds with the defaults, and every subscriber sorted after it
waited that long. A user holding many subscriptions to a slow endpoint
could delay everyone else's events by minutes per emission. A transform
that produced an invalid payload for one subscriber also threw only after
the subscribers sorted before it had already been sent the event.

emit_event! now builds and validates every body before sending any, then
delivers each owner's subscriptions in order and different owners on a
pool of 16 tasks. It still returns only after every delivery finishes,
and a failing task no longer stops other owners' deliveries.

The test receiver now answers 401 to a bad signature instead of calling
@test, because deliveries run on worker tasks where @test results do not
reach the enclosing testset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ttlMs: null asks for no expiry. The server never grants that, so it now
  returns the longest finite lifetime rather than the short default.
  Oversized integers are clamped instead of rejected, because the draft
  says TTL values are clamped, never rejected.
- maxAgeMs only bounds replay, and emit-only events never replay, so the
  draft says to ignore it. Validating it rejected clients that send null.
- The request-ID checks in dispatch_events could not fail: the envelope
  parser and the notification check already enforce them.
- Verbose logs no longer label an empty body as invalid JSON-RPC.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The store writes a temporary file and renames it over the old one. Without
an fsync, a power loss soon after an update can leave an empty file on some
filesystems, and the store deliberately refuses to load invalid state, so
the server would not start until someone removed the file. Flushing the
bytes first leaves either the old or the new file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Describe consent from a live subscription and the per-subject cooldown,
concurrent per-owner delivery with up-front validation, the null-TTL
grant, and ignored maxAgeMs. Drop the record-lock rule for custom stores.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A missing directory was detected only on the first write, so a client's
first subscribe failed with -32602 Invalid params for a server setup
error. The store constructor now rejects a missing directory at startup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The PR removed events from the 2025-11-25 initialize result, but the
discovery manifest, which also describes the 2025-11-25 endpoint, still
advertised them. Both now use one version-aware helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cross-review found two races that the lock rework introduced, where the
original per-record lock had prevented them:

- A retry still in backoff reached a subscription that was unsubscribed and
  created again with the same key, even though the new subscription asked to
  start from now. Each record now carries a random instance token that a
  refresh keeps and a new subscription replaces; retries stop when it
  changes.
- An access check that returned "revoked" after access was restored and the
  subscription refreshed deleted the refreshed subscription, so the client
  believed it was subscribed and got nothing until its next refresh. Revoked
  records are now deleted only if unchanged since the check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cooldown key used the host exactly as written, so one principal could
send a challenge to the same destination under Receiver.example,
receiver.example., or another spelling of an IPv6 literal without waiting.
The key now ignores case and trailing dots and uses the canonical IPv6 form.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Downloads.jl enables libcurl's netrc lookup on every request. A webhook POST
therefore carried an Authorization header from the server account's
~/.netrc whenever it had an entry for the callback host, or any host with a
default entry. Any authenticated user could collect those credentials by
subscribing with a URL they control. The sender now disables netrc. Found
by the Codex cross-review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Downloads rounds the timeout to whole milliseconds, and libcurl treats zero
as no limit. When a delivery's remaining deadline was under half a
millisecond, the request could therefore wait forever, and an unsubscribe
would wait with it. The sender now passes at least one millisecond. Found by
the Codex cross-review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
In-flight attempts were counted per subscription ID. If the same key was
subscribed again while an unsubscribe was draining, the unsubscribe also
waited for deliveries to the new subscription. Attempts are now counted per
instance, and unsubscribe waits only for those running when it deleted the
record. Found by the Codex cross-review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The draft says a clamped grant is not an error and that TTL values have no
rejection path. ttlMs of 0 or -1 now gets the one-millisecond minimum;
non-integer values are still invalid params. Found by the Codex
cross-review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The default sender now rounds a tiny timeout up, but a replacement request
hook could still receive one. The delivery path now stops before calling any
transport when less than a millisecond of its deadline remains. Suggested
by the Codex cross-review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two problems found by the Codex cross-review:

- The subscription quota was checked against saved subscriptions only, so
  concurrent requests from one principal to different hosts each sent a
  challenge. With a quota of one, eight requests sent eight.
- A request waiting on another request's verification of the same URL woke
  before that subscription was saved, missed the new consent, and sent a
  second challenge. If the endpoint then failed, the second subscribe failed
  for no reason.

Each principal now runs one verification at a time and keeps its turn until
its subscription is saved. Waiting requests recheck the quota and existing
consent when they wake, so they fail on quota or reuse the verification
without sending a challenge. Other principals are not affected, and the
per-host cooldown after a failure is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Syncing the temporary file keeps its bytes safe, but the rename that
publishes it lives in the directory. After a power loss the old file could
come back, dropping a new subscription or reviving an unsubscribed one. The
store now syncs the directory after the rename. This is best effort: some
filesystems cannot sync a directory, and failing there would make every
update error after its file was already in place. Found by the Codex
cross-review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The test set https_proxy and HTTPS_PROXY with withenv to prove that
deliveries ignore proxies. Windows environment names ignore case, so both
keys name one variable. withenv saved the first value as the second key's
previous value and restored it after deleting the first, leaving
HTTPS_PROXY=http://127.0.0.1:1 set. The later JuliaC trim test's Pkg
subprocess inherited it and failed to reach pkg.julialang.org and GitHub
("via 127.0.0.1") whenever it had to download a package. Windows now sets a
single variable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A subscribe that waited for the same principal's verification woke up,
found the new subscription's consent, and skipped verification, and with
it the access recheck. If access to its filters was revoked during the wait,
it saved an unauthorized subscription and returned success. Deliveries
recheck access, so no data leaked, but the client was told the subscribe
succeeded. Access is now rechecked whenever the request waited or verified.
Found by the Codex cross-review, round 2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
b7b9111 rechecked access only when a subscribe had verified a callback or
waited for its principal's turn. A request can also stall on a slow store
(file fsync, a database) or on lock contention while neither is true, and a
revocation in that window still produced a saved subscription. Access is now
checked again just before every save. It costs one extra call to the
application's authorize hook per refresh, and removes the waited flag.
Found by the Codex cross-review, round 3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A stale revocation check deleted a record only if it still equalled the
checked one field for field. A refresh with the same secret and TTL within
one clock tick (a whole-second custom clock, for example) produced an
identical record, so an older denial still deleted a refresh that had just
succeeded. Each save now stamps a random revision, and a stale check deletes
only if the revision is unchanged. The lifetime instance token still survives
refreshes. Found by the Codex cross-review, round 3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant