Skip to content

v0.10.0: unattended operations - #15

Merged
renezander030 merged 1 commit into
masterfrom
release/v0.10.0
Oct 8, 2026
Merged

renezander030 merged 1 commit into
masterfrom
release/v0.10.0

Conversation

@renezander030

Copy link
Copy Markdown
Owner

v0.10.0: unattended operations

Calendar schedules, a preflight, a clean shutdown, spend warnings and credentials that stay out of every log: the settings that matter once Draftcat runs as a service.

Changes, in review order

# Change Why it matters
1 draftcat doctor: read-only preflight of config, credentials, operator access, state store, listener ports and schedules. Gives a fix for each finding, supports --json, and exits 1 on failure Every install and every deploy; can gate a restart (draftcat doctor && systemctl restart draftcat)
2 Shutdown drain: SIGINT/SIGTERM refuses new runs, closes the webhook listener and waits up to timeouts.shutdown_grace (default 30s) Every redeploy; approval taps keep arriving while it waits
3 Credential redaction: the values of all configured secret variables become [REDACTED:<VARIABLE>] in the log, operator notifications, stored run and webhook errors, JSON spans and OTLP exports Credentials never leave the engine through its own output, including transport errors that quote request URLs
4 pause_after_failures: a timer pipeline that fails N runs in a row is paused and the operator is told once. The streak is restored at start; /cron resume clears it A broken upstream stops costing budget and alerts
5 Calendar schedules: five-field cron and @hourly/@daily/@weekly/@monthly in a per-pipeline timezone, DST-aware, with optional catch_up Business-hours pipelines (weekday 08:00 digests, follow-ups)
6 Schedules continue across restarts: intervals resume from the recorded run history A 24h pipeline keeps its cadence on machines that restart often
7 budgets.alert_at: one notification per threshold, cap and UTC day, persisted across restarts Act before the daily cap closes the gate
8 Current-state gauges on /metrics: running and paused pipelines, failure streaks, open approvals, today's tokens and spend against the caps Dashboards and alerting on state, not only totals
9 One claim for every run: timer, /run, Run-now and webhook all go through the same atomic claim, so each pipeline has at most one run in progress /run while a run is in progress answers Not started (pipeline is already running)

/cron set accepts cron expressions; /cron shows full next and last run times. Validation covers schedules, time zones, pause_after_failures, alert_at and shutdown_grace. New guide: docs/operations.md.

Compatibility

All new settings are optional. Configurations without them behave as before, except that interval pipelines continue from their last recorded run, shutdown waits up to 30s for running pipelines (shutdown_grace: 0s exits at once), and /run of a pipeline that is already running is refused. Versions are synchronized to 0.10.0 in version.go, package.json and package-lock.json.

Verification

  • go build ./..., go vet ./..., gofmt: clean
  • go test -short -count=1 ./...: 16 packages ok; 451 top-level tests pass (399 on master), 0 failures
  • go test -short -tags voice -count=1 ./...: ok
  • go test -race -count=1 ./...: ok
  • golangci-lint run --new-from-rev=master (v2.5.0, repo config): 0 issues
  • go run . validate: 0 errors
  • npm test: 3/3 pass; npm pack --dry-run: draftcat-0.10.0.tgz, 5 files
  • Local smoke with a built binary: doctor exits 0 with the next cron run shown in its zone; the engine starts, a webhook trigger returns 202 and the run completes, SIGTERM exits cleanly, and no configured secret appears in the engine log

New tests cover cron parsing and DST, interval restore, catch-up, the claim (manual vs. timer vs. webhook, drain refusal), the auto-pause streak and restart restore, drain with a live listener and grace expiry, alert thresholds (once, highest-of-jump, persisted across a store reopen), gauges, redaction of the Telegram transport error, notices, spans and run errors, and doctor (ready, missing credentials, read-only store inspection, JSON, busy port).

Review notes

  • Base: master @ 5986efa. No other open PRs.
  • Touched: scheduler (moved to scheduler.go), engine start/stop (engine_lifecycle.go), budget ledger hook (budget_alerts.go), Telegram/relay notification text, obs span/OTLP writers and /metrics, validation, config structs.
  • New packages: internal/schedule, internal/redact.
  • Untouched: approval gates and receipts, tool gate, webhook admission and signatures, the hitl/v0 protocol, ZK/FHE commands, the state schema (alert marks use the existing dedup table).
  • Approval drafts are shown unredacted so the payload hash covers exactly what the operator approves; the guide points to model_policy for holding drafts that contain credentials.

- draftcat doctor: read-only preflight of config, credentials, operator
  access, state store, listener ports and schedules, with --json.
- Cron expressions and @daily-style schedules in a per-pipeline time zone,
  with optional catch_up for slots missed during downtime.
- Interval schedules continue from the recorded run history.
- One claim admits every run (timer, /run, Run-now, webhook).
- pause_after_failures pauses a failing timer pipeline and notifies once.
- SIGINT/SIGTERM drain within timeouts.shutdown_grace.
- budgets.alert_at notifies at fractions of the daily caps.
- Configured credentials are redacted from logs, notifications, stored
  errors, spans and OTLP exports.
- Current-state Prometheus gauges.
@renezander030
renezander030 merged commit 6f6c5aa into master Oct 8, 2026
1 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant