Skip to content

DNS-01 orders can exceed hard-coded 2-minute RequestTimeout #64

Description

@rootdet

Summary

DNS-01 certificate issuance can fail when the upstream DNS provider takes longer than two minutes to synchronize or propagate a challenge record because acmeProxyConfig.RequestTimeout() is hard-coded to:

func (c *acmeProxyConfig) RequestTimeout() time.Duration {
    return 2 * time.Minute
}

That timeout wraps the complete external certificate issuance operation. This can expire while the DNS-01 solver is still operating normally.

This is related to, but distinct from, issue #56. Issue #56 covers Lego's separate 30-second post-finalization certificate wait. This issue concerns the outer two-minute request context.

Environment

  • acme-proxy built from current main
  • authority.type: externalcas
  • Upstream CA: Let's Encrypt production
  • DNS provider: ClouDNS through Lego DNS-01
  • Downstream client: Certbot 5.8.0 using HTTP-01 against acme-proxy
  • DNS propagation timeout: 6 minutes
  • DNS propagation polling interval: 10 seconds

Observed behavior

The downstream ACME authorization succeeds quickly. acme-proxy then begins the upstream DNS-01 process.

In one reproducible case:

15:33:59  Upstream certificate request started
15:35:43  DNS provider synchronization reached 8/8
15:35:44  DNS propagation checking began
15:35:54  Lego: Wait for propagation [timeout: 6m0s, interval: 10s]
15:35:59  acme-proxy: context deadline exceeded, downstream order becomes invalid
15:36:00  Let's Encrypt validates the DNS-01 challenge
15:36:02  Let's Encrypt returns the certificate

The upstream certificate was successfully issued only a few seconds after acme-proxy had already invalidated the downstream order.

Representative error:

ERROR async finalization failed ... error creating certificate:
certificate request timed out: context deadline exceeded

Why this is a problem

The DNS-01 solver is configured to allow up to six minutes for propagation, but the parent certificate request context is limited to two minutes.

For DNS providers with slower synchronization or propagation, the outer timeout can therefore expire before the DNS solver's own allowed propagation window has elapsed.

This can cause otherwise-successful orders to be marked invalid and can result in unnecessary issuance retries.

Confirmed workaround

For testing, I changed:

return 2 * time.Minute

to:

return 10 * time.Minute

I rebuilt acme-proxy and repeated the same certificate request. The issuance completed successfully through the same ClouDNS and Let's Encrypt path.

Proposed improvement

Would it make sense to make the external certificate request timeout configurable rather than hard-coded?

For example, conceptually:

{
  "request_timeout": "10m"
}

A configurable value would allow operators using slower DNS providers to select a timeout appropriate for their provider while preserving whatever default the maintainers consider appropriate.

Another option may be ensuring that the outer request timeout cannot expire before the configured DNS propagation timeout.

I would be happy to submit a pull request once there is agreement on the preferred configuration format and default behavior.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions