Skip to content

Promotion-to-voter attempts fail whenever the catch-up log size exceeds 500MiB #910

Description

@jrtorsella

Hi all, I have encountered a set of issues involving promotion/demotion in dqlite that appear to impact fault tolerance. The issue reported here is essentially: a promotion will fail whenever the portion of the log required to catch-up the potential-promotee exceeds MAX_PAYLOAD_LEN, but there is no handling of this requirement on the promoter's side. This has the consequence of making it possible for all nodes except for the leader and voters, including stand-bys, to be unpromotable-to-voter without warning.

I encountered this issue when attempting to administer clusters with sizes larger than 50 nodes (average log entry size is a factor in how large the unpromotability window is and scales significantly with the number of nodes at least in the users of dqlite I have been testing). I've attached a reproducer below which demonstrates this problem with a synthetically-inflated raft log, both in a scenario that shows an unpromotable spare and in one where this problem makes all nodes including standbys unpromotable except for active voters such that any failure cannot promote a voter. I will talk a little bit about how this is showing up in our scaling tests after explaining the mechanism in case there is some concern that the reproducer is producing artificially large log files.

uv_recv.c defines limit macros:

#define MAX_HEADER_LEN (4 * 1024 * 1024) /* 4 MiB */
#define MAX_PAYLOAD_LEN (500 * 1024 * 1024) /* 500 MiB */

in case the message exceeds MAX_PAYLOAD_LEN, the receiver aborts the communication:

dqlite/src/raft/uv_recv.c

Lines 283 to 292 in 9055088

} else if (s->payload.len > MAX_PAYLOAD_LEN) {
tracef("message payload too long: %" PRIu64 " bytes",
(uint64_t)s->payload.len);
if (s->message.type == RAFT_IO_APPEND_ENTRIES) {
raft_free(s->message.append_entries.entries);
s->message.append_entries.entries = NULL;
s->message.append_entries.n_entries = 0;
}
goto abort;
}

The sender places no limit on the catch-up size and does not otherwise handle this case:

/* TODO: implement a limit to the total size of the entries being sent
*/
rv = logAcquire(r->log, next_index, &args->entries, &args->n_entries);
if (rv != 0) {
goto err;
}

and then sends the potentially-too-large catch-up to the server:
req->raft = r;
req->index = args->prev_log_index + 1;
req->entries = args->entries;
req->n = args->n_entries;
req->server_id = server->id;
req->send.data = req;
rv = r->io->send(r->io, &req->send, &message, sendAppendEntriesCb);
if (rv != 0) {
goto err_after_req_alloc;
}

The consequence is that the promotion hangs for max_catch_up_round_duration (50s), during which time role adjustment is frozen and the leader's configuration change slot is held.

This also affects stand-bys: catch-up does not prevent them from being promoted to stand-by from spare, and as stand-bys they silently fail to catch up due to the 500MiB limit such that they cannot be promoted to voter.s.

This is a window of unpromotability that lasts until the last seen sequence number of the unpromotable node is compacted out from the logs, but the window can reopen again any time a node is demoted (for that node).

In real-world scenarios, catch-up logs do reach significantly more than 500MiB. In a microcloud/microovn deployment at 100 nodes, I observed an average log entry size of about 240KiB per entry at the full deployment scale. This issue interacts with a go-dqlite bug which causes every new joiner to self-promote to create a situation where a very large proportion of the cluster becomes unpromotable during an overlapping period of time, during which fault tolerance is correspondingly seriously impaired. In a large enough cluster, any demotion starts a clock essentially: after 500MiB of logs have been written, the node becomes unpromotable until compaction removes the last seen seqno from the log file.

This issue would also affect any situations where the snapshot file itself is larger than 500MiB, although I have not reproduced that and don't know if it's possible - just that there's nothing I can see preventing it. If it is possible it would cover the inverse of the current window essentially, and might make promotion impossible generally except for a small window.

I've attached reproducers that demonstrate this problem synthetically against the go-dqlite demo. Note that this has artificially inflated log entry sizes and that in general never-voter spares are probably not vulnerable to this until cluster sizes of 200+ members (at least in microovn/microcloud clusters), but demoted-voter spares/standbys are starting at a much lower level.

dqlite-issue.tar.gz

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions