Skip to content

Latest commit

 

History

History
76 lines (45 loc) · 6.91 KB

File metadata and controls

76 lines (45 loc) · 6.91 KB

Scaling Polis

The ⚡ Running Polis instructions in the main README set up a system with only a single running instance of the server service (Compose:52). For very small engagements of a few hundred participants, this may be sufficient. But for even moderate scale in terms of size and number of concurrent conversations, it will be necessary to run multiple polis-server instances to handle the number of web requests.

Scaling approaches

On a single machine

Compose's replica option uses the current service name: docker compose up --scale server=N. Review the selected overlays, fixed host-port mappings and reverse-proxy routing before using multiple replicas; the root service definition is the source for this topology. Typically, this would be run on a single machine, and thus for even moderate scale requires that the machine being deployed on have room to accomodate running a large number of containers, and an active math worker. This complicates the process of adjusting in real time to changing demand in a cost and resource effective manner.

For small to medium size deployments with rather steady or predictable participation rates, this may be a tenable solution. But for deployments which expect exceptionally high, uneven and/or unpredictable participation rates, it may prove too costly in terms of computational and operations personnel resources.

Docker Compose over Docker Swarm

This section records the historical Docker Swarm exploration. The checked-in Compose file and CDK entry point do not establish a tested Swarm deployment procedure; the current CDK capacity reference is below. Preserve this as context for evaluating a separate deployment topology, not a claim that the Compose development setup is already validated on Swarm.

What's most uncertain at present is whether it will be possible to automatically scale servers based on demand. A local Compose --scale server=N operation is distinct from a Swarm service update. No Docker Cloud or demand-based Swarm autoscaling workflow is validated by the current source.

Note: use the Compose command documented by the selected deployment tooling. The choice of docker compose spelling alone does not establish Swarm compatibility.

Scaling to the limits

For Polis to scale it's most performant potential, you may need to consider additional infrastructure beyond what Docker & Docker Compose provide.

Alternative solutions you might consider:

These are alternative deployment approaches to evaluate against the actual service requirements; the links are historical context, not a statement of current hosting or supported capacity. That having been said, we'd like to be able to support scalable deployments out of the box, and are happy to accept pull requests which get us closer to this goal.

With all that out of the way, deploying a small Polis instance using the docker-compose infrastructure looks more or less like the development environment setup below, with one exception: Instead of running docker compose -f docker-compose.yml -f docker-compose.dev.yml ..., you run docker compose -f docker-compose.yml ... (or simply docker compose, since -f defaults to docker-compose.yml). The development overlay contains development-specific settings, but simply omitting it is not a complete production configuration. Review Makefile profile/overlay selection, TLS, and deployment configuration.

Provisioning compute power for the math worker

Regardless of which method you use above, you'll need to make sure that you provision a large enough node for the math worker to do it's business effectively.

For many simultaneous conversations, increasing the number of cores available will improve overall throughput. For large conversations, you'll also need to consider how much RAM is available to process the data.

Keep in mind that the Polis vote matrix has dimensions p * c, where c is the number of comments, and p is the number of participants, and the amount of memory required to process a conversation grows proportional to this (and computational time increases with higher c). For larger conversations with tens of thousands of participants and thousands of comments, you may need in the dozens of GB of RAM available.

Unfortunately, scaling the size of a worker node is not typically very easy, but this is where more advanced solutions such as Kubernetes could potentially provide additional value.


Current CDK capacity reference

These are checked-in CDK defaults at source snapshot 0985a1a58; they do not establish the capacity currently deployed. Environment and service inputs are in deployment configuration.

Tier Instance type Minimum / desired / maximum Source
Web t3.medium 2 / 2 / 10 instance type, ASG
Clojure math (retired; the Python engine runs on the Delphi tier) r8g.2xlarge 0 / 0 / 0 instance type, ASG
Delphi small r7i.2xlarge 1 / 1 / 7 instance type, ASG
Delphi large c7i.8xlarge 0 / 0 / 3 instance type, ASG
Ollama, when enabled g4dn.xlarge 1 / 1 / 3 instance type, ASG

Ollama infrastructure is off by default via CDK_ENABLE_OLLAMA; see entry point. Delphi CPU target tracking targets 60%; the separate high-CPU alarm threshold is 80%. Those are not symmetric scale-in/scale-out thresholds. See scaling policy and alarms.

The legacy math ASG is capped at one because another instance would repeat work and writes. The Python poller validates explicit shard index/count settings; it is not made safely parallel merely by duplicating the container. See math cap and Python shard validation.

Delphi's worker classification still sends conversations above 5,000 comments to the large class. A large-tier desired capacity of zero therefore matters to routing; see job classification. The import-worker ECS service has desired count zero and separate queue-based scaling rules in import-worker-service.