SoroScan handles a high-volume, write-heavy ingestion workload by continually pulling and indexing block events from the Soroban network. The default PostgreSQL configuration settings are optimized for small-footprint environments and will quickly bottleneck the ingestion pipelines under sustained throughput.
This document describes the recommended database tuning configurations for optimizing write-throughput, smoothing out checkpoint spikes, and speeding up index execution.
SoroScan derives its connection target from process concurrency:
max connections = min(DB_POOL_HARD_LIMIT, WEB_CONCURRENCY × DB_CONNECTIONS_PER_WORKER)
The defaults produce a target range of 2–16 connections for four Gunicorn
workers, and automatically increase as the deployment scales. PostgreSQL must
reserve capacity for migrations, Celery, and administration; keep
DB_POOL_HARD_LIMIT below the database server's max_connections.
| Variable | Default | Purpose |
|---|---|---|
WEB_CONCURRENCY |
4 |
Number of API worker processes |
DB_CONNECTIONS_PER_WORKER |
4 |
Connection budget per worker |
DB_POOL_MIN_SIZE |
2 |
Warm connection target |
DB_POOL_HARD_LIMIT |
40 |
Safety cap during scale-out |
DB_CONNECT_TIMEOUT |
5 |
Seconds allowed to establish a connection |
DB_CONN_MAX_AGE |
300 |
Recycle persistent Django connections after five minutes |
CONN_HEALTH_CHECKS is enabled so recycled connections are checked before
reuse. In larger deployments, point DATABASE_URL at PgBouncer in transaction
mode and apply the calculated maximum there.
Prometheus exports:
soroscan_db_pool_connections{state="total|active|idle|wait_queue"}soroscan_db_pool_configured_limit{limit="min|max"}
Alert when active connections exceed 80% of the configured maximum or when
wait_queue remains non-zero.
For self-managed setups or dedicated database instances (e.g., standard containers/VMs), update your postgresql.conf or cloud parameter group with these settings based on server memory:
| Parameter | 8GB RAM Recommended | 16GB RAM Recommended | Scope / Type | Description |
|---|---|---|---|---|
shared_buffers |
2GB (25%) |
4GB (25%) |
Memory (Requires Restart) | Cache for database data pages. Smooths out write peaks. |
effective_cache_size |
6GB (75%) |
12GB (75%) |
Memory (Reload) | Tells the query planner how much total memory is available for caching. |
work_mem |
64MB |
128MB |
Memory (Reload) | Max memory per sort/hash join operation. Prevents disk spilling during complex queries. |
maintenance_work_mem |
512MB |
1GB |
Memory (Reload) | Speeds up CREATE INDEX and VACUUM processes. |
max_wal_size |
16GB |
32GB |
WAL (Reload) | Max size WAL can grow before forcing checkpoint. |
checkpoint_timeout |
15min |
30min |
WAL (Reload) | Extends time between checkpoints. |
checkpoint_completion_target |
0.9 |
0.9 |
WAL (Reload) | Spreads checkpoint writes over 90% of the timeout window. |
wal_buffers |
64MB |
128MB |
WAL (Requires Restart) | Buffers transaction log data before writing to disk. |
In default configurations, PostgreSQL triggers a checkpoint every 5 minutes or after 1GB of Write-Ahead Logs (max_wal_size) are generated. SoroScan creates massive amounts of WAL records during peak event synchronization loops, causing rapid back-to-back checkpoints that choke disk I/O.
- Raising
max_wal_sizeto16GB(8GB RAM) or32GB(16GB RAM) and extendingcheckpoint_timeoutto15minor30minmakes checkpoints predictable and strictly timed. - Setting
checkpoint_completion_target = 0.9guarantees that instead of flushing data to disk all at once, PostgreSQL trickles the data down smoothly over the length of the timeout.
As tables like ContractEvent approach millions of rows, background data maintenance and index updates require ample working overhead.
- A healthy
maintenance_work_memsetting (512MB/1GB) allows index optimization processes to run entirely inside RAM instead of exhausting disk read/write loops, maximizing the performance of concurrent worker threads.
If your underlying instance uses solid-state storage (such as AWS EBS gp3/io2 or NVMe volumes), ensure you update:
random_page_cost = 1.1
This informs the PostgreSQL engine that random index lookups are virtually as cheap as sequential disk scans, favoring smarter index-driven plans.
Smart contract events emit unstructured payloads that contain dynamic dictionary shapes. SoroScan optimizes search performance on these datasets through a hybrid indexing strategy:
Use standard B-Tree indexes for columns with high-cardinality values, exact match searches, and range filtering:
- Composite Index:
(contract_id, timestamp)allows fast querying of a single contract's chronological event history. - Ingest Index:
(contract_id, ledger, event_index)is marked unique to enforce fast membership checks and upsert validation. - Avoid over-indexing: Avoid adding B-Tree indexes on low-cardinality boolean columns (e.g.
is_active) unless they are part of a composite index.
SoroScan uses GIN (Generalized Inverted Index) to perform lightning-fast lookup operations on key-value pairs stored within event payloads:
- JSONB Inverted Index: An index on the
payloadJSONB column (jsonb_path_opsor default class) enables querying nested sub-attributes without full-table scans. - Usage: Ideal for search queries mapping dot-notation field paths (e.g.
payload__transfer__amount). - Warning: GIN indexes have higher write overhead. For extremely high write workloads, verify that your disk write IOPS can handle the volume, or optimize the set of indexed fields.
SoroScan is write-heavy and continually inserts event records. Under this ingestion rate, table bloat accumulates rapidly as deleted or updated rows (such as those replaced during re-orgs or status updates) leave dead tuples.
To prevent performance degradation, tune the PostgreSQL Autovacuum daemon specifically for SoroScan's heavy workload:
- Increase Autovacuum Concurrency:
autovacuum_max_workers = 4(Allows up to 4 parallel worker threads to run vacuum and analyze tasks simultaneously).
- Smooth Autovacuum Costs:
autovacuum_vacuum_cost_limit = 1000(Increase from default 200 to allow autovacuum to complete faster on large tables before hitting resource throttling boundaries).autovacuum_vacuum_cost_delay = 2ms(Reduce sleep cost to allow workers to clear dead tuples more aggressively).
- Aggressive Ingest Tables Autovacuum Tuning:
For the
ContractEventandWebhookDeliveryLogtables, configure table-level overrides to trigger vacuums more frequently:This triggers vacuuming when 5% of rows have changed, keeping tables clean and preventing database query plans from becoming stale.ALTER TABLE ingest_contractevent SET ( autovacuum_vacuum_scale_factor = 0.05, autovacuum_vacuum_threshold = 1000, autovacuum_analyze_scale_factor = 0.02, autovacuum_analyze_threshold = 500 );