This repo captures the infrastructure, configuration, and process for migrating the
higlass.io production service off its original 2020-era single EC2 instance into a
hardened, version-controlled setup in a new non-default VPC.
The original instance was set up in 2020 by engineers no longer at the lab, with no IaC and undocumented local changes. This repo exists so that never happens again everything needed to rebuild this service from scratch lives here, in version control, reviewed via pull request.
This phase is a redeploy to a non-default VPC, not an architecture change. The service stays on EC2 with Docker Compose, rebuilt on current Ubuntu (24.04) with proper hardening, IMDSv2 enforcement, and SSM-only access (no SSH keys).
A move to ECS/Fargate (with EBS-backed task storage, given the ~4 TB data volume) is a candidate for a later modernization phase, informed by what's learned operating this rebuilt instance. It is explicitly out of scope for this cutover.
higlass-iac/
├── cloudformation/
│ ├── vpc.yaml # VPC, subnets, routing, flow logs (Lisa's base template)
│ ├── sns-cert-alerts.yaml # SNS topic for renewal-failure alerts (deploy before iam/)
│ └── higlass-ec2.yaml # t3.xlarge instance, IMDSv2 hop limit 2, SG, /data volume from legacy snapshot
├── compose/
│ ├── docker-compose.yml # HiGlass (with bundled nginx), Redis, certbot service definitions
│ ├── .env.example # template for image tag / SITE_URL — copy to .env, never commit .env
│ ├── init-letsencrypt.sh # first-run TLS cert bootstrap (recovered from legacy instance)
│ ├── redis-context/
│ │ └── redis.conf # recovered production redis config
│ └── web-context/
│ └── sites-enabled/
│ └── hgserver_nginx.conf # recovered production nginx config (runs inside the higlass container)
├── iam/
│ └── higlass-instance-role.yaml # Least-privilege IAM role + instance profile (SSM + monitoring only), deployable via CloudFormation
├── docs/
│ ├── runbook.md # Step-by-step migration runbook (roles, phases, validation, cutover)
│ ├── ci-cd.md # GitHub Actions + OIDC deploy setup, the two AWS roles, what CI does and doesn't deploy
│ └── legacy-findings.md # What was recovered from the legacy instance's files/git history, and why the repo is structured the way it is
├── .github/
│ └── workflows/
│ ├── validate.yml # CI: cfn-lint on templates, docker-compose config validation, nginx syntax check
│ └── deploy.yml # CD: manual-only deploy of the SNS, IAM and (plan/deploy) EC2 stacks
├── params/ # per-stack CloudFormation parameters used by deploy.yml
└── .gitignore
- Infrastructure (VPC, IAM, security groups): Lisa Choy
- Security review, hardening validation, cutover checkpoints: Kelly Briceno
See docs/legacy-findings.md for the full detail behind these:
- Fix TLS renewal to actually rely on the
certbotcontainer's built-in loop (confirmed by Lisa: the legacy quarterly cron job did NOT successfully renew certs — renewal has been manual) — add real alerting on renewal failure, not a silent log. Done in config: failedcertbot renewpublishes to SNS;higlassreloads nginx every 12h so renewed certs are actually served. - Set up the
/datamount via/etc/fstabor user-data on the new instance, not a manualmountcommand (cloudformation/higlass-ec2.yamluser data: UUID-based fstab entry, Docker waits for/data) -
higlass/higlass-docker:0.10.4is confirmed the newest published tag — no upstream upgrade path exists. Decide: accept as-is (lean on network controls), or patch known-vulnerable dependencies (Django, Pillow) directly in the Dockerfile - Layer nginx hardening (HSTS,
server_tokens off, HTTP→HTTPS redirect) onto the recovered config without breaking its working location blocks- After the post-cutover soak, raise HSTS
max-agefrom 86400 to 31536000
- After the post-cutover soak, raise HSTS
See docs/runbook.md for the full migration sequence, validation checklist, and
rollback plan.