Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

HiGlass Migration

This repo captures the infrastructure, configuration, and process for migrating the higlass.io production service off its original 2020-era single EC2 instance into a hardened, version-controlled setup in a new non-default VPC.

Why this exists

The original instance was set up in 2020 by engineers no longer at the lab, with no IaC and undocumented local changes. This repo exists so that never happens again everything needed to rebuild this service from scratch lives here, in version control, reviewed via pull request.

Scope of this phase

This phase is a redeploy to a non-default VPC, not an architecture change. The service stays on EC2 with Docker Compose, rebuilt on current Ubuntu (24.04) with proper hardening, IMDSv2 enforcement, and SSM-only access (no SSH keys).

A move to ECS/Fargate (with EBS-backed task storage, given the ~4 TB data volume) is a candidate for a later modernization phase, informed by what's learned operating this rebuilt instance. It is explicitly out of scope for this cutover.

Repo layout

higlass-iac/
├── cloudformation/
│   ├── vpc.yaml                          # VPC, subnets, routing, flow logs (Lisa's base template)
│   ├── sns-cert-alerts.yaml              # SNS topic for renewal-failure alerts (deploy before iam/)
│   └── higlass-ec2.yaml                  # t3.xlarge instance, IMDSv2 hop limit 2, SG, /data volume from legacy snapshot
├── compose/
│   ├── docker-compose.yml                # HiGlass (with bundled nginx), Redis, certbot service definitions
│   ├── .env.example                      # template for image tag / SITE_URL — copy to .env, never commit .env
│   ├── init-letsencrypt.sh               # first-run TLS cert bootstrap (recovered from legacy instance)
│   ├── redis-context/
│   │   └── redis.conf                    # recovered production redis config
│   └── web-context/
│       └── sites-enabled/
│           └── hgserver_nginx.conf       # recovered production nginx config (runs inside the higlass container)
├── iam/
│   └── higlass-instance-role.yaml        # Least-privilege IAM role + instance profile (SSM + monitoring only), deployable via CloudFormation
├── docs/
│   ├── runbook.md                        # Step-by-step migration runbook (roles, phases, validation, cutover)
│   ├── ci-cd.md                          # GitHub Actions + OIDC deploy setup, the two AWS roles, what CI does and doesn't deploy
│   └── legacy-findings.md                # What was recovered from the legacy instance's files/git history, and why the repo is structured the way it is
├── .github/
│   └── workflows/
│       ├── validate.yml                  # CI: cfn-lint on templates, docker-compose config validation, nginx syntax check
│       └── deploy.yml                    # CD: manual-only deploy of the SNS, IAM and (plan/deploy) EC2 stacks
├── params/                               # per-stack CloudFormation parameters used by deploy.yml
└── .gitignore

Ownership

  • Infrastructure (VPC, IAM, security groups): Lisa Choy
  • Security review, hardening validation, cutover checkpoints: Kelly Briceno

Open items before first deploy

See docs/legacy-findings.md for the full detail behind these:

  • Fix TLS renewal to actually rely on the certbot container's built-in loop (confirmed by Lisa: the legacy quarterly cron job did NOT successfully renew certs — renewal has been manual) — add real alerting on renewal failure, not a silent log. Done in config: failed certbot renew publishes to SNS; higlass reloads nginx every 12h so renewed certs are actually served.
  • Set up the /data mount via /etc/fstab or user-data on the new instance, not a manual mount command (cloudformation/higlass-ec2.yaml user data: UUID-based fstab entry, Docker waits for /data)
  • higlass/higlass-docker:0.10.4 is confirmed the newest published tag — no upstream upgrade path exists. Decide: accept as-is (lean on network controls), or patch known-vulnerable dependencies (Django, Pillow) directly in the Dockerfile
  • Layer nginx hardening (HSTS, server_tokens off, HTTP→HTTPS redirect) onto the recovered config without breaking its working location blocks
    • After the post-cutover soak, raise HSTS max-age from 86400 to 31536000

See docs/runbook.md for the full migration sequence, validation checklist, and rollback plan.

About

Infrastructure as Code

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages