Skip to content

Set up ingestion framework and ETL #695

Description

@glenflorendo

User Story

No response

Description

Deploy the ingestion pipelines to capture external and internal data from application databases and APIs.

Acceptance Criteria

  • ETL/ELT pipelines automated and scheduled
  • Incremental load support and retry mechanisms
  • Pipeline monitoring dashboard

References (Design, Technical, etc.)

No response

Additional Information

No response

Activity

  1. gregpawin commented on May 19, 2026

    @gregpawin
    Member

    Working on early setup for pipeline as Dockerized database with Python toolset.
    Link to my fork's code

  2. gregpawin commented on Jun 9, 2026

    @gregpawin
    Member

    Added a data cleaning pipeline to my db updating code: Link to my fork's code.

  3. gregpawin commented on Jul 22, 2026

    @gregpawin
    Member

    New PostGIS Docker setup being worked on here: link to my fork's code. Contains neighborhood council, zipcodes, and complete dataset--work in progress.

  4. gregpawin commented on Jul 28, 2026

    @gregpawin
    Member

    Enhance data contract and loading scripts for boundary polygons and place points

    Updated datacontract.yaml to include new valid values: "Neighborhood" and "City Council District".
    Modified Dockerfile to load additional GeoJSON files for council districts, neighborhoods, and places.
    Refactored 02_load_boundaries.sh to streamline loading of boundary polygons and place points, adding functions for better code organization and including new data sources.
    link to my fork's code

  5. gregpawin commented on Aug 4, 2026

    @gregpawin
    Member

    Enhance PostGIS database setup and API for parking citations

    • Updated docker-compose.yml to include SKIP_CITATIONS_LOAD environment variable for controlling CSV loading on first boot.
    • Modified Dockerfile to install additional Python dependencies and include new scripts for loading contract citations.
    • Added 03_load_citations.sh script to handle loading of parking citations from CSV files.
    • Introduced FastAPI application in api/main.py for querying parking citation data.
    • Created query_contract.py CLI for executing data-contract queries against the PostGIS database.
    • Implemented error handling and validation in the new API and CLI.
    • Added tests for FastAPI routes to ensure functionality and reliability.
      link to my fork's code
  6. gregpawin commented on Aug 11, 2026

    @gregpawin
    Member

    Citation explorer UI

    • Built/expanded the local web UI (web_sheet) for querying citations by region type, region, and date range
    • Added live region autocomplete (top 5 alphabetical matches, including empty-box focus)
    • Added Chart (sheet) vs Map view, with Leaflet markers on citation coordinates (Leaflet vendored locally after a CDN/SRI load issue)
    • Added Single-region / Compare mode toggle, with Region 1 + Region 2 inputs and side-by-side sheet/map results

    Boundaries & Docker

    • Made boundary loading portable (load_boundaries.sh, normalize_boundaries.sql, check/reload helpers)
    • Wired first-boot Docker init to load all boundary tables (councils, zips, districts, neighborhoods, places)
    • Reloaded missing layers on the existing DB volume (neighborhoods / districts / places had been absent)

    Docs

    • Added a Quick start section and updated README for the explorer UI, boundary reload, and portable scripts

    link to my fork's code

  7. gregpawin commented on Aug 25, 2026

    @gregpawin
    Member

    Add production configuration and environment setup for PostGIS and API

    • Introduced docker-compose.prod.yml for production deployment of PostGIS, API, and explorer UI.
    • Created .env.example to provide a template for environment variables required in production.
    • Updated Dockerfile and Dockerfile.api to support production builds with necessary dependencies.
    • Enhanced README.md with detailed instructions for building and deploying the application on a VPS.
    • Added deployment documentation for IONOS and general VPS setups, including steps for database initialization and restoration.
    • Implemented health checks for services to ensure proper startup and readiness.
    • Removed obsolete postgisdb.zip file to streamline the project structure

    link to my fork's code

  8. gregpawin commented on Sep 1, 2026

    @gregpawin
    Member
    • Enhanced the "Quick start" section with a direct link to the Parking Citations dataset on the LA Open Data Portal.
    • Revised instructions for downloading the citations CSV to specify the required format and source for clarity.
    • Update healthcheck command in docker-compose.yml to ensure compatibility with host bind-mounts

    link to my fork's code

  9. gregpawin commented on Sep 14, 2026

    @gregpawin
    Member

    I've opened the work for this as three PRs rather than one, since it turned into three separable pieces and bundling them
    would have meant one reviewer evaluating a Docker/PostGIS deployment, a Polars pipeline, and a documentation set
    together:

    All three are additive only — no deletions, nothing existing rewritten — and CI is green on each.

    Where this leaves the acceptance criteria

    I want to be straightforward that this does not close the issue:

    • ETL/ELT pipelines automated and scheduled — partially done. Loading is automated and repeatable (one command, or
      automatically on container start), but nothing schedules it. There's no cron, Airflow, or Prefect component.
    • Incremental load support and retry mechanisms — not done. The loader does a full refresh via a staging table and an
      atomic swap, so a failed load never leaves the served table in a mixed state, but there's no watermark, no upsert, and
      no retry logic.
    • Pipeline monitoring dashboard — not done. The explorer UI queries citation data; it doesn't report on load runs or
      pipeline health.

    So roughly: the destination, the schema, the data contract, and a working bulk load are in place; the orchestration
    half of this issue is not.

    What I'd suggest

    Scheduling, incremental loading, and monitoring each need a decision I don't think I should make alone — particularly
    which orchestrator the project wants to commit to, since that choice is hard to reverse and affects deployment. Happy to
    split them into three follow-up issues, or to keep this one open and scope it down to just the orchestrator decision,
    whichever you prefer.

    Two smaller questions while I'm here:

    1. Add architecture and contributor documentation #758 is documentation and has no issue of its own. Should I open one, or does contributor documentation belong in the
      wiki by convention? Happy to close it if so.
    2. docs/CONTRIBUTING.md says to branch from and target dev, but that branch doesn't exist — the remote branches are
      main, stable, citation-analysis, fixes, 710/node-v24, and stale/cicd. I targeted main for all three. If
      that's wrong I can retarget. Might be worth updating the guide either way.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions