Skip to content

Add Polars ingestion pipeline for parking citations - #756

Merged
gregpawin merged 4 commits into
hackforla:mainfrom
gregpawin:feat/beta-pipeline
Sep 16, 2026
Merged

gregpawin merged 4 commits into
hackforla:mainfrom
gregpawin:feat/beta-pipeline

Conversation

@gregpawin

Copy link
Copy Markdown
Member

Description

Adds data-science/beta_pipeline/, a Polars-based ingestion pipeline that reads the raw LA Parking Citations export,
cleans it, and loads it into SQLite or PostGIS.

  • parking_clean.py — normalization and type coercion, including dropping records too incomplete to be analytically
    useful rather than carrying nulls downstream.
  • parking_db.py / parking_postgis.py — schema creation and loading for the two targets.
  • parking_pipeline.py — orchestration entry point tying the stages together.
  • ARCHITECTURE.md — module-level reference; README.md — quick start.
  • parking_db_explore.ipynb and violation_analysis.ipynb — the exploratory work the violation-code and column
    mappings were derived from.

Polars was chosen over pandas because the raw export is ~6 GB and does not fit comfortably in memory with pandas'
row-oriented overhead.

Also ignores .venv/ and raw_data/ at the repo root, so the multi-GB raw CSV can be staged locally without becoming
committable.

Related Issues

Refs #695

This is the ingestion/ETL half of #695. Measured against its acceptance criteria:

  • ETL/ELT pipelines automated and scheduled — partial. The pipeline runs end to end from one entry point, but
    scheduling is not implemented.
  • Incremental load support and retry mechanisms — not addressed. Loads are full-refresh with no retry logic.
  • Pipeline monitoring dashboard — not addressed.

I would suggest treating scheduling, incremental loading, and monitoring as follow-up issues rather than expanding this
PR, since each needs a decision this PR should not make unilaterally — in particular which orchestrator the project
wants to commit to.

Testing

Exercised manually against the full ~6 GB citations export and against truncated samples, loading into both SQLite and
PostGIS, with row counts and null rates compared before and after cleaning.

This code has no automated test suite yet, which is the main thing I would flag for review. It predates the tested
postgis_db work and I did not want to bundle a test-writing effort into this PR. If you would prefer tests before this
lands, say so and I will add them.

Checklist

  • I have followed all conventions outlined in our documentation.
  • I have fully tested my changes and confirmed that all new and existing tests pass.
  • I have written meaningful commit messages for all changes.
  • I have linted and formatted my changes to follow the code style of this repository.
  • I have updated our documentation, accordingly.
  • I have checked currently opened pull requests to ensure that there are no pending pull request for the same
    changes or issues.
  • I have confirmed that this pull request fully meets the acceptance criteria for all related issues listed above.

Unchecked, deliberately: there is no automated test coverage here (see Testing), Python is outside the repo's
Prettier/ESLint scope, and #695's criteria are only partially met (see Related Issues).

gregpawin and others added 4 commits September 14, 2026 16:47
…or parking citations

- Introduced `ARCHITECTURE.md` to detail project structure, data flow, and key functions.
- Added `parking_pipeline.py` to orchestrate the loading of raw citations and rebuilding of the cleaned table.
- Created `parking_clean.py` for managing the cleaned citations table.
- Updated permissions for several scripts and configuration files.
- Enhanced `parking_postgis.py` to support rebuilding the cleaned table after loading CSV data.
Pin the beta_pipeline requirements and add the violation-code analysis
notebook used to derive the reference mappings.

Ignore .venv/ and raw_data/ at the repo root: the pipeline expects the
raw citations CSV to be staged locally and it is multi-GB, so it must
never be committable.

Co-authored-by: Cursor <cursoragent@cursor.com>
- Implemented `drop_incomplete` function in `parking_clean.py` to remove rows missing `issue_datetime` or `loc_lat`.
- Updated `rebuild_clean` to call `drop_incomplete` after datetime creation.
- Modified `parking_pipeline.py` to include dropping incomplete rows in the pipeline process.
- Adjusted `parking_postgis.py` to ensure `drop_incomplete` is executed after rebuilding the cleaned table.
@github-actions

Copy link
Copy Markdown

@gregpawin, this Pull Request is not linked to a valid issue. Please provide a valid linked issue in "Related Issues" above, using the format of "Resolves #" + issue number.

@gregpawin
gregpawin merged commit aad9a0b into hackforla:main Sep 16, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants