Repository navigation
Add Polars ingestion pipeline for parking citations - #756
Merged
Merged
Conversation
…or parking citations - Introduced `ARCHITECTURE.md` to detail project structure, data flow, and key functions. - Added `parking_pipeline.py` to orchestrate the loading of raw citations and rebuilding of the cleaned table. - Created `parking_clean.py` for managing the cleaned citations table. - Updated permissions for several scripts and configuration files. - Enhanced `parking_postgis.py` to support rebuilding the cleaned table after loading CSV data.
Pin the beta_pipeline requirements and add the violation-code analysis notebook used to derive the reference mappings. Ignore .venv/ and raw_data/ at the repo root: the pipeline expects the raw citations CSV to be staged locally and it is multi-GB, so it must never be committable. Co-authored-by: Cursor <cursoragent@cursor.com>
- Implemented `drop_incomplete` function in `parking_clean.py` to remove rows missing `issue_datetime` or `loc_lat`. - Updated `rebuild_clean` to call `drop_incomplete` after datetime creation. - Modified `parking_pipeline.py` to include dropping incomplete rows in the pipeline process. - Adjusted `parking_postgis.py` to ensure `drop_incomplete` is executed after rebuilding the cleaned table.
|
@gregpawin, this Pull Request is not linked to a valid issue. Please provide a valid linked issue in "Related Issues" above, using the format of "Resolves #" + issue number. |
glenflorendo
approved these changes
Sep 15, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Adds
data-science/beta_pipeline/, a Polars-based ingestion pipeline that reads the raw LA Parking Citations export,cleans it, and loads it into SQLite or PostGIS.
parking_clean.py— normalization and type coercion, including dropping records too incomplete to be analyticallyuseful rather than carrying nulls downstream.
parking_db.py/parking_postgis.py— schema creation and loading for the two targets.parking_pipeline.py— orchestration entry point tying the stages together.ARCHITECTURE.md— module-level reference;README.md— quick start.parking_db_explore.ipynbandviolation_analysis.ipynb— the exploratory work the violation-code and columnmappings were derived from.
Polars was chosen over pandas because the raw export is ~6 GB and does not fit comfortably in memory with pandas'
row-oriented overhead.
Also ignores
.venv/andraw_data/at the repo root, so the multi-GB raw CSV can be staged locally without becomingcommittable.
Related Issues
Refs #695
This is the ingestion/ETL half of #695. Measured against its acceptance criteria:
scheduling is not implemented.
I would suggest treating scheduling, incremental loading, and monitoring as follow-up issues rather than expanding this
PR, since each needs a decision this PR should not make unilaterally — in particular which orchestrator the project
wants to commit to.
Testing
Exercised manually against the full ~6 GB citations export and against truncated samples, loading into both SQLite and
PostGIS, with row counts and null rates compared before and after cleaning.
This code has no automated test suite yet, which is the main thing I would flag for review. It predates the tested
postgis_dbwork and I did not want to bundle a test-writing effort into this PR. If you would prefer tests before thislands, say so and I will add them.
Checklist
changes or issues.
Unchecked, deliberately: there is no automated test coverage here (see Testing), Python is outside the repo's
Prettier/ESLint scope, and #695's criteria are only partially met (see Related Issues).