diff --git a/docs/01-overview.md b/docs/01-overview.md new file mode 100644 index 00000000..5e79d742 --- /dev/null +++ b/docs/01-overview.md @@ -0,0 +1,160 @@ +# Overview + +## Mission + +**Lucky Parking** is a [Hack for LA](https://www.hackforla.org/) project that helps city planners and community members explore Los Angeles parking citation data and make informed decisions about parking policy. + +The repository combines: + +- A **web map** for interactive exploration of citations by date and geography +- A **data pipeline** for ingesting the city's multi-gigabyte citation export and keeping it fresh via API sync +- A **legacy backend and data-science stack** from earlier project phases + +The project is **not complete**. The web app, API, and new data pipeline were built at different times and are not fully integrated. See [Roadmap & open questions](./08-roadmap-and-open-questions.md) for the current gap list. + +## Who is this for? + +| Audience | Primary entry point | +|----------|---------------------| +| Frontend / full-stack developers | [Web application](./03-web-application.md), [Monorepo structure](./02-monorepo-structure.md) | +| Data / analytics contributors | [Data pipeline](./05-data-pipeline.md), [Data sources & schemas](./07-data-sources-and-schemas.md) | +| API / infra contributors | [Backend API](./04-backend-api.md), [Monorepo structure](./02-monorepo-structure.md) | +| New contributors | This page → [Monorepo structure](./02-monorepo-structure.md) → area of interest | + +## System at a glance + +```mermaid +flowchart LR + subgraph user [User] + Browser[Browser] + end + + subgraph web [Web tier — active] + Next["Next.js app
/parking-insights"] + Config["/api/v1/public/config"] + end + + subgraph external [External services] + Socrata["LA Open Data
Socrata 4f5p-udkv"] + Mapbox["Mapbox
maps + geocoding"] + end + + subgraph api [API tier — legacy / unused by web] + Express["Express API"] + MongoDB[(MongoDB)] + end + + subgraph data [Data tier — beta pipeline] + Pipeline["beta_pipeline
Python + Polars"] + SQLite[(SQLite)] + PostGIS[(PostGIS)] + end + + Browser --> Next + Next --> Config + Next --> Socrata + Next --> Mapbox + + Express --> MongoDB + Socrata --> Pipeline + Pipeline --> SQLite + Pipeline --> PostGIS + + Next -.->|"planned integration"| Express +``` + +## Three parallel data paths + +The same underlying citation records can flow through three different paths today: + +### 1. Web app → Socrata (live, limited) + +The Next.js app queries the [Socrata Query API v3](https://dev.socrata.com/docs/queries/) directly. Users pick a date range and up to two geographic areas; the app fetches matching citations and renders them on a Mapbox map. + +**Limitation:** Results are capped at **50 rows** per query (pagination not implemented). See [Web application](./03-web-application.md). + +### 2. Beta pipeline → SQLite / PostGIS (local analytics) + +The `data-science/beta_pipeline/` Python tools stream the ~6 GB CSV into SQLite or PostGIS using Polars batches. SQLite supports **incremental API sync**; PostGIS supports **spatial queries** and a cleaned analytics table. + +**Limitation:** PostGIS does not yet have API sync. The web app does not read from these databases. See [Data pipeline](./05-data-pipeline.md). + +### 3. Legacy pipeline → normalized PostGIS (older dataset) + +The `data-science/src/data/` scripts and Makefile target an **older dataset ID** (`wjz9-h9np`) and a normalized relational schema (`citation`, `vehicle`, `make`, etc.). This path predates the beta pipeline. + +**Limitation:** Dataset ID mismatch with current web app and beta pipeline. See [Legacy data science](./06-legacy-data-science.md). + +## Technology summary + +| Layer | Stack | +|-------|-------| +| Monorepo | Turborepo, pnpm workspaces | +| Web | Next.js 16, React 19, TypeScript, Tailwind 4, Mapbox GL | +| Shared UI | `@lucky-parking/design` (Radix-based components) | +| API | Express 4, MongoDB driver, Zod validation | +| Beta pipeline | Python 3.12, Polars, psycopg, uv (recommended) | +| Legacy data science | pandas, geopandas, SQLAlchemy, Conda/Makefile | +| Spatial DB | PostGIS 16 (Docker), SQLite (file) | +| CI | GitHub Actions (lint, format, build, test) | + +## Primary dataset + +All current-facing work uses the LA City Open Data dataset **[Parking Citations (`4f5p-udkv`)](https://data.lacity.org/Transportation/Parking-Citations/4f5p-udkv)** — millions of rows, ~6 GB as CSV, updated on a rolling basis by the city. + +Details: [Data sources & schemas](./07-data-sources-and-schemas.md). + +## Getting started (short) + +**Web app only:** + +```bash +pnpm install +cp apps/web/.env.schema apps/web/.env # add Mapbox + Socrata tokens +cd apps/web && pnpm dev +``` + +Open `/parking-insights`. + +**Data pipeline:** + +```bash +cd data-science/beta_pipeline +uv venv --python 3.12 .venv +uv pip install -r requirements.txt jupyter ipykernel --python .venv/bin/python +``` + +See [Data pipeline](./05-data-pipeline.md) for full workflows. + +## Document map + +```mermaid +mindmap + root((Lucky Parking docs)) + Overview + Mission + Three data paths + Monorepo + apps/web + apps/api + packages + Web app + Map + filters + Socrata queries + Zustand state + API + MongoDB GeoJSON + OpenAPI spec + Beta pipeline + SQLite sync + PostGIS clean table + Legacy DS + Makefile ETL + Old dataset ID + Data dictionary + 23 columns + Clean schema + Roadmap + TODOs + Integration gaps +``` diff --git a/docs/02-monorepo-structure.md b/docs/02-monorepo-structure.md new file mode 100644 index 00000000..da324807 --- /dev/null +++ b/docs/02-monorepo-structure.md @@ -0,0 +1,186 @@ +# Monorepo structure + +Lucky Parking is a [Turborepo](https://turbo.build/) monorepo managed with **pnpm workspaces**. One install at the root pulls dependencies for all apps and packages. + +## Directory layout + +``` +lucky-parking/ +├── apps/ +│ ├── web/ # Next.js frontend (primary user-facing app) +│ └── api/ # Express REST API (MongoDB) +├── packages/ +│ ├── design/ # Shared React UI component library +│ └── configs/ # Shared ESLint, Prettier, TypeScript, Tailwind configs +├── data-science/ +│ ├── beta_pipeline/ # Modern Polars + SQLite/PostGIS pipeline +│ ├── src/data/ # Legacy ETL scripts +│ ├── notebooks/ # Jupyter notebooks (exploratory + archived) +│ ├── references/ # Lookup tables (makes, violation codes, regex) +│ ├── db/ # Legacy PostGIS schema dump +│ └── docker/ # Conda + Jupyter Lab container +├── docs/ # This documentation set +├── scripts/ # Shared tooling (e.g. typecheck.mjs) +├── .github/workflows/ # CI (integration, compliance) +├── package.json # Root scripts via Turbo +├── pnpm-workspace.yaml +└── turbo.json +``` + +## Workspace packages + +```mermaid +graph TB + subgraph apps [apps/] + WEB["@lucky-parking/web"] + API["@lucky-parking/api"] + end + + subgraph packages [packages/] + DESIGN["@lucky-parking/design"] + CONFIGS["@lucky-parking/configs"] + end + + WEB --> DESIGN + WEB --> CONFIGS + API --> CONFIGS + DESIGN --> CONFIGS +``` + +| Package | Path | Role | +|---------|------|------| +| `@lucky-parking/web` | `apps/web` | Next.js map application | +| `@lucky-parking/api` | `apps/api` | Express citations API | +| `@lucky-parking/design` | `packages/design` | Buttons, sidebar, calendar, dialog, etc. | +| `@lucky-parking/configs` | `packages/configs` | Shared lint/format/TS/tailwind presets | + +## Turbo task pipeline + +Root `package.json` delegates to Turbo: + +| Command | Effect | +|---------|--------| +| `pnpm install` | Install all workspace dependencies | +| `pnpm dev` | Start dev servers (persistent, uncached) | +| `pnpm build` | Build apps/packages; depends on upstream `^build` | +| `pnpm lint` | ESLint across workspaces | +| `pnpm format` | Prettier | +| `pnpm test` | Test task (limited coverage today) | +| `pnpm check-types` | TypeScript checking | + +`turbo.json` defines task dependencies — for example, `build` waits for dependencies' builds and outputs `.next/` or `dist/`. + +## Prerequisites + +| Tool | Version (as documented) | Used for | +|------|---------------------------|----------| +| Node.js | 22+ (root); API Docker uses 22 | Web, API, tooling | +| pnpm | 9 | Package management | +| Python | 3.10+ (3.12 recommended) | `beta_pipeline` | +| uv | latest | Recommended Python env manager | +| Docker | optional | PostGIS for beta pipeline; legacy Jupyter image | + +## Environment files + +Each app documents its own env vars via `.env.schema`: + +| App / area | Schema file | Key variables | +|------------|-------------|---------------| +| Web | `apps/web/.env.schema` | `MAPBOX_ACCESS_TOKEN`, `SOCRATA_APP_TOKEN` | +| API | `apps/api/.env.schema` | `DB_*`, `API_VERSION`, `COL_NAME_CITATIONS` | +| Beta pipeline | documented in README | `SOCRATA_APP_TOKEN`, `DATABASE_URL` | + +**Note:** `.env` files are gitignored. Copy from `.env.schema` and fill in tokens. + +**Known issue:** The API code reads `COL_CITATIONS` but the schema documents `COL_NAME_CITATIONS`. See [Backend API](./04-backend-api.md). + +## Git hooks and CI + +```mermaid +flowchart LR + Commit[git commit] --> Husky[Husky pre-commit] + Husky --> LS[lint-staged] + LS --> Prettier + LS --> ESLint + LS --> Typecheck + + Push[git push / PR] --> GHA[GitHub Actions] + GHA --> Format + GHA --> Lint + GHA --> Build + GHA --> Test +``` + +- **Pre-commit:** Husky runs lint-staged (Prettier, ESLint, typecheck on staged files) +- **CI:** `.github/workflows/integration.yaml` — format, lint, build, test on PRs +- **Compliance:** `.github/workflows/compliance.yaml` — PR/issue linking rules + +## Running locally + +### Web (most common) + +```bash +pnpm install +cp apps/web/.env.schema apps/web/.env +# Edit .env with Mapbox and Socrata tokens +cd apps/web && pnpm dev +``` + +The root route redirects to `/parking-insights` (`apps/web/next.config.ts`). + +### API (standalone) + +```bash +cd apps/api +# Configure MongoDB connection via .env +pnpm dev # or see apps/api package.json +``` + +The web app does **not** call this API today. + +### Data pipeline + +See [Data pipeline (beta)](./05-data-pipeline.md). Lives outside the Node workspace but shares the repo. + +## What is not in the monorepo + +| Item | Location / note | +|------|-----------------| +| Citation CSV files | Downloaded locally; gitignored (`raw_data/`, etc.) | +| SQLite DB files | Generated by pipeline; not committed | +| Python `.venv` | Created per contributor; gitignored | +| Production deploy config | Referenced in OpenAPI (`luckyparking.org`) but not fully documented in repo | + +## Package boundaries (design intent) + +```mermaid +flowchart TB + subgraph presentation [Presentation] + WEB[apps/web] + end + + subgraph shared [Shared libraries] + DESIGN[packages/design] + end + + subgraph services [Services — partially used] + API[apps/api] + end + + subgraph analytics [Analytics — offline] + BETA[beta_pipeline] + LEG[legacy data-science] + end + + WEB --> DESIGN + WEB --> Socrata[Socrata API] + API --> Mongo[(MongoDB)] + BETA --> SQLite[(SQLite)] + BETA --> PostGIS[(PostGIS)] + LEG --> PostGIS + + WEB -.-> API + WEB -.-> BETA +``` + +The monorepo currently optimizes for **shared frontend tooling** and **co-located data work**. Full-stack integration (web ↔ API ↔ local DB) remains future work. diff --git a/docs/03-web-application.md b/docs/03-web-application.md new file mode 100644 index 00000000..e12f44ba --- /dev/null +++ b/docs/03-web-application.md @@ -0,0 +1,218 @@ +# Web application + +The primary user-facing application lives in `apps/web` — a **Next.js 16** app that visualizes LA parking citations on an interactive Mapbox map. + +**Route:** `/parking-insights` (root `/` redirects here) + +## User experience + +Users can: + +1. Select a **date range** (default: last 7 days) +2. Search and add up to **two places** (neighborhood councils, addresses via Mapbox geocoding) +3. View citations as **circle markers** and an optional **heatmap** on the map +4. See **aggregate stats** (total citations, total/average fines) in a sidebar panel +5. Click map features for citation detail popups + +```mermaid +flowchart TB + subgraph sidebar [Sidebar — DataPanel] + Places[Place search + list] + Dates[Date range picker] + Stats[DataVisuals — totals] + Legend[DataLegend] + About[About section] + end + + subgraph map [Map — ParkingCitationsMap] + Source[MapSourceCitations — GeoJSON] + Circles[MapLayerCircles] + Heat[MapLayerHeatmap] + Popup[Click popups] + end + + subgraph state [Client state] + Store[Zustand store] + RQ[React Query — useCitations] + end + + Places --> Store + Dates --> Store + Store --> RQ + RQ --> Source + Source --> Circles + Source --> Heat + Circles --> Popup +``` + +## Key files + +| Path | Purpose | +|------|---------| +| `src/app/parking-insights/page.tsx` | Main page layout (sidebar + map) | +| `src/app/layout.tsx` | Root layout, fonts, global styles | +| `src/app/api/v1/public/config/route.ts` | Serves public tokens to the client | +| `src/store.ts` | Zustand store (query, range, places) | +| `src/hooks/use-citations.tsx` | React Query hook wrapping Socrata fetch | +| `src/hooks/use-geocoder.tsx` | Mapbox + neighborhood council search | +| `src/hooks/use-public-config.ts` | Fetches Mapbox/Socrata tokens | +| `src/lib/socrata/parking-citations.ts` | Builds SoQL query, POSTs to Socrata | +| `src/components/map.tsx` | Map container | +| `src/components/map-source-citations.tsx` | GeoJSON source layer | +| `src/components/data-visuals.tsx` | Citation/fine statistics | +| `src/data/los-angeles*.json` | City/county boundary GeoJSON | + +## State management + +The app uses **Zustand** with Immer, persist, and devtools middleware (`src/store.ts`). + +| State | Type | Default | Persisted? | +|-------|------|---------|------------| +| `query` | string | `""` | No | +| `range` | `{ from, to }` | Last 7 days | Yes (localStorage) | +| `places` | `Map` | empty | Yes | +| `isHydrated` | boolean | false | No | + +**Constraints:** + +- Maximum **2 places** (`MAX_PLACES = 2`) +- Date range normalized to start/end of day via `date-fns` + +Persistence key: `luckyparking` in `localStorage`. + +## Data fetching flow + +```mermaid +sequenceDiagram + participant User + participant Store as Zustand store + participant Hook as useCitations + participant Config as /api/v1/public/config + participant Socrata as Socrata Query API v3 + + User->>Store: Set date range / places + Store->>Hook: places + range change + Hook->>Config: GET tokens (on load) + Config-->>Hook: mapboxAccessToken, socrataAppToken + Hook->>Socrata: POST query (GeoJSON) + Note over Socrata: pageSize: 50 (hard limit) + Socrata-->>Hook: FeatureCollection + Hook-->>User: Map layers update +``` + +### Socrata integration + +**Endpoint:** `https://data.lacity.org/api/v3/views/4f5p-udkv/query` + +**Query construction** (`parking-citations.ts`): + +- Date filter: `issue_date BETWEEN 'YYYY-MM-DD' AND 'YYYY-MM-DD'` +- Optional geo filter: `within_polygon(geocodelocation, 'WKT')` when places are selected +- Combines multiple place geometries into a multipolygon via Turf.js + +**Headers:** `Accept: application/vnd.geo+json`, `X-App-Token` + +**Validation:** Response parsed with Zod (`ParkingCitationFeatureCollectionSchema`) + +### Public config API + +`GET /api/v1/public/config` exposes server-side env vars to the browser: + +- `MAPBOX_ACCESS_TOKEN` +- `SOCRATA_APP_TOKEN` + +The client will not fetch citations unless both a Socrata token and at least one place are present. + +## Map layers + +| Component | Layer type | Behavior | +|-----------|------------|----------| +| `MapSourceCitations` | GeoJSON source | Driven by `useCitations` data | +| `MapLayerCircles` | Circle layer | Individual citation points | +| `MapLayerHeatmap` | Heatmap layer | Density visualization | +| Base map | Mapbox GL | LA county bounds enforced | + +Static boundary data under `src/data/` supports geocoding context and map constraints. + +## Component hierarchy + +```mermaid +graph TD + Page[parking-insights/page.tsx] + Page --> Sidebar[AppSidebar] + Page --> Header[AppHeader] + Page --> Map[ParkingCitationsMap] + + Sidebar --> DataPanel[DataPanel] + DataPanel --> SearchGeocoder + DataPanel --> SearchDateRange + DataPanel --> DataPlaceList + DataPanel --> DataVisuals + DataPanel --> DataLegend + + Map --> MapSourceCitations + Map --> MapLayerCircles + Map --> MapLayerHeatmap +``` + +UI primitives come from `@lucky-parking/design` (sidebar, accordion, calendar, button, etc.). + +## Environment setup + +```bash +cp apps/web/.env.schema apps/web/.env +``` + +| Variable | Required | Purpose | +|----------|----------|---------| +| `MAPBOX_ACCESS_TOKEN` | Yes | Map tiles and geocoding | +| `SOCRATA_APP_TOKEN` | Yes (for data) | Raises rate limits; required for queries in current logic | + +Free tokens: [Mapbox](https://account.mapbox.com/), [LA City Data / Socrata](https://data.lacity.org/login). + +## Known limitations (unfinished) + +| Issue | Location | Impact | +|-------|----------|--------| +| **50-row pagination cap** | `parking-citations.ts` FIXME | Map shows at most 50 citations per query | +| **No backend API usage** | architecture | Cannot leverage MongoDB-preloaded data | +| **Legend placeholder** | `data-legend.tsx` TODO | Legend may not reflect real violation categories | +| **Geocoder edge cases** | `use-geocoder.tsx` TODO | Address/postcode result types not fully handled | +| **No offline/local DB mode** | — | Pipeline databases not wired to web | + +See [Roadmap](./08-roadmap-and-open-questions.md) for planned fixes and integration options. + +## Possible future directions + +```mermaid +flowchart LR + Today[Today: Socrata direct
50 rows max] + + OptA[Option A: Pagination
fetch all pages client-side] + OptB[Option B: API proxy
apps/api + MongoDB/PostGIS] + OptC[Option C: Vector tiles
pre-aggregated map layers] + OptD[Option D: Hybrid
Socrata for recent, DB for bulk] + + Today --> OptA + Today --> OptB + Today --> OptC + Today --> OptD +``` + +| Direction | Pros | Cons | +|-----------|------|------| +| Client pagination | Smallest change | Slow for large date ranges; rate limits | +| Express API + DB | Full control, fast queries | Requires ingestion + deployment | +| Vector tiles / MVT | Best map performance at scale | New infra pipeline | +| Hybrid recent + historical | Good UX split | Two code paths to maintain | + +## Running and building + +```bash +cd apps/web +pnpm dev # development server +pnpm build # production build +pnpm lint # ESLint +``` + +From repo root: `pnpm dev` starts all configured dev tasks via Turbo. diff --git a/docs/04-backend-api.md b/docs/04-backend-api.md new file mode 100644 index 00000000..b9336d07 --- /dev/null +++ b/docs/04-backend-api.md @@ -0,0 +1,181 @@ +# Backend API + +The `apps/api` package is an **Express 4** REST service that queries parking citations from **MongoDB** using geospatial and date filters. + +**Important:** The current Next.js web app does **not** use this API. It queries Socrata directly. Treat this service as **legacy / alternate backend** until integration work is done. + +## Purpose (intended) + +Provide a stable API for filtering citations by: + +- **Date range** — `issue_date` between two ISO datetimes +- **Geography** — GeoJSON polygon (`$geoWithin`) + +This would allow pre-indexed queries over a full local copy of citations instead of live Socrata calls. + +## Architecture + +```mermaid +flowchart LR + Client[HTTP client] + Express[Express app] + Validator[Zod middleware] + Controller[CitationController] + Service[CitationService] + Mongo[(MongoDB collection)] + + Client -->|POST /v1/citations| Express + Express --> Validator + Validator --> Controller + Controller --> Service + Service --> Mongo +``` + +## Routes + +| Method | Path | Handler | Description | +|--------|------|---------|-------------| +| GET | `/{API_VERSION}/` | inline | Health/hello | +| POST | `/{API_VERSION}/citations` | `CitationController.listCitations` | Filter citations | + +`API_VERSION` defaults from env (see `.env.schema`). + +## Request contract + +**Body** (validated by `CitationFiltersSchema`): + +```json +{ + "dates": ["2025-01-01T00:00:00.000Z", "2025-01-31T23:59:59.999Z"], + "geometry": { + "type": "Polygon", + "coordinates": [[[...]]] + } +} +``` + +Both fields are optional. Empty filters return broader result sets (subject to MongoDB query logic). + +## MongoDB query logic + +`CitationService.findCitations()` (`src/services/citations.ts`): + +```javascript +{ + $and: [ + { issue_date: { $gte: dates[0], $lte: dates[1] } }, // if dates provided + { geometry: { $geoWithin: { $geometry: geometry } } }, // if geometry provided + { "geometry.coordinates": { $nin: [null] } } + ] +} +``` + +Documents are expected to store citation geometry in a GeoJSON-compatible shape. + +## OpenAPI specification + +Full contract: [`apps/api/src/docs/specs-v1.yaml`](../apps/api/src/docs/specs-v1.yaml) + +- Documented server: `https://luckyparking.org/api/v1` +- Response shape: `{ data: CitationFeatureCollection }` (GeoJSON FeatureCollection) +- Includes RFC 7946 GeoJSON schema components + +**Stale reference:** External docs link may still point to older dataset `wjz9-h9np` while the web app and beta pipeline use `4f5p-udkv`. + +## Environment variables + +From `apps/api/.env.schema`: + +| Variable | Purpose | +|----------|---------| +| `API_VERSION` | URL prefix (e.g. `v1`) | +| `DB_USERNAME` | MongoDB user | +| `DB_PASSWORD` | MongoDB password | +| `DB_HOST` | MongoDB host | +| `DB_NAME` | Database name | +| `COL_NAME_CITATIONS` | Collection name (documented) | + +**Bug:** Runtime code reads `COL_CITATIONS`, not `COL_NAME_CITATIONS`: + +```typescript +const { COL_CITATIONS } = process.env; +// ... +db.collection(COL_CITATIONS as string) +``` + +Contributors must set the variable name the code actually uses, or fix the mismatch. + +## Key source files + +| File | Role | +|------|------| +| `src/index.ts` | Server entry, MongoDB connect, listen | +| `src/app.ts` | Express app setup, routes, middleware | +| `src/controllers/citations.ts` | HTTP handler | +| `src/services/citations.ts` | MongoDB query | +| `src/database/client.ts` | Mongo client | +| `src/middleware/validator.ts` | Zod request validation | +| `src/utilities/schemas.ts` | Zod schemas (GeoPolygon is `z.any()` — TODO) | + +## Relationship to other systems + +```mermaid +flowchart TB + subgraph current [Current production path] + WEB[apps/web] + SOC[Socrata 4f5p-udkv] + WEB --> SOC + end + + subgraph api_path [API path — not connected] + API[apps/api] + MONGO[(MongoDB)] + API --> MONGO + end + + subgraph pipeline [Ingestion options] + BETA[beta_pipeline] + LEG[legacy upload scripts] + BETA --> SQLITE[(SQLite)] + BETA --> PG[(PostGIS)] + LEG --> PG + end + + WEB -.->|"future"| API + BETA -.->|"export / ETL"| MONGO + LEG -.-> MONGO +``` + +No automated job today loads `beta_pipeline` output into MongoDB for the API. + +## Unfinished / open items + +| Item | Status | +|------|--------| +| Web app integration | Not started | +| Env var naming (`COL_CITATIONS`) | Bug / inconsistency | +| Full GeoPolygon Zod schema | TODO in `schemas.ts` | +| Dataset alignment (`4f5p-udkv` vs `wjz9-h9np`) | Needs migration plan | +| Ingestion pipeline → MongoDB | Not implemented in beta_pipeline | +| Automated tests | Minimal / absent at app level | + +## Possible directions + +1. **Retire the API** — If Socrata + local PostGIS meet all needs, deprecate MongoDB path and document removal. +2. **Revive as BFF** — Express proxies to PostGIS (or SQLite) with the same filter contract as OpenAPI; web app switches from Socrata. +3. **MongoDB as cache** — Scheduled job syncs Socrata or CSV into Mongo for GeoJSON-native `$geoWithin` queries. +4. **Unify on PostGIS** — Single spatial database serves both analytics notebooks and a thin API layer (PostgREST, custom Express, or Next.js route handlers). + +Each option trades operational complexity against query performance and data freshness. See [Roadmap](./08-roadmap-and-open-questions.md). + +## Running locally + +Configure MongoDB connection and collection name, then: + +```bash +cd apps/api +pnpm install # from root is preferred +pnpm dev # see package.json for exact script +``` + +Test with `POST /v1/citations` and a JSON body matching the OpenAPI spec. diff --git a/docs/05-data-pipeline.md b/docs/05-data-pipeline.md new file mode 100644 index 00000000..09e059e0 --- /dev/null +++ b/docs/05-data-pipeline.md @@ -0,0 +1,274 @@ +# Data pipeline (beta) + +The **`data-science/beta_pipeline/`** directory contains the project's **modern data ingestion stack** — Python tools built on **Polars** for streaming multi-gigabyte CSV files into **SQLite** or **PostGIS**, with optional API sync (SQLite only). + +For module-level function reference, also see [`data-science/beta_pipeline/ARCHITECTURE.md`](../data-science/beta_pipeline/ARCHITECTURE.md). + +## Design goals + +| Goal | How it's achieved | +|------|-------------------| +| Handle ~6 GB CSV without OOM | Polars batched reads (100k rows default) | +| Safe re-runs | `INSERT OR IGNORE` / `ON CONFLICT DO NOTHING` | +| Keep data fresh | Socrata incremental sync (SQLite) | +| Spatial analytics | PostGIS geometry + GiST indexes + clean table | +| Simple local dev | Docker Compose for PostGIS; SQLite is a single file | + +## High-level architecture + +```mermaid +flowchart TB + subgraph sources [Sources] + CSV["Parking_Citations_*.csv"] + API["Socrata Resource API
4f5p-udkv.json"] + end + + subgraph shared [Shared parsing — parking_db.py] + BATCH["_csv_batches()"] + NORM["_normalize_csv_batch()"] + COLS["COLUMNS — 23 fields"] + end + + subgraph sqlite [SQLite path] + PDB[parking_db.py] + DB[(parking_citations.db)] + SYNC[update_from_api] + end + + subgraph postgis [PostGIS path] + PPG[parking_postgis.py] + RAW[(citations)] + PCL[parking_clean.py] + CLEAN[(citations_clean)] + PPL[parking_pipeline.py] + end + + CSV --> BATCH + BATCH --> NORM + NORM --> PDB + NORM --> PPG + API --> SYNC + PDB --> DB + SYNC --> DB + PPG --> RAW + RAW --> PCL + PCL --> CLEAN + PPL --> PPG + PPL --> PCL +``` + +## Modules + +### `parking_db.py` — SQLite + API sync + +Standalone module: CSV load, SQLite schema, Socrata incremental sync. + +| CLI command | Action | +|-------------|--------| +| `init` | Create tables and indexes | +| `load-csv FILE` | Stream CSV → SQLite | +| `sync` | Incremental API sync (newest first) | +| `stats` | Row counts, max date, last sync | + +**Sync algorithm:** + +```mermaid +flowchart TD + Start[offset = 0] --> Fetch[Fetch page ORDER BY :updated_at DESC] + Fetch --> Empty{Page empty?} + Empty -->|yes| Done[Caught up] + Empty -->|no| Check[Find existing ticket_numbers] + Check --> Insert[INSERT OR IGNORE new rows] + Insert --> Match{Any ticket on page
already in DB?} + Match -->|yes| Done + Match -->|no| Next[offset += page_size] + Next --> Fetch +``` + +### `parking_postgis.py` — Raw PostGIS store + +Reuses CSV parsing from `parking_db.py`. Adds `geom geometry(Point, 4326)` built at insert time: + +1. Prefer WKT from `geocodelocation` +2. Fallback to `ST_MakePoint(loc_long, loc_lat)` + +| CLI command | Action | +|-------------|--------| +| `init` | Enable PostGIS extension, create tables | +| `load-csv FILE [--clean]` | Bulk load; optional clean rebuild | +| `stats` | Counts, geometry coverage | + +### `parking_clean.py` — Analytics table + +Builds slim `citations_clean` from raw `citations`: + +- Combines `issue_date` + `issue_time` (HHMM) → `issue_datetime` +- Trims violation text fields +- Keeps ticket, violation, fine, geometry + +| CLI command | Action | +|-------------|--------| +| `init` | Create clean table | +| `rebuild` | `TRUNCATE` + full `INSERT` from raw | +| `stats` | Row count, datetime range, geom coverage | + +### `parking_pipeline.py` — Orchestration + +| CLI command | Action | +|-------------|--------| +| `run FILE` | Load CSV → rebuild clean → print stats | +| `clean` | Rebuild clean only (no CSV load) | + +## PostGIS full pipeline + +```mermaid +sequenceDiagram + participant User + participant Docker as docker compose + participant PPL as parking_pipeline.py + participant PPG as parking_postgis.py + participant PCL as parking_clean.py + participant PG as PostGIS + + User->>Docker: docker compose up -d + User->>PPL: run Parking_Citations.csv + PPL->>PPG: bulk_load_csv() + loop Each 100k batch + PPG->>PG: INSERT citations + geom + end + PPL->>PCL: rebuild_clean() + PCL->>PG: TRUNCATE citations_clean + PCL->>PG: INSERT cleaned rows + PPL-->>User: JSON stats (raw + clean) +``` + +## Docker setup + +`docker-compose.yml` runs **PostGIS 16**: + +| Setting | Value | +|---------|-------| +| Image | `postgis/postgis:16-3.4` | +| Port | `5432:5432` | +| Credentials | `parking` / `parking` / `parking` | +| Volume | `postgis_data` (persistent) | +| Healthcheck | `pg_isready` every 5s | + +Default DSN: `postgresql://parking:parking@localhost:5432/parking` + +## Environment + +| Variable | Used by | Default | +|----------|---------|---------| +| `SOCRATA_APP_TOKEN` | `parking_db.py sync` | none (anonymous API) | +| `DATABASE_URL` | PostGIS modules | local Docker DSN above | + +**Gap:** README references `.env.example` but the file is not in the repo yet. + +## Python environment + +Recommended setup with **uv**: + +```bash +cd data-science/beta_pipeline +uv venv --python 3.12 .venv +uv pip install -r requirements.txt jupyter ipykernel --python .venv/bin/python +``` + +**Dependencies** (`requirements.txt`): + +| Package | Role | +|---------|------| +| `polars` | CSV streaming, analytics | +| `python-dotenv` | Load `.env` | +| `psycopg[binary]` | PostGIS driver | +| `pandas`, `numpy`, `matplotlib`, `seaborn`, `geopandas` | Notebooks | + +`.venv/` is gitignored. + +## Notebooks + +| Notebook | Status | +|----------|--------| +| `parking_db_explore.ipynb` | **Active** — Polars exploration of CSV; documents SQLite workflow | +| `violation_analysis.ipynb` | **Stub** — imports only; analysis not written yet | + +## Module dependency graph + +``` +parking_db.py (standalone) + ↑ +parking_postgis.py (imports CSV helpers) + ↑ +parking_clean.py (imports connect, init_db) + ↑ +parking_pipeline.py (orchestrates load + clean) +``` + +## Design tradeoffs + +| Decision | Benefit | Cost | +|----------|---------|------| +| `INSERT OR IGNORE` | Idempotent bulk loads | Ticket corrections never update existing rows | +| Separate `citations_clean` | Simple analytics schema | Full rebuild on each clean (no incremental upsert) | +| SQLite fast pragmas | Faster bulk load | Less crash safety during load | +| PostGIS `synchronous_commit=OFF` | Faster ingest | Recent commits may be lost on crash | + +## Unfinished work + +| Feature | Status | Notes | +|---------|--------|-------| +| PostGIS API sync | **Not implemented** | Copy/adapt `update_from_api` from SQLite | +| Incremental clean rebuild | **Not implemented** | Today: full `TRUNCATE` + insert | +| Scheduled automation | **Not implemented** | cron/launchd after CSV download | +| Web app integration | **Not implemented** | No API reads from SQLite/PostGIS | +| `.env.example` | **Missing file** | Documented in README only | +| `violation_analysis.ipynb` | **Empty** | Planned analysis TBD | + +## Possible directions + +```mermaid +mindmap + root((beta_pipeline future)) + Sync + PostGIS API sync + Unified sync module + Clean layer + Incremental upsert + Materialized views + Integration + Next.js API routes + GraphQL over PostGIS + Export to Mongo for legacy API + Ops + GitHub Action nightly sync + Cloud RDS / Supabase + dbt models on citations_clean + Analysis + violation_analysis notebook + dbt metrics + ML feature store +``` + +## Quick command reference + +**SQLite (no Docker):** + +```bash +.venv/bin/python parking_db.py init +.venv/bin/python parking_db.py load-csv Parking_Citations_20260426.csv +.venv/bin/python parking_db.py sync +.venv/bin/python parking_db.py stats +``` + +**PostGIS:** + +```bash +docker compose up -d +.venv/bin/python parking_pipeline.py run Parking_Citations_20250811.csv +.venv/bin/python parking_clean.py stats +``` + +Replace CSV filenames with your local download from [data.lacity.org](https://data.lacity.org/Transportation/Parking-Citations/4f5p-udkv). + +Schema details: [Data sources & schemas](./07-data-sources-and-schemas.md). diff --git a/docs/06-legacy-data-science.md b/docs/06-legacy-data-science.md new file mode 100644 index 00000000..f86d6e58 --- /dev/null +++ b/docs/06-legacy-data-science.md @@ -0,0 +1,203 @@ +# Legacy data science + +Before `beta_pipeline/`, the project used a **Cookiecutter-style data science layout** under `data-science/` with Makefile-driven ETL, normalized PostGIS schemas, and extensive Jupyter notebooks. + +This stack is **still in the repo** but is largely **superseded** by the beta pipeline for new work. It targets a **different Socrata dataset ID** and a different database schema. + +## Legacy vs beta pipeline + +```mermaid +flowchart LR + subgraph legacy [Legacy — data-science/] + OLD_DS["Dataset wjz9-h9np"] + MK[make_dataset.py] + UP[upload_serial.py] + NORM[(Normalized schema
citation, vehicle, make...)] + end + + subgraph beta [Beta — beta_pipeline/] + NEW_DS["Dataset 4f5p-udkv"] + PDB[parking_db.py / parking_postgis.py] + FLAT[(Flat 23-column schema
citations / citations_clean)] + end + + subgraph apps [Applications] + WEB[apps/web → 4f5p-udkv] + API[apps/api spec → wjz9-h9np reference] + end + + OLD_DS --> MK --> UP --> NORM + NEW_DS --> PDB --> FLAT + WEB --> NEW_DS + API -.-> OLD_DS +``` + +| Aspect | Legacy | Beta pipeline | +|--------|--------|---------------| +| Dataset ID | `wjz9-h9np` | `4f5p-udkv` | +| Primary tool | pandas / geopandas | Polars | +| DB schema | Normalized relational | Flat citation table + clean layer | +| API sync | Via older scripts / manual | SQLite `sync` command | +| Recommended for new work? | No | **Yes** | + +## Directory layout + +``` +data-science/ +├── src/data/ # ETL scripts (click/Makefile driven) +├── notebooks/ +│ ├── exploratory/ # Active exploration notebooks +│ └── archived_notebooks/ # Older ML, viz, upload experiments +├── references/ # Lookup tables and regex rules +├── db/ # db_dev.sql — legacy schema dump +├── docker/ # Conda + Jupyter Lab image +├── old_docker/ # Deprecated Dockerfiles +├── docs/ # Sphinx documentation (partially stale) +├── Makefile # Primary automation interface +├── requirements.txt # Legacy Python deps +└── beta_pipeline/ # Modern pipeline (see separate doc) +``` + +## Makefile workflow + +The Makefile (`data-science/Makefile`) is the main entry point for legacy tasks: + +| Target | Action | +|--------|--------| +| `make requirements` | Install Python dependencies | +| `make data` | Run `make_dataset.py` (raw → processed) | +| `make sample` | Create sample datasets | +| `make serial_data` | Build serial-friendly output | +| `make upload_serial` | Upload to PostGIS | +| `make upload_geojson` | Upload GeoJSON via `upload.py` | +| `make upload_zip` | Upload zipcode boundaries | +| `make upload_neighborhood` | Upload neighborhood councils | +| `make lint` | flake8 on `src/` | +| `make clean` / `make clean_data` | Remove caches / data files | + +Requires Conda or system Python (`PYTHON_INTERPRETER = python3`). + +## Key ETL scripts (`src/data/`) + +| Script | Purpose | +|--------|---------| +| `make_dataset.py` | Download/process raw citation CSV | +| `make_dataset_dask.py` | Dask-based variant for large files | +| `make_serial_data.py` | Prepare serial upload format | +| `upload_serial.py` | Load processed data into PostGIS | +| `upload.py` | Upload GeoJSON layers | +| `upload_neighborhood.py` | Neighborhood council boundaries | +| `get_zipcodes.py` | Zipcode boundary data | +| `sample.py` | Generate sample subsets | +| `date_threshold.py` | Filter by date threshold | + +These scripts expect `.env` with `DB_USER`, `DB_PASSWORD`, `DB_HOST`, `DB_PORT`, `DB_DATABASE`. + +## Legacy PostGIS schema + +`db/db_dev.sql` defines a **normalized** model: + +```mermaid +erDiagram + citation ||--o| vehicle : has + citation }o--|| codes : violation + vehicle }o--o| make : references + citation }o--o| zipcodes : located_in + neighborhood_councils ||--o{ citation : contains + + citation { + text ticket_number PK + timestamp issue_date + geometry geom + } + vehicle { + text vin + text make + text color + } + codes { + text violation_code + text description + } + make { + text make_name + } +``` + +This differs from beta_pipeline's flat `citations` table with 23 city-export columns. + +## Reference data (`references/`) + +Lookup and normalization files used by legacy cleaning: + +| File | Purpose | +|------|---------| +| `make.csv`, `makes.json`, `top_makes.txt` | Vehicle make normalization | +| `violation_codes.json`, `violation_descriptions.json` | Violation lookup | +| `vio_regex.csv` | Regex rules for violation descriptions | +| `column_names.json` | Output column mapping | +| `top_violation_codes.txt` | Common codes list | + +These may be useful for future cleaning logic in `beta_pipeline` or `violation_analysis.ipynb` but are not wired into the beta pipeline today. + +## Notebooks + +### Exploratory (`notebooks/exploratory/`) + +Active-ish exploration notebooks (e.g. `1-gp-explore_raw.ipynb`, `1-fl-analysis.ipynb`). + +### Archived (`notebooks/archived_notebooks/`) + +Historical work including: + +- Google Maps citation visualization +- Random Forest / zip code models +- Reddit data (PRAW) +- Server upload experiments +- Regulation sweeping exploration + +Treat archived notebooks as **historical context**, not current runbooks. + +## Docker (legacy Jupyter) + +`data-science/docker/` provides a Conda-based Jupyter Lab image (port **8888**). See `data-science/docker/README.md`. + +`old_docker/` contains deprecated Dockerfiles — do not use for new work. + +## Sphinx docs + +`data-science/docs/` contains Sphinx scaffolding. Some documented commands (e.g. S3 sync) are **not present in the Makefile** — docs may be stale. + +## When to use legacy vs beta + +| Use case | Recommendation | +|----------|----------------| +| New CSV ingestion for LA citations | **beta_pipeline** | +| Spatial analytics on flat schema | **beta_pipeline** + PostGIS | +| Incremental Socrata sync | **beta_pipeline** SQLite | +| Understanding old ML/viz experiments | Legacy notebooks | +| Normalized vehicle/violation schema | Legacy (or redesign on top of beta) | + +## Migration considerations (unfinished) + +No automated migration exists between: + +- Legacy normalized PostGIS ↔ beta flat PostGIS +- Either database ↔ MongoDB (API) +- Either database ↔ web app + +A full migration plan would need to address: + +1. **Dataset ID** alignment (`wjz9-h9np` → `4f5p-udkv`) +2. **Schema mapping** (normalized tables vs 23-column flat + clean) +3. **Reference data** port (make/violation normalization into clean step) +4. **Retiring or repointing** legacy Makefile scripts + +See [Roadmap](./08-roadmap-and-open-questions.md). + +## Possible directions + +- **Deprecate legacy ETL** — Archive `src/data/` scripts, keep notebooks for reference only +- **Port normalization into beta clean step** — Use `references/` in `parking_clean.py` +- **Unified Makefile or uv project** — Single Python entry point wrapping beta_pipeline CLI +- **Revive ML notebooks** — Re-run archived models against `citations_clean` in PostGIS diff --git a/docs/07-data-sources-and-schemas.md b/docs/07-data-sources-and-schemas.md new file mode 100644 index 00000000..24a62292 --- /dev/null +++ b/docs/07-data-sources-and-schemas.md @@ -0,0 +1,229 @@ +# Data sources and schemas + +This document describes where citation data comes from, how it is shaped in each system, and known data quality issues. + +## Primary dataset (current) + +**[Parking Citations — `4f5p-udkv`](https://data.lacity.org/Transportation/Parking-Citations/4f5p-udkv)** + +Published by the City of Los Angeles on the Socrata open data portal. + +| Property | Value | +|----------|-------| +| Approximate size | ~6 GB CSV | +| Row count | Millions of citations | +| Update frequency | Rolling updates by city | +| Used by | Web app, beta_pipeline | + +### Access methods + +```mermaid +flowchart TB + DS["Dataset 4f5p-udkv"] + + DS --> CSV["CSV export
(Transportation → Parking Citations)"] + DS --> RES["Resource API
/resource/4f5p-udkv.json"] + DS --> QRY["Query API v3
/api/v3/views/4f5p-udkv/query"] + + CSV --> BETA_LOAD["beta_pipeline load-csv"] + RES --> BETA_SYNC["beta_pipeline sync"] + QRY --> WEB["apps/web fetchParkingCitations"] +``` + +| Method | URL pattern | Consumer | +|--------|-------------|----------| +| CSV download | data.lacity.org export | `parking_db.py`, `parking_postgis.py` | +| Resource API | `https://data.lacity.org/resource/4f5p-udkv.json` | `parking_db.py sync` | +| Query API v3 | `https://data.lacity.org/api/v3/views/4f5p-udkv/query` | Next.js web app | + +Optional **`X-App-Token`** header improves rate limits (free registration at data.lacity.org). + +## Legacy dataset + +**[`wjz9-h9np`](https://data.lacity.org/)** — older parking citations dataset referenced by: + +- Legacy `make_dataset.py` and related ETL +- OpenAPI external docs in `apps/api` (may be stale) + +New work should use **`4f5p-udkv`** unless explicitly migrating historical comparisons. + +## Canonical column set (23 fields) + +Both SQLite and PostGIS raw tables in `beta_pipeline` store these columns, defined as `COLUMNS` in `parking_db.py`: + +| Column | Typical type | Description | +|--------|--------------|-------------| +| `ticket_number` | string | Primary key — unique citation ID | +| `issue_date` | datetime/text | Citation date (CSV: `"2025 Apr 26 12:00:00 AM"`) | +| `issue_time` | string | Time as HHMM without leading zeros (e.g. `"904"`, `"1430"`) | +| `meter_id` | string | Parking meter identifier | +| `marked_time` | string | Marked time field from source | +| `rp_state_plate` | string | Registered plate state | +| `plate_expiry_date` | string | Plate expiration | +| `vin` | string | Vehicle VIN | +| `make` | string | Vehicle make | +| `body_style` | string | Body style code | +| `color` | string | Vehicle color code | +| `location` | string | Street location description | +| `route` | string | Route identifier | +| `agency` | integer | Issuing agency code | +| `violation_code` | string | Violation code | +| `violation_description` | string | Human-readable violation | +| `fine_amount` | float | Fine in dollars | +| `agency_desc` | string | Agency description | +| `color_desc` | string | Color description | +| `body_style_desc` | string | Body style description | +| `loc_lat` | float | Latitude | +| `loc_long` | float | Longitude | +| `geocodelocation` | string | WKT POINT geometry from CSV | + +### CSV parsing notes + +- Polars schema assigns explicit types per column +- `null_values`: `""`, `"NA"`, `"N/A"` +- `ignore_errors=True` — malformed rows skipped, not fatal +- Batch size default: **100,000 rows** + +## PostGIS raw table (`citations`) + +All 23 columns plus: + +```sql +geom geometry(Point, 4326) +``` + +**Geometry construction priority:** + +1. `ST_GeomFromText(geocodelocation)` when WKT present +2. Else `ST_MakePoint(loc_long, loc_lat)` +3. Else `NULL` + +**Indexes:** `issue_date`, `violation_code`, `make`, GiST on `geom` + +## PostGIS clean table (`citations_clean`) + +Slim analytics schema: + +| Column | Type | Notes | +|--------|------|-------| +| `ticket_number` | TEXT PK | | +| `issue_datetime` | TIMESTAMPTZ NOT NULL | Combined date + parsed HHMM time | +| `violation_code` | TEXT | Trimmed | +| `violation_description` | TEXT | Trimmed | +| `fine_amount` | DOUBLE PRECISION | | +| `geom` | geometry(Point, 4326) | Copied from raw | + +**Datetime parsing:** + +``` +issue_date: 2025-04-26T00:00:00.000 (midnight from CSV) +issue_time: "904" → pad → "0904" → 09:04 → 2025-04-26 09:04:00+TZ +``` + +Missing or non-numeric `issue_time` defaults to midnight on issue date. + +## SQLite schema + +### `citations` + +Same 23 columns, `ticket_number TEXT PRIMARY KEY`, `WITHOUT ROWID`. + +### `sync_log` + +| Column | Type | Purpose | +|--------|------|---------| +| `id` | INTEGER PK | Auto-increment | +| `started_at` | TEXT | ISO UTC | +| `finished_at` | TEXT | ISO UTC | +| `source` | TEXT | `'csv'` or `'api'` | +| `rows_inserted` | INTEGER | Rows processed | +| `notes` | TEXT | e.g. `'caught_up'` | + +## API record shape (Socrata sync) + +Socrata returns GeoJSON for location; SQLite sync converts to WKT for `geocodelocation` consistency with CSV rows. + +Coerced fields: `agency`, `fine_amount`, `loc_lat`, `loc_long`. + +## MongoDB documents (legacy API) + +Expected shape (GeoJSON-centric, per OpenAPI): + +- `issue_date` — filterable datetime +- `geometry` — GeoJSON Point or similar for `$geoWithin` + +Exact document schema is not fully documented in repo — inferred from `CitationService` query logic. + +## Auxiliary / boundary data + +| Data | Location | Used for | +|------|----------|----------| +| LA city boundary | `apps/web/src/data/los-angeles.json` | Map constraints | +| LA county boundary | `apps/web/src/data/los-angeles-county.json` | Map bounds | +| Neighborhood councils | Referenced in geocoder hook | Place search | +| Mock citations / geocoder | `mock-*.json` | Development / testing | +| Make / violation lookups | `data-science/references/` | Legacy normalization | +| Zipcodes | Legacy upload scripts | Spatial joins | + +## Data quality quirks + +```mermaid +flowchart TD + RAW[Raw citation row] + RAW --> Q1{Future issue_date?} + RAW --> Q2{Valid issue_time?} + RAW --> Q3{Has coordinates?} + RAW --> Q4{Duplicate ticket_number?} + + Q1 -->|some rows| FILTER1[Filter in queries] + Q2 -->|missing / odd| DEFAULT[Default to midnight in clean] + Q3 -->|often no| NULLGEOM[geom IS NULL] + Q4 -->|re-load / sync| SKIP[INSERT OR IGNORE — no update] +``` + +| Quirk | Detail | Mitigation | +|-------|--------|------------| +| Future dates | Some `issue_date` values years ahead | Filter in analysis queries | +| `issue_time` format | HHMM, no leading zeros; `"0"` → midnight | Handled in clean rebuild | +| Missing geometry | Not all rows geocoded | `WHERE geom IS NOT NULL` for maps | +| API vs CSV location | GeoJSON vs WKT | Normalized on API ingest | +| No ticket updates | `INSERT OR IGNORE` | Switch to upsert if corrections needed | +| Web app row cap | 50 rows per Socrata query | Pagination or DB backend | + +## Example queries + +**PostGIS — recent geocoded citations:** + +```sql +SELECT ticket_number, issue_datetime, violation_description, ST_AsText(geom) +FROM citations_clean +WHERE geom IS NOT NULL +ORDER BY issue_datetime DESC +LIMIT 10; +``` + +**SQLite + Polars:** + +```python +import sqlite3, polars as pl + +with sqlite3.connect("parking_citations.db") as conn: + df = pl.read_database(""" + SELECT violation_description, COUNT(*) AS n, + ROUND(AVG(fine_amount), 2) AS avg_fine + FROM citations + WHERE violation_description IS NOT NULL + GROUP BY violation_description + ORDER BY n DESC + LIMIT 15 + """, connection=conn) +``` + +## File artifacts (not in git) + +| Artifact | Typical path | Created by | +|----------|--------------|------------| +| Citation CSV | `Parking_Citations_*.csv` | Manual download | +| SQLite DB | `parking_citations.db` | `parking_db.py` | +| PostGIS volume | Docker `postgis_data` | `docker compose` | +| Raw data dir | `raw_data/` (gitignored) | Various | diff --git a/docs/08-roadmap-and-open-questions.md b/docs/08-roadmap-and-open-questions.md new file mode 100644 index 00000000..3e72a25a --- /dev/null +++ b/docs/08-roadmap-and-open-questions.md @@ -0,0 +1,205 @@ +# Roadmap and open questions + +Lucky Parking is an active Hack for LA project with **multiple partially overlapping systems**. This page catalogs what is unfinished, known bugs, and reasonable directions the project could take. + +**Status key:** ✅ Done · 🟡 Partial · 🔴 Not started / broken · 📋 Documented but missing + +## Integration map (today) + +```mermaid +flowchart TB + subgraph done [Working today] + WEB[Web app map + filters] + SOC[Socrata live queries] + BETA_CSV[beta_pipeline CSV load] + BETA_SQL[beta_pipeline SQLite sync] + BETA_PG[beta_pipeline PostGIS + clean] + end + + subgraph partial [Partial / limited] + WEB50[Web 50-row cap] + NB1[parking_db_explore notebook] + end + + subgraph missing [Not connected / missing] + WEB_API[Web → Express API] + API_DB[API → MongoDB populated] + PG_SYNC[PostGIS API sync] + WEB_DB[Web → local DB] + VIO_NB[violation_analysis notebook] + ENV_EX[.env.example in beta_pipeline] + end + + WEB --> SOC + WEB50 --> WEB + BETA_CSV --> BETA_PG + BETA_SQL --> BETA_CSV +``` + +## Unfinished by area + +### Web application + +| Item | Status | Location | Notes | +|------|--------|----------|-------| +| Socrata pagination | 🔴 | `apps/web/src/lib/socrata/parking-citations.ts:73` | Hard `pageSize: 50`; FIXME comment | +| Map legend accuracy | 🔴 | `apps/web/src/components/data-legend.tsx` | TODO: use real violation categories | +| Geocoder address/postcode handling | 🟡 | `apps/web/src/hooks/use-geocoder.tsx:65` | TODO for some result types | +| Connect to backend API | 🔴 | architecture | Web bypasses `apps/api` entirely | +| Connect to local PostGIS/SQLite | 🔴 | architecture | Pipeline output unused by web | +| Automated tests | 🔴 | `apps/web` | No substantive test suite | + +### Backend API + +| Item | Status | Location | Notes | +|------|--------|----------|-------| +| Used by frontend | 🔴 | — | Express API is orphaned | +| Env var naming bug | 🔴 | `citations.ts` vs `.env.schema` | Code uses `COL_CITATIONS`, schema says `COL_NAME_CITATIONS` | +| GeoPolygon validation | 🟡 | `apps/api/src/utilities/schemas.ts` | TODO: proper Zod schema (`z.any()` today) | +| Dataset ID alignment | 🔴 | OpenAPI external docs | May reference `wjz9-h9np` vs current `4f5p-udkv` | +| Ingestion into MongoDB | 🔴 | — | No pipeline loads current dataset into API DB | +| Automated tests | 🔴 | `apps/api` | Minimal coverage | + +### Beta data pipeline + +| Item | Status | Location | Notes | +|------|--------|----------|-------| +| SQLite CSV + API sync | ✅ | `parking_db.py` | Production-ready for local use | +| PostGIS CSV + clean | ✅ | `parking_postgis.py`, `parking_clean.py` | Full pipeline works | +| PostGIS API sync | 🔴 | ARCHITECTURE.md | Copy from SQLite, add geom handling | +| Incremental clean rebuild | 🔴 | `parking_clean.py` | Full TRUNCATE each time | +| Scheduled sync / load | 🔴 | — | No cron/CI automation | +| `.env.example` file | 📋 | README references it | File not in repo | +| `violation_analysis.ipynb` | 🔴 | stub imports only | Analysis not written | +| Web/API consumption | 🔴 | — | Databases are offline analytics only | + +### Legacy data science + +| Item | Status | Notes | +|------|--------|-------| +| Makefile ETL | 🟡 | Works for old dataset/schema; not maintained for `4f5p-udkv` | +| Sphinx docs | 🟡 | Some commands documented but not in Makefile | +| S3 sync | 📋 | Referenced in docs, not implemented in Makefile | +| Cookiecutter `src/models`, `src/features` | 🔴 | Never added; only `src/data/` exists | +| Archived ML notebooks | 🟡 | Historical; not validated against current data | + +### Repository / ops + +| Item | Status | Notes | +|------|--------|-------| +| Root `pnpm test` | 🟡 | Task exists; apps lack meaningful tests | +| Branch naming docs | 🟡 | CONTRIBUTING mentions `master`/`dev`; CI may use `main`/`stable` | +| Monorepo Python tooling | 🔴 | Python lives outside pnpm; no unified root Python project | + +## Code-tracked TODOs / FIXMEs + +| File | Marker | Summary | +|------|--------|---------| +| `parking-citations.ts` | FIXME | Remove 50-row pagination limit | +| `use-geocoder.tsx` | TODO | Handle address and postcode geocoder types | +| `data-legend.tsx` | TODO | Refactor legend to real data | +| `schemas.ts` (API) | TODO | Full GeoPolygon Zod schema | + +## Strategic forks (where the project could go) + +### Path A — "Live Socrata first" (minimal change) + +Improve the current web → Socrata path: + +- Implement client-side pagination in `fetchParkingCitations` +- Add loading states and rate-limit handling +- Fix legend and geocoder edge cases + +**Best if:** Team lacks infra for hosted DB; citations needed are small (recent + localized). + +**Risk:** Socrata rate limits and latency at scale; still no offline analysis parity with pipeline. + +### Path B — "PostGIS as source of truth" + +Make `beta_pipeline` + PostGIS the backend for everything: + +- Add PostGIS API sync +- Expose citations via Next.js route handlers or revived Express API querying PostGIS +- Point web app at internal API instead of Socrata +- Deploy PostGIS (RDS, Supabase, etc.) + +**Best if:** Full-city or long date-range queries matter; team can run a database. + +**Risk:** Ops cost; need ingestion monitoring and auth for public API. + +### Path C — "Revive MongoDB API" + +Populate MongoDB from pipeline exports; web switches to existing OpenAPI contract: + +- ETL job: PostGIS/SQLite → GeoJSON documents → MongoDB +- Fix env var bug and GeoPolygon schema +- Update dataset to `4f5p-udkv` + +**Best if:** Team prefers document store + existing API spec. + +**Risk:** Two storage systems (PostGIS for analytics, Mongo for API) unless Mongo becomes sole store. + +### Path D — "Analytics focus" + +Prioritize data science deliverables over web integration: + +- Complete `violation_analysis.ipynb` +- Build scheduled sync + clean jobs +- Publish insights/reports; web remains demo with 50-row cap + +**Best if:** Primary stakeholders are researchers/policy analysts, not public map users. + +### Path E — "Consolidate and deprecate" + +Remove or archive legacy paths to reduce contributor confusion: + +- Mark `data-science/src/data/` and Mongo API as deprecated +- Single Python package (`beta_pipeline`) + single dataset ID +- Document one golden path in README + +**Best if:** Maintainer bandwidth is limited. + +## Suggested near-term priorities + +A pragmatic sequence many teams would follow: + +```mermaid +gantt + title Possible near-term sequence + dateFormat YYYY-MM + section Quick wins + Add .env.example to beta_pipeline :a1, 2026-06, 1w + Fix API COL_CITATIONS env bug :a2, 2026-06, 1w + Document dataset ID in OpenAPI :a3, 2026-06, 1w + section Web UX + Socrata pagination OR raise limit :b1, 2026-07, 2w + Fix data legend :b2, 2026-07, 1w + section Data + violation_analysis notebook :c1, 2026-07, 3w + PostGIS API sync :c2, 2026-08, 3w + section Integration + Choose BFF strategy (B vs C) :d1, 2026-08, 2w + Wire web to internal API :d2, 2026-09, 4w +``` + +*Timeline is illustrative — not a committed project plan.* + +## Open questions for the team + +1. **Which dataset is canonical going forward?** Assume `4f5p-udkv` unless legacy comparisons require `wjz9-h9np`. +2. **Should the Express API survive?** Or replace with Next.js server routes / PostgREST? +3. **Is 50 rows acceptable temporarily?** Or is pagination/blocking for launch? +4. **Who hosts PostGIS in production?** Docker locally only vs cloud managed service. +5. **Are ticket corrections important?** If yes, move from `INSERT OR IGNORE` to upsert. +6. **Should violation normalization (`references/`) feed into `citations_clean`?** +7. **What is the public deployment target?** OpenAPI references `luckyparking.org` — document actual infra. + +## How to update this doc + +When closing a gap: + +1. Change status in the tables above +2. Remove or resolve corresponding TODO/FIXME in code +3. Link to PR or issue if tracked on GitHub + +When adding new scope, append to the strategic forks section with pros/cons so future contributors understand why a path was chosen or rejected. diff --git a/docs/README.md b/docs/README.md index 1738e47a..ba024a0c 100644 --- a/docs/README.md +++ b/docs/README.md @@ -103,6 +103,65 @@ Run these from the repository root. | `pnpm verify` | Run type checks, linting, formatting, tests, and builds | | `pnpm clean` | Remove generated workspace artifacts and dependencies | +## Detailed documentation + +| Document | What you'll learn | +|----------|-------------------| +| [Overview](./01-overview.md) | Project mission, high-level architecture, and how the pieces relate | +| [Monorepo structure](./02-monorepo-structure.md) | Turborepo layout, packages, tooling, and local dev workflow | +| [Web application](./03-web-application.md) | Next.js map app, state, Socrata integration, components | +| [Backend API](./04-backend-api.md) | Express/MongoDB API — current status and contract | +| [Data pipeline (beta)](./05-data-pipeline.md) | Polars + SQLite/PostGIS ingestion, CLI, schemas | +| [Legacy data science](./06-legacy-data-science.md) | Older ETL, notebooks, and normalized PostGIS schema | +| [Data sources & schemas](./07-data-sources-and-schemas.md) | Dataset IDs, column dictionary, reference files | +| [Roadmap & open questions](./08-roadmap-and-open-questions.md) | Unfinished work, known gaps, and possible directions | + +## Quick reference + +```mermaid +flowchart TB + subgraph sources [External data] + CSV["Parking Citations CSV (~6 GB)"] + SOC["Socrata API (4f5p-udkv)"] + MAP["Mapbox Geocoding"] + end + + subgraph monorepo [Lucky Parking monorepo] + WEB["apps/web — Next.js map"] + API["apps/api — Express + MongoDB"] + DS["data-science/beta_pipeline"] + LEG["data-science/ (legacy)"] + end + + subgraph storage [Local / cloud storage] + SQLITE[(SQLite)] + PG[(PostGIS)] + MONGO[(MongoDB)] + end + + CSV --> DS + SOC --> DS + SOC --> WEB + MAP --> WEB + + DS --> SQLITE + DS --> PG + API --> MONGO + + WEB -.->|"not connected today"| API + LEG -.->|"older dataset (wjz9-h9np)"| PG +``` + +## Related docs elsewhere in the repo + +- [`data-science/beta_pipeline/README.md`](../data-science/beta_pipeline/README.md) — quick start for the Python pipeline +- [`data-science/beta_pipeline/ARCHITECTURE.md`](../data-science/beta_pipeline/ARCHITECTURE.md) — module-level code reference (complements [05-data-pipeline.md](./05-data-pipeline.md)) +- [`apps/api/src/docs/specs-v1.yaml`](../apps/api/src/docs/specs-v1.yaml) — OpenAPI spec for the citations API + +## Documentation status + +This documentation was written to reflect the repository as of mid-2026. Where behavior is uncertain or in flux, see [Roadmap & open questions](./08-roadmap-and-open-questions.md). + ## Contributing Contributions are welcome. Start with Hack for LA's [onboarding guide](https://www.hackforla.org/getting-started), then