diff --git a/docs/01-overview.md b/docs/01-overview.md
new file mode 100644
index 00000000..5e79d742
--- /dev/null
+++ b/docs/01-overview.md
@@ -0,0 +1,160 @@
+# Overview
+
+## Mission
+
+**Lucky Parking** is a [Hack for LA](https://www.hackforla.org/) project that helps city planners and community members explore Los Angeles parking citation data and make informed decisions about parking policy.
+
+The repository combines:
+
+- A **web map** for interactive exploration of citations by date and geography
+- A **data pipeline** for ingesting the city's multi-gigabyte citation export and keeping it fresh via API sync
+- A **legacy backend and data-science stack** from earlier project phases
+
+The project is **not complete**. The web app, API, and new data pipeline were built at different times and are not fully integrated. See [Roadmap & open questions](./08-roadmap-and-open-questions.md) for the current gap list.
+
+## Who is this for?
+
+| Audience | Primary entry point |
+|----------|---------------------|
+| Frontend / full-stack developers | [Web application](./03-web-application.md), [Monorepo structure](./02-monorepo-structure.md) |
+| Data / analytics contributors | [Data pipeline](./05-data-pipeline.md), [Data sources & schemas](./07-data-sources-and-schemas.md) |
+| API / infra contributors | [Backend API](./04-backend-api.md), [Monorepo structure](./02-monorepo-structure.md) |
+| New contributors | This page → [Monorepo structure](./02-monorepo-structure.md) → area of interest |
+
+## System at a glance
+
+```mermaid
+flowchart LR
+ subgraph user [User]
+ Browser[Browser]
+ end
+
+ subgraph web [Web tier — active]
+ Next["Next.js app
/parking-insights"]
+ Config["/api/v1/public/config"]
+ end
+
+ subgraph external [External services]
+ Socrata["LA Open Data
Socrata 4f5p-udkv"]
+ Mapbox["Mapbox
maps + geocoding"]
+ end
+
+ subgraph api [API tier — legacy / unused by web]
+ Express["Express API"]
+ MongoDB[(MongoDB)]
+ end
+
+ subgraph data [Data tier — beta pipeline]
+ Pipeline["beta_pipeline
Python + Polars"]
+ SQLite[(SQLite)]
+ PostGIS[(PostGIS)]
+ end
+
+ Browser --> Next
+ Next --> Config
+ Next --> Socrata
+ Next --> Mapbox
+
+ Express --> MongoDB
+ Socrata --> Pipeline
+ Pipeline --> SQLite
+ Pipeline --> PostGIS
+
+ Next -.->|"planned integration"| Express
+```
+
+## Three parallel data paths
+
+The same underlying citation records can flow through three different paths today:
+
+### 1. Web app → Socrata (live, limited)
+
+The Next.js app queries the [Socrata Query API v3](https://dev.socrata.com/docs/queries/) directly. Users pick a date range and up to two geographic areas; the app fetches matching citations and renders them on a Mapbox map.
+
+**Limitation:** Results are capped at **50 rows** per query (pagination not implemented). See [Web application](./03-web-application.md).
+
+### 2. Beta pipeline → SQLite / PostGIS (local analytics)
+
+The `data-science/beta_pipeline/` Python tools stream the ~6 GB CSV into SQLite or PostGIS using Polars batches. SQLite supports **incremental API sync**; PostGIS supports **spatial queries** and a cleaned analytics table.
+
+**Limitation:** PostGIS does not yet have API sync. The web app does not read from these databases. See [Data pipeline](./05-data-pipeline.md).
+
+### 3. Legacy pipeline → normalized PostGIS (older dataset)
+
+The `data-science/src/data/` scripts and Makefile target an **older dataset ID** (`wjz9-h9np`) and a normalized relational schema (`citation`, `vehicle`, `make`, etc.). This path predates the beta pipeline.
+
+**Limitation:** Dataset ID mismatch with current web app and beta pipeline. See [Legacy data science](./06-legacy-data-science.md).
+
+## Technology summary
+
+| Layer | Stack |
+|-------|-------|
+| Monorepo | Turborepo, pnpm workspaces |
+| Web | Next.js 16, React 19, TypeScript, Tailwind 4, Mapbox GL |
+| Shared UI | `@lucky-parking/design` (Radix-based components) |
+| API | Express 4, MongoDB driver, Zod validation |
+| Beta pipeline | Python 3.12, Polars, psycopg, uv (recommended) |
+| Legacy data science | pandas, geopandas, SQLAlchemy, Conda/Makefile |
+| Spatial DB | PostGIS 16 (Docker), SQLite (file) |
+| CI | GitHub Actions (lint, format, build, test) |
+
+## Primary dataset
+
+All current-facing work uses the LA City Open Data dataset **[Parking Citations (`4f5p-udkv`)](https://data.lacity.org/Transportation/Parking-Citations/4f5p-udkv)** — millions of rows, ~6 GB as CSV, updated on a rolling basis by the city.
+
+Details: [Data sources & schemas](./07-data-sources-and-schemas.md).
+
+## Getting started (short)
+
+**Web app only:**
+
+```bash
+pnpm install
+cp apps/web/.env.schema apps/web/.env # add Mapbox + Socrata tokens
+cd apps/web && pnpm dev
+```
+
+Open `/parking-insights`.
+
+**Data pipeline:**
+
+```bash
+cd data-science/beta_pipeline
+uv venv --python 3.12 .venv
+uv pip install -r requirements.txt jupyter ipykernel --python .venv/bin/python
+```
+
+See [Data pipeline](./05-data-pipeline.md) for full workflows.
+
+## Document map
+
+```mermaid
+mindmap
+ root((Lucky Parking docs))
+ Overview
+ Mission
+ Three data paths
+ Monorepo
+ apps/web
+ apps/api
+ packages
+ Web app
+ Map + filters
+ Socrata queries
+ Zustand state
+ API
+ MongoDB GeoJSON
+ OpenAPI spec
+ Beta pipeline
+ SQLite sync
+ PostGIS clean table
+ Legacy DS
+ Makefile ETL
+ Old dataset ID
+ Data dictionary
+ 23 columns
+ Clean schema
+ Roadmap
+ TODOs
+ Integration gaps
+```
diff --git a/docs/02-monorepo-structure.md b/docs/02-monorepo-structure.md
new file mode 100644
index 00000000..da324807
--- /dev/null
+++ b/docs/02-monorepo-structure.md
@@ -0,0 +1,186 @@
+# Monorepo structure
+
+Lucky Parking is a [Turborepo](https://turbo.build/) monorepo managed with **pnpm workspaces**. One install at the root pulls dependencies for all apps and packages.
+
+## Directory layout
+
+```
+lucky-parking/
+├── apps/
+│ ├── web/ # Next.js frontend (primary user-facing app)
+│ └── api/ # Express REST API (MongoDB)
+├── packages/
+│ ├── design/ # Shared React UI component library
+│ └── configs/ # Shared ESLint, Prettier, TypeScript, Tailwind configs
+├── data-science/
+│ ├── beta_pipeline/ # Modern Polars + SQLite/PostGIS pipeline
+│ ├── src/data/ # Legacy ETL scripts
+│ ├── notebooks/ # Jupyter notebooks (exploratory + archived)
+│ ├── references/ # Lookup tables (makes, violation codes, regex)
+│ ├── db/ # Legacy PostGIS schema dump
+│ └── docker/ # Conda + Jupyter Lab container
+├── docs/ # This documentation set
+├── scripts/ # Shared tooling (e.g. typecheck.mjs)
+├── .github/workflows/ # CI (integration, compliance)
+├── package.json # Root scripts via Turbo
+├── pnpm-workspace.yaml
+└── turbo.json
+```
+
+## Workspace packages
+
+```mermaid
+graph TB
+ subgraph apps [apps/]
+ WEB["@lucky-parking/web"]
+ API["@lucky-parking/api"]
+ end
+
+ subgraph packages [packages/]
+ DESIGN["@lucky-parking/design"]
+ CONFIGS["@lucky-parking/configs"]
+ end
+
+ WEB --> DESIGN
+ WEB --> CONFIGS
+ API --> CONFIGS
+ DESIGN --> CONFIGS
+```
+
+| Package | Path | Role |
+|---------|------|------|
+| `@lucky-parking/web` | `apps/web` | Next.js map application |
+| `@lucky-parking/api` | `apps/api` | Express citations API |
+| `@lucky-parking/design` | `packages/design` | Buttons, sidebar, calendar, dialog, etc. |
+| `@lucky-parking/configs` | `packages/configs` | Shared lint/format/TS/tailwind presets |
+
+## Turbo task pipeline
+
+Root `package.json` delegates to Turbo:
+
+| Command | Effect |
+|---------|--------|
+| `pnpm install` | Install all workspace dependencies |
+| `pnpm dev` | Start dev servers (persistent, uncached) |
+| `pnpm build` | Build apps/packages; depends on upstream `^build` |
+| `pnpm lint` | ESLint across workspaces |
+| `pnpm format` | Prettier |
+| `pnpm test` | Test task (limited coverage today) |
+| `pnpm check-types` | TypeScript checking |
+
+`turbo.json` defines task dependencies — for example, `build` waits for dependencies' builds and outputs `.next/` or `dist/`.
+
+## Prerequisites
+
+| Tool | Version (as documented) | Used for |
+|------|---------------------------|----------|
+| Node.js | 22+ (root); API Docker uses 22 | Web, API, tooling |
+| pnpm | 9 | Package management |
+| Python | 3.10+ (3.12 recommended) | `beta_pipeline` |
+| uv | latest | Recommended Python env manager |
+| Docker | optional | PostGIS for beta pipeline; legacy Jupyter image |
+
+## Environment files
+
+Each app documents its own env vars via `.env.schema`:
+
+| App / area | Schema file | Key variables |
+|------------|-------------|---------------|
+| Web | `apps/web/.env.schema` | `MAPBOX_ACCESS_TOKEN`, `SOCRATA_APP_TOKEN` |
+| API | `apps/api/.env.schema` | `DB_*`, `API_VERSION`, `COL_NAME_CITATIONS` |
+| Beta pipeline | documented in README | `SOCRATA_APP_TOKEN`, `DATABASE_URL` |
+
+**Note:** `.env` files are gitignored. Copy from `.env.schema` and fill in tokens.
+
+**Known issue:** The API code reads `COL_CITATIONS` but the schema documents `COL_NAME_CITATIONS`. See [Backend API](./04-backend-api.md).
+
+## Git hooks and CI
+
+```mermaid
+flowchart LR
+ Commit[git commit] --> Husky[Husky pre-commit]
+ Husky --> LS[lint-staged]
+ LS --> Prettier
+ LS --> ESLint
+ LS --> Typecheck
+
+ Push[git push / PR] --> GHA[GitHub Actions]
+ GHA --> Format
+ GHA --> Lint
+ GHA --> Build
+ GHA --> Test
+```
+
+- **Pre-commit:** Husky runs lint-staged (Prettier, ESLint, typecheck on staged files)
+- **CI:** `.github/workflows/integration.yaml` — format, lint, build, test on PRs
+- **Compliance:** `.github/workflows/compliance.yaml` — PR/issue linking rules
+
+## Running locally
+
+### Web (most common)
+
+```bash
+pnpm install
+cp apps/web/.env.schema apps/web/.env
+# Edit .env with Mapbox and Socrata tokens
+cd apps/web && pnpm dev
+```
+
+The root route redirects to `/parking-insights` (`apps/web/next.config.ts`).
+
+### API (standalone)
+
+```bash
+cd apps/api
+# Configure MongoDB connection via .env
+pnpm dev # or see apps/api package.json
+```
+
+The web app does **not** call this API today.
+
+### Data pipeline
+
+See [Data pipeline (beta)](./05-data-pipeline.md). Lives outside the Node workspace but shares the repo.
+
+## What is not in the monorepo
+
+| Item | Location / note |
+|------|-----------------|
+| Citation CSV files | Downloaded locally; gitignored (`raw_data/`, etc.) |
+| SQLite DB files | Generated by pipeline; not committed |
+| Python `.venv` | Created per contributor; gitignored |
+| Production deploy config | Referenced in OpenAPI (`luckyparking.org`) but not fully documented in repo |
+
+## Package boundaries (design intent)
+
+```mermaid
+flowchart TB
+ subgraph presentation [Presentation]
+ WEB[apps/web]
+ end
+
+ subgraph shared [Shared libraries]
+ DESIGN[packages/design]
+ end
+
+ subgraph services [Services — partially used]
+ API[apps/api]
+ end
+
+ subgraph analytics [Analytics — offline]
+ BETA[beta_pipeline]
+ LEG[legacy data-science]
+ end
+
+ WEB --> DESIGN
+ WEB --> Socrata[Socrata API]
+ API --> Mongo[(MongoDB)]
+ BETA --> SQLite[(SQLite)]
+ BETA --> PostGIS[(PostGIS)]
+ LEG --> PostGIS
+
+ WEB -.-> API
+ WEB -.-> BETA
+```
+
+The monorepo currently optimizes for **shared frontend tooling** and **co-located data work**. Full-stack integration (web ↔ API ↔ local DB) remains future work.
diff --git a/docs/03-web-application.md b/docs/03-web-application.md
new file mode 100644
index 00000000..e12f44ba
--- /dev/null
+++ b/docs/03-web-application.md
@@ -0,0 +1,218 @@
+# Web application
+
+The primary user-facing application lives in `apps/web` — a **Next.js 16** app that visualizes LA parking citations on an interactive Mapbox map.
+
+**Route:** `/parking-insights` (root `/` redirects here)
+
+## User experience
+
+Users can:
+
+1. Select a **date range** (default: last 7 days)
+2. Search and add up to **two places** (neighborhood councils, addresses via Mapbox geocoding)
+3. View citations as **circle markers** and an optional **heatmap** on the map
+4. See **aggregate stats** (total citations, total/average fines) in a sidebar panel
+5. Click map features for citation detail popups
+
+```mermaid
+flowchart TB
+ subgraph sidebar [Sidebar — DataPanel]
+ Places[Place search + list]
+ Dates[Date range picker]
+ Stats[DataVisuals — totals]
+ Legend[DataLegend]
+ About[About section]
+ end
+
+ subgraph map [Map — ParkingCitationsMap]
+ Source[MapSourceCitations — GeoJSON]
+ Circles[MapLayerCircles]
+ Heat[MapLayerHeatmap]
+ Popup[Click popups]
+ end
+
+ subgraph state [Client state]
+ Store[Zustand store]
+ RQ[React Query — useCitations]
+ end
+
+ Places --> Store
+ Dates --> Store
+ Store --> RQ
+ RQ --> Source
+ Source --> Circles
+ Source --> Heat
+ Circles --> Popup
+```
+
+## Key files
+
+| Path | Purpose |
+|------|---------|
+| `src/app/parking-insights/page.tsx` | Main page layout (sidebar + map) |
+| `src/app/layout.tsx` | Root layout, fonts, global styles |
+| `src/app/api/v1/public/config/route.ts` | Serves public tokens to the client |
+| `src/store.ts` | Zustand store (query, range, places) |
+| `src/hooks/use-citations.tsx` | React Query hook wrapping Socrata fetch |
+| `src/hooks/use-geocoder.tsx` | Mapbox + neighborhood council search |
+| `src/hooks/use-public-config.ts` | Fetches Mapbox/Socrata tokens |
+| `src/lib/socrata/parking-citations.ts` | Builds SoQL query, POSTs to Socrata |
+| `src/components/map.tsx` | Map container |
+| `src/components/map-source-citations.tsx` | GeoJSON source layer |
+| `src/components/data-visuals.tsx` | Citation/fine statistics |
+| `src/data/los-angeles*.json` | City/county boundary GeoJSON |
+
+## State management
+
+The app uses **Zustand** with Immer, persist, and devtools middleware (`src/store.ts`).
+
+| State | Type | Default | Persisted? |
+|-------|------|---------|------------|
+| `query` | string | `""` | No |
+| `range` | `{ from, to }` | Last 7 days | Yes (localStorage) |
+| `places` | `Map` | empty | Yes |
+| `isHydrated` | boolean | false | No |
+
+**Constraints:**
+
+- Maximum **2 places** (`MAX_PLACES = 2`)
+- Date range normalized to start/end of day via `date-fns`
+
+Persistence key: `luckyparking` in `localStorage`.
+
+## Data fetching flow
+
+```mermaid
+sequenceDiagram
+ participant User
+ participant Store as Zustand store
+ participant Hook as useCitations
+ participant Config as /api/v1/public/config
+ participant Socrata as Socrata Query API v3
+
+ User->>Store: Set date range / places
+ Store->>Hook: places + range change
+ Hook->>Config: GET tokens (on load)
+ Config-->>Hook: mapboxAccessToken, socrataAppToken
+ Hook->>Socrata: POST query (GeoJSON)
+ Note over Socrata: pageSize: 50 (hard limit)
+ Socrata-->>Hook: FeatureCollection
+ Hook-->>User: Map layers update
+```
+
+### Socrata integration
+
+**Endpoint:** `https://data.lacity.org/api/v3/views/4f5p-udkv/query`
+
+**Query construction** (`parking-citations.ts`):
+
+- Date filter: `issue_date BETWEEN 'YYYY-MM-DD' AND 'YYYY-MM-DD'`
+- Optional geo filter: `within_polygon(geocodelocation, 'WKT')` when places are selected
+- Combines multiple place geometries into a multipolygon via Turf.js
+
+**Headers:** `Accept: application/vnd.geo+json`, `X-App-Token`
+
+**Validation:** Response parsed with Zod (`ParkingCitationFeatureCollectionSchema`)
+
+### Public config API
+
+`GET /api/v1/public/config` exposes server-side env vars to the browser:
+
+- `MAPBOX_ACCESS_TOKEN`
+- `SOCRATA_APP_TOKEN`
+
+The client will not fetch citations unless both a Socrata token and at least one place are present.
+
+## Map layers
+
+| Component | Layer type | Behavior |
+|-----------|------------|----------|
+| `MapSourceCitations` | GeoJSON source | Driven by `useCitations` data |
+| `MapLayerCircles` | Circle layer | Individual citation points |
+| `MapLayerHeatmap` | Heatmap layer | Density visualization |
+| Base map | Mapbox GL | LA county bounds enforced |
+
+Static boundary data under `src/data/` supports geocoding context and map constraints.
+
+## Component hierarchy
+
+```mermaid
+graph TD
+ Page[parking-insights/page.tsx]
+ Page --> Sidebar[AppSidebar]
+ Page --> Header[AppHeader]
+ Page --> Map[ParkingCitationsMap]
+
+ Sidebar --> DataPanel[DataPanel]
+ DataPanel --> SearchGeocoder
+ DataPanel --> SearchDateRange
+ DataPanel --> DataPlaceList
+ DataPanel --> DataVisuals
+ DataPanel --> DataLegend
+
+ Map --> MapSourceCitations
+ Map --> MapLayerCircles
+ Map --> MapLayerHeatmap
+```
+
+UI primitives come from `@lucky-parking/design` (sidebar, accordion, calendar, button, etc.).
+
+## Environment setup
+
+```bash
+cp apps/web/.env.schema apps/web/.env
+```
+
+| Variable | Required | Purpose |
+|----------|----------|---------|
+| `MAPBOX_ACCESS_TOKEN` | Yes | Map tiles and geocoding |
+| `SOCRATA_APP_TOKEN` | Yes (for data) | Raises rate limits; required for queries in current logic |
+
+Free tokens: [Mapbox](https://account.mapbox.com/), [LA City Data / Socrata](https://data.lacity.org/login).
+
+## Known limitations (unfinished)
+
+| Issue | Location | Impact |
+|-------|----------|--------|
+| **50-row pagination cap** | `parking-citations.ts` FIXME | Map shows at most 50 citations per query |
+| **No backend API usage** | architecture | Cannot leverage MongoDB-preloaded data |
+| **Legend placeholder** | `data-legend.tsx` TODO | Legend may not reflect real violation categories |
+| **Geocoder edge cases** | `use-geocoder.tsx` TODO | Address/postcode result types not fully handled |
+| **No offline/local DB mode** | — | Pipeline databases not wired to web |
+
+See [Roadmap](./08-roadmap-and-open-questions.md) for planned fixes and integration options.
+
+## Possible future directions
+
+```mermaid
+flowchart LR
+ Today[Today: Socrata direct
50 rows max]
+
+ OptA[Option A: Pagination
fetch all pages client-side]
+ OptB[Option B: API proxy
apps/api + MongoDB/PostGIS]
+ OptC[Option C: Vector tiles
pre-aggregated map layers]
+ OptD[Option D: Hybrid
Socrata for recent, DB for bulk]
+
+ Today --> OptA
+ Today --> OptB
+ Today --> OptC
+ Today --> OptD
+```
+
+| Direction | Pros | Cons |
+|-----------|------|------|
+| Client pagination | Smallest change | Slow for large date ranges; rate limits |
+| Express API + DB | Full control, fast queries | Requires ingestion + deployment |
+| Vector tiles / MVT | Best map performance at scale | New infra pipeline |
+| Hybrid recent + historical | Good UX split | Two code paths to maintain |
+
+## Running and building
+
+```bash
+cd apps/web
+pnpm dev # development server
+pnpm build # production build
+pnpm lint # ESLint
+```
+
+From repo root: `pnpm dev` starts all configured dev tasks via Turbo.
diff --git a/docs/04-backend-api.md b/docs/04-backend-api.md
new file mode 100644
index 00000000..b9336d07
--- /dev/null
+++ b/docs/04-backend-api.md
@@ -0,0 +1,181 @@
+# Backend API
+
+The `apps/api` package is an **Express 4** REST service that queries parking citations from **MongoDB** using geospatial and date filters.
+
+**Important:** The current Next.js web app does **not** use this API. It queries Socrata directly. Treat this service as **legacy / alternate backend** until integration work is done.
+
+## Purpose (intended)
+
+Provide a stable API for filtering citations by:
+
+- **Date range** — `issue_date` between two ISO datetimes
+- **Geography** — GeoJSON polygon (`$geoWithin`)
+
+This would allow pre-indexed queries over a full local copy of citations instead of live Socrata calls.
+
+## Architecture
+
+```mermaid
+flowchart LR
+ Client[HTTP client]
+ Express[Express app]
+ Validator[Zod middleware]
+ Controller[CitationController]
+ Service[CitationService]
+ Mongo[(MongoDB collection)]
+
+ Client -->|POST /v1/citations| Express
+ Express --> Validator
+ Validator --> Controller
+ Controller --> Service
+ Service --> Mongo
+```
+
+## Routes
+
+| Method | Path | Handler | Description |
+|--------|------|---------|-------------|
+| GET | `/{API_VERSION}/` | inline | Health/hello |
+| POST | `/{API_VERSION}/citations` | `CitationController.listCitations` | Filter citations |
+
+`API_VERSION` defaults from env (see `.env.schema`).
+
+## Request contract
+
+**Body** (validated by `CitationFiltersSchema`):
+
+```json
+{
+ "dates": ["2025-01-01T00:00:00.000Z", "2025-01-31T23:59:59.999Z"],
+ "geometry": {
+ "type": "Polygon",
+ "coordinates": [[[...]]]
+ }
+}
+```
+
+Both fields are optional. Empty filters return broader result sets (subject to MongoDB query logic).
+
+## MongoDB query logic
+
+`CitationService.findCitations()` (`src/services/citations.ts`):
+
+```javascript
+{
+ $and: [
+ { issue_date: { $gte: dates[0], $lte: dates[1] } }, // if dates provided
+ { geometry: { $geoWithin: { $geometry: geometry } } }, // if geometry provided
+ { "geometry.coordinates": { $nin: [null] } }
+ ]
+}
+```
+
+Documents are expected to store citation geometry in a GeoJSON-compatible shape.
+
+## OpenAPI specification
+
+Full contract: [`apps/api/src/docs/specs-v1.yaml`](../apps/api/src/docs/specs-v1.yaml)
+
+- Documented server: `https://luckyparking.org/api/v1`
+- Response shape: `{ data: CitationFeatureCollection }` (GeoJSON FeatureCollection)
+- Includes RFC 7946 GeoJSON schema components
+
+**Stale reference:** External docs link may still point to older dataset `wjz9-h9np` while the web app and beta pipeline use `4f5p-udkv`.
+
+## Environment variables
+
+From `apps/api/.env.schema`:
+
+| Variable | Purpose |
+|----------|---------|
+| `API_VERSION` | URL prefix (e.g. `v1`) |
+| `DB_USERNAME` | MongoDB user |
+| `DB_PASSWORD` | MongoDB password |
+| `DB_HOST` | MongoDB host |
+| `DB_NAME` | Database name |
+| `COL_NAME_CITATIONS` | Collection name (documented) |
+
+**Bug:** Runtime code reads `COL_CITATIONS`, not `COL_NAME_CITATIONS`:
+
+```typescript
+const { COL_CITATIONS } = process.env;
+// ...
+db.collection(COL_CITATIONS as string)
+```
+
+Contributors must set the variable name the code actually uses, or fix the mismatch.
+
+## Key source files
+
+| File | Role |
+|------|------|
+| `src/index.ts` | Server entry, MongoDB connect, listen |
+| `src/app.ts` | Express app setup, routes, middleware |
+| `src/controllers/citations.ts` | HTTP handler |
+| `src/services/citations.ts` | MongoDB query |
+| `src/database/client.ts` | Mongo client |
+| `src/middleware/validator.ts` | Zod request validation |
+| `src/utilities/schemas.ts` | Zod schemas (GeoPolygon is `z.any()` — TODO) |
+
+## Relationship to other systems
+
+```mermaid
+flowchart TB
+ subgraph current [Current production path]
+ WEB[apps/web]
+ SOC[Socrata 4f5p-udkv]
+ WEB --> SOC
+ end
+
+ subgraph api_path [API path — not connected]
+ API[apps/api]
+ MONGO[(MongoDB)]
+ API --> MONGO
+ end
+
+ subgraph pipeline [Ingestion options]
+ BETA[beta_pipeline]
+ LEG[legacy upload scripts]
+ BETA --> SQLITE[(SQLite)]
+ BETA --> PG[(PostGIS)]
+ LEG --> PG
+ end
+
+ WEB -.->|"future"| API
+ BETA -.->|"export / ETL"| MONGO
+ LEG -.-> MONGO
+```
+
+No automated job today loads `beta_pipeline` output into MongoDB for the API.
+
+## Unfinished / open items
+
+| Item | Status |
+|------|--------|
+| Web app integration | Not started |
+| Env var naming (`COL_CITATIONS`) | Bug / inconsistency |
+| Full GeoPolygon Zod schema | TODO in `schemas.ts` |
+| Dataset alignment (`4f5p-udkv` vs `wjz9-h9np`) | Needs migration plan |
+| Ingestion pipeline → MongoDB | Not implemented in beta_pipeline |
+| Automated tests | Minimal / absent at app level |
+
+## Possible directions
+
+1. **Retire the API** — If Socrata + local PostGIS meet all needs, deprecate MongoDB path and document removal.
+2. **Revive as BFF** — Express proxies to PostGIS (or SQLite) with the same filter contract as OpenAPI; web app switches from Socrata.
+3. **MongoDB as cache** — Scheduled job syncs Socrata or CSV into Mongo for GeoJSON-native `$geoWithin` queries.
+4. **Unify on PostGIS** — Single spatial database serves both analytics notebooks and a thin API layer (PostgREST, custom Express, or Next.js route handlers).
+
+Each option trades operational complexity against query performance and data freshness. See [Roadmap](./08-roadmap-and-open-questions.md).
+
+## Running locally
+
+Configure MongoDB connection and collection name, then:
+
+```bash
+cd apps/api
+pnpm install # from root is preferred
+pnpm dev # see package.json for exact script
+```
+
+Test with `POST /v1/citations` and a JSON body matching the OpenAPI spec.
diff --git a/docs/05-data-pipeline.md b/docs/05-data-pipeline.md
new file mode 100644
index 00000000..09e059e0
--- /dev/null
+++ b/docs/05-data-pipeline.md
@@ -0,0 +1,274 @@
+# Data pipeline (beta)
+
+The **`data-science/beta_pipeline/`** directory contains the project's **modern data ingestion stack** — Python tools built on **Polars** for streaming multi-gigabyte CSV files into **SQLite** or **PostGIS**, with optional API sync (SQLite only).
+
+For module-level function reference, also see [`data-science/beta_pipeline/ARCHITECTURE.md`](../data-science/beta_pipeline/ARCHITECTURE.md).
+
+## Design goals
+
+| Goal | How it's achieved |
+|------|-------------------|
+| Handle ~6 GB CSV without OOM | Polars batched reads (100k rows default) |
+| Safe re-runs | `INSERT OR IGNORE` / `ON CONFLICT DO NOTHING` |
+| Keep data fresh | Socrata incremental sync (SQLite) |
+| Spatial analytics | PostGIS geometry + GiST indexes + clean table |
+| Simple local dev | Docker Compose for PostGIS; SQLite is a single file |
+
+## High-level architecture
+
+```mermaid
+flowchart TB
+ subgraph sources [Sources]
+ CSV["Parking_Citations_*.csv"]
+ API["Socrata Resource API
4f5p-udkv.json"]
+ end
+
+ subgraph shared [Shared parsing — parking_db.py]
+ BATCH["_csv_batches()"]
+ NORM["_normalize_csv_batch()"]
+ COLS["COLUMNS — 23 fields"]
+ end
+
+ subgraph sqlite [SQLite path]
+ PDB[parking_db.py]
+ DB[(parking_citations.db)]
+ SYNC[update_from_api]
+ end
+
+ subgraph postgis [PostGIS path]
+ PPG[parking_postgis.py]
+ RAW[(citations)]
+ PCL[parking_clean.py]
+ CLEAN[(citations_clean)]
+ PPL[parking_pipeline.py]
+ end
+
+ CSV --> BATCH
+ BATCH --> NORM
+ NORM --> PDB
+ NORM --> PPG
+ API --> SYNC
+ PDB --> DB
+ SYNC --> DB
+ PPG --> RAW
+ RAW --> PCL
+ PCL --> CLEAN
+ PPL --> PPG
+ PPL --> PCL
+```
+
+## Modules
+
+### `parking_db.py` — SQLite + API sync
+
+Standalone module: CSV load, SQLite schema, Socrata incremental sync.
+
+| CLI command | Action |
+|-------------|--------|
+| `init` | Create tables and indexes |
+| `load-csv FILE` | Stream CSV → SQLite |
+| `sync` | Incremental API sync (newest first) |
+| `stats` | Row counts, max date, last sync |
+
+**Sync algorithm:**
+
+```mermaid
+flowchart TD
+ Start[offset = 0] --> Fetch[Fetch page ORDER BY :updated_at DESC]
+ Fetch --> Empty{Page empty?}
+ Empty -->|yes| Done[Caught up]
+ Empty -->|no| Check[Find existing ticket_numbers]
+ Check --> Insert[INSERT OR IGNORE new rows]
+ Insert --> Match{Any ticket on page
already in DB?}
+ Match -->|yes| Done
+ Match -->|no| Next[offset += page_size]
+ Next --> Fetch
+```
+
+### `parking_postgis.py` — Raw PostGIS store
+
+Reuses CSV parsing from `parking_db.py`. Adds `geom geometry(Point, 4326)` built at insert time:
+
+1. Prefer WKT from `geocodelocation`
+2. Fallback to `ST_MakePoint(loc_long, loc_lat)`
+
+| CLI command | Action |
+|-------------|--------|
+| `init` | Enable PostGIS extension, create tables |
+| `load-csv FILE [--clean]` | Bulk load; optional clean rebuild |
+| `stats` | Counts, geometry coverage |
+
+### `parking_clean.py` — Analytics table
+
+Builds slim `citations_clean` from raw `citations`:
+
+- Combines `issue_date` + `issue_time` (HHMM) → `issue_datetime`
+- Trims violation text fields
+- Keeps ticket, violation, fine, geometry
+
+| CLI command | Action |
+|-------------|--------|
+| `init` | Create clean table |
+| `rebuild` | `TRUNCATE` + full `INSERT` from raw |
+| `stats` | Row count, datetime range, geom coverage |
+
+### `parking_pipeline.py` — Orchestration
+
+| CLI command | Action |
+|-------------|--------|
+| `run FILE` | Load CSV → rebuild clean → print stats |
+| `clean` | Rebuild clean only (no CSV load) |
+
+## PostGIS full pipeline
+
+```mermaid
+sequenceDiagram
+ participant User
+ participant Docker as docker compose
+ participant PPL as parking_pipeline.py
+ participant PPG as parking_postgis.py
+ participant PCL as parking_clean.py
+ participant PG as PostGIS
+
+ User->>Docker: docker compose up -d
+ User->>PPL: run Parking_Citations.csv
+ PPL->>PPG: bulk_load_csv()
+ loop Each 100k batch
+ PPG->>PG: INSERT citations + geom
+ end
+ PPL->>PCL: rebuild_clean()
+ PCL->>PG: TRUNCATE citations_clean
+ PCL->>PG: INSERT cleaned rows
+ PPL-->>User: JSON stats (raw + clean)
+```
+
+## Docker setup
+
+`docker-compose.yml` runs **PostGIS 16**:
+
+| Setting | Value |
+|---------|-------|
+| Image | `postgis/postgis:16-3.4` |
+| Port | `5432:5432` |
+| Credentials | `parking` / `parking` / `parking` |
+| Volume | `postgis_data` (persistent) |
+| Healthcheck | `pg_isready` every 5s |
+
+Default DSN: `postgresql://parking:parking@localhost:5432/parking`
+
+## Environment
+
+| Variable | Used by | Default |
+|----------|---------|---------|
+| `SOCRATA_APP_TOKEN` | `parking_db.py sync` | none (anonymous API) |
+| `DATABASE_URL` | PostGIS modules | local Docker DSN above |
+
+**Gap:** README references `.env.example` but the file is not in the repo yet.
+
+## Python environment
+
+Recommended setup with **uv**:
+
+```bash
+cd data-science/beta_pipeline
+uv venv --python 3.12 .venv
+uv pip install -r requirements.txt jupyter ipykernel --python .venv/bin/python
+```
+
+**Dependencies** (`requirements.txt`):
+
+| Package | Role |
+|---------|------|
+| `polars` | CSV streaming, analytics |
+| `python-dotenv` | Load `.env` |
+| `psycopg[binary]` | PostGIS driver |
+| `pandas`, `numpy`, `matplotlib`, `seaborn`, `geopandas` | Notebooks |
+
+`.venv/` is gitignored.
+
+## Notebooks
+
+| Notebook | Status |
+|----------|--------|
+| `parking_db_explore.ipynb` | **Active** — Polars exploration of CSV; documents SQLite workflow |
+| `violation_analysis.ipynb` | **Stub** — imports only; analysis not written yet |
+
+## Module dependency graph
+
+```
+parking_db.py (standalone)
+ ↑
+parking_postgis.py (imports CSV helpers)
+ ↑
+parking_clean.py (imports connect, init_db)
+ ↑
+parking_pipeline.py (orchestrates load + clean)
+```
+
+## Design tradeoffs
+
+| Decision | Benefit | Cost |
+|----------|---------|------|
+| `INSERT OR IGNORE` | Idempotent bulk loads | Ticket corrections never update existing rows |
+| Separate `citations_clean` | Simple analytics schema | Full rebuild on each clean (no incremental upsert) |
+| SQLite fast pragmas | Faster bulk load | Less crash safety during load |
+| PostGIS `synchronous_commit=OFF` | Faster ingest | Recent commits may be lost on crash |
+
+## Unfinished work
+
+| Feature | Status | Notes |
+|---------|--------|-------|
+| PostGIS API sync | **Not implemented** | Copy/adapt `update_from_api` from SQLite |
+| Incremental clean rebuild | **Not implemented** | Today: full `TRUNCATE` + insert |
+| Scheduled automation | **Not implemented** | cron/launchd after CSV download |
+| Web app integration | **Not implemented** | No API reads from SQLite/PostGIS |
+| `.env.example` | **Missing file** | Documented in README only |
+| `violation_analysis.ipynb` | **Empty** | Planned analysis TBD |
+
+## Possible directions
+
+```mermaid
+mindmap
+ root((beta_pipeline future))
+ Sync
+ PostGIS API sync
+ Unified sync module
+ Clean layer
+ Incremental upsert
+ Materialized views
+ Integration
+ Next.js API routes
+ GraphQL over PostGIS
+ Export to Mongo for legacy API
+ Ops
+ GitHub Action nightly sync
+ Cloud RDS / Supabase
+ dbt models on citations_clean
+ Analysis
+ violation_analysis notebook
+ dbt metrics
+ ML feature store
+```
+
+## Quick command reference
+
+**SQLite (no Docker):**
+
+```bash
+.venv/bin/python parking_db.py init
+.venv/bin/python parking_db.py load-csv Parking_Citations_20260426.csv
+.venv/bin/python parking_db.py sync
+.venv/bin/python parking_db.py stats
+```
+
+**PostGIS:**
+
+```bash
+docker compose up -d
+.venv/bin/python parking_pipeline.py run Parking_Citations_20250811.csv
+.venv/bin/python parking_clean.py stats
+```
+
+Replace CSV filenames with your local download from [data.lacity.org](https://data.lacity.org/Transportation/Parking-Citations/4f5p-udkv).
+
+Schema details: [Data sources & schemas](./07-data-sources-and-schemas.md).
diff --git a/docs/06-legacy-data-science.md b/docs/06-legacy-data-science.md
new file mode 100644
index 00000000..f86d6e58
--- /dev/null
+++ b/docs/06-legacy-data-science.md
@@ -0,0 +1,203 @@
+# Legacy data science
+
+Before `beta_pipeline/`, the project used a **Cookiecutter-style data science layout** under `data-science/` with Makefile-driven ETL, normalized PostGIS schemas, and extensive Jupyter notebooks.
+
+This stack is **still in the repo** but is largely **superseded** by the beta pipeline for new work. It targets a **different Socrata dataset ID** and a different database schema.
+
+## Legacy vs beta pipeline
+
+```mermaid
+flowchart LR
+ subgraph legacy [Legacy — data-science/]
+ OLD_DS["Dataset wjz9-h9np"]
+ MK[make_dataset.py]
+ UP[upload_serial.py]
+ NORM[(Normalized schema
citation, vehicle, make...)]
+ end
+
+ subgraph beta [Beta — beta_pipeline/]
+ NEW_DS["Dataset 4f5p-udkv"]
+ PDB[parking_db.py / parking_postgis.py]
+ FLAT[(Flat 23-column schema
citations / citations_clean)]
+ end
+
+ subgraph apps [Applications]
+ WEB[apps/web → 4f5p-udkv]
+ API[apps/api spec → wjz9-h9np reference]
+ end
+
+ OLD_DS --> MK --> UP --> NORM
+ NEW_DS --> PDB --> FLAT
+ WEB --> NEW_DS
+ API -.-> OLD_DS
+```
+
+| Aspect | Legacy | Beta pipeline |
+|--------|--------|---------------|
+| Dataset ID | `wjz9-h9np` | `4f5p-udkv` |
+| Primary tool | pandas / geopandas | Polars |
+| DB schema | Normalized relational | Flat citation table + clean layer |
+| API sync | Via older scripts / manual | SQLite `sync` command |
+| Recommended for new work? | No | **Yes** |
+
+## Directory layout
+
+```
+data-science/
+├── src/data/ # ETL scripts (click/Makefile driven)
+├── notebooks/
+│ ├── exploratory/ # Active exploration notebooks
+│ └── archived_notebooks/ # Older ML, viz, upload experiments
+├── references/ # Lookup tables and regex rules
+├── db/ # db_dev.sql — legacy schema dump
+├── docker/ # Conda + Jupyter Lab image
+├── old_docker/ # Deprecated Dockerfiles
+├── docs/ # Sphinx documentation (partially stale)
+├── Makefile # Primary automation interface
+├── requirements.txt # Legacy Python deps
+└── beta_pipeline/ # Modern pipeline (see separate doc)
+```
+
+## Makefile workflow
+
+The Makefile (`data-science/Makefile`) is the main entry point for legacy tasks:
+
+| Target | Action |
+|--------|--------|
+| `make requirements` | Install Python dependencies |
+| `make data` | Run `make_dataset.py` (raw → processed) |
+| `make sample` | Create sample datasets |
+| `make serial_data` | Build serial-friendly output |
+| `make upload_serial` | Upload to PostGIS |
+| `make upload_geojson` | Upload GeoJSON via `upload.py` |
+| `make upload_zip` | Upload zipcode boundaries |
+| `make upload_neighborhood` | Upload neighborhood councils |
+| `make lint` | flake8 on `src/` |
+| `make clean` / `make clean_data` | Remove caches / data files |
+
+Requires Conda or system Python (`PYTHON_INTERPRETER = python3`).
+
+## Key ETL scripts (`src/data/`)
+
+| Script | Purpose |
+|--------|---------|
+| `make_dataset.py` | Download/process raw citation CSV |
+| `make_dataset_dask.py` | Dask-based variant for large files |
+| `make_serial_data.py` | Prepare serial upload format |
+| `upload_serial.py` | Load processed data into PostGIS |
+| `upload.py` | Upload GeoJSON layers |
+| `upload_neighborhood.py` | Neighborhood council boundaries |
+| `get_zipcodes.py` | Zipcode boundary data |
+| `sample.py` | Generate sample subsets |
+| `date_threshold.py` | Filter by date threshold |
+
+These scripts expect `.env` with `DB_USER`, `DB_PASSWORD`, `DB_HOST`, `DB_PORT`, `DB_DATABASE`.
+
+## Legacy PostGIS schema
+
+`db/db_dev.sql` defines a **normalized** model:
+
+```mermaid
+erDiagram
+ citation ||--o| vehicle : has
+ citation }o--|| codes : violation
+ vehicle }o--o| make : references
+ citation }o--o| zipcodes : located_in
+ neighborhood_councils ||--o{ citation : contains
+
+ citation {
+ text ticket_number PK
+ timestamp issue_date
+ geometry geom
+ }
+ vehicle {
+ text vin
+ text make
+ text color
+ }
+ codes {
+ text violation_code
+ text description
+ }
+ make {
+ text make_name
+ }
+```
+
+This differs from beta_pipeline's flat `citations` table with 23 city-export columns.
+
+## Reference data (`references/`)
+
+Lookup and normalization files used by legacy cleaning:
+
+| File | Purpose |
+|------|---------|
+| `make.csv`, `makes.json`, `top_makes.txt` | Vehicle make normalization |
+| `violation_codes.json`, `violation_descriptions.json` | Violation lookup |
+| `vio_regex.csv` | Regex rules for violation descriptions |
+| `column_names.json` | Output column mapping |
+| `top_violation_codes.txt` | Common codes list |
+
+These may be useful for future cleaning logic in `beta_pipeline` or `violation_analysis.ipynb` but are not wired into the beta pipeline today.
+
+## Notebooks
+
+### Exploratory (`notebooks/exploratory/`)
+
+Active-ish exploration notebooks (e.g. `1-gp-explore_raw.ipynb`, `1-fl-analysis.ipynb`).
+
+### Archived (`notebooks/archived_notebooks/`)
+
+Historical work including:
+
+- Google Maps citation visualization
+- Random Forest / zip code models
+- Reddit data (PRAW)
+- Server upload experiments
+- Regulation sweeping exploration
+
+Treat archived notebooks as **historical context**, not current runbooks.
+
+## Docker (legacy Jupyter)
+
+`data-science/docker/` provides a Conda-based Jupyter Lab image (port **8888**). See `data-science/docker/README.md`.
+
+`old_docker/` contains deprecated Dockerfiles — do not use for new work.
+
+## Sphinx docs
+
+`data-science/docs/` contains Sphinx scaffolding. Some documented commands (e.g. S3 sync) are **not present in the Makefile** — docs may be stale.
+
+## When to use legacy vs beta
+
+| Use case | Recommendation |
+|----------|----------------|
+| New CSV ingestion for LA citations | **beta_pipeline** |
+| Spatial analytics on flat schema | **beta_pipeline** + PostGIS |
+| Incremental Socrata sync | **beta_pipeline** SQLite |
+| Understanding old ML/viz experiments | Legacy notebooks |
+| Normalized vehicle/violation schema | Legacy (or redesign on top of beta) |
+
+## Migration considerations (unfinished)
+
+No automated migration exists between:
+
+- Legacy normalized PostGIS ↔ beta flat PostGIS
+- Either database ↔ MongoDB (API)
+- Either database ↔ web app
+
+A full migration plan would need to address:
+
+1. **Dataset ID** alignment (`wjz9-h9np` → `4f5p-udkv`)
+2. **Schema mapping** (normalized tables vs 23-column flat + clean)
+3. **Reference data** port (make/violation normalization into clean step)
+4. **Retiring or repointing** legacy Makefile scripts
+
+See [Roadmap](./08-roadmap-and-open-questions.md).
+
+## Possible directions
+
+- **Deprecate legacy ETL** — Archive `src/data/` scripts, keep notebooks for reference only
+- **Port normalization into beta clean step** — Use `references/` in `parking_clean.py`
+- **Unified Makefile or uv project** — Single Python entry point wrapping beta_pipeline CLI
+- **Revive ML notebooks** — Re-run archived models against `citations_clean` in PostGIS
diff --git a/docs/07-data-sources-and-schemas.md b/docs/07-data-sources-and-schemas.md
new file mode 100644
index 00000000..24a62292
--- /dev/null
+++ b/docs/07-data-sources-and-schemas.md
@@ -0,0 +1,229 @@
+# Data sources and schemas
+
+This document describes where citation data comes from, how it is shaped in each system, and known data quality issues.
+
+## Primary dataset (current)
+
+**[Parking Citations — `4f5p-udkv`](https://data.lacity.org/Transportation/Parking-Citations/4f5p-udkv)**
+
+Published by the City of Los Angeles on the Socrata open data portal.
+
+| Property | Value |
+|----------|-------|
+| Approximate size | ~6 GB CSV |
+| Row count | Millions of citations |
+| Update frequency | Rolling updates by city |
+| Used by | Web app, beta_pipeline |
+
+### Access methods
+
+```mermaid
+flowchart TB
+ DS["Dataset 4f5p-udkv"]
+
+ DS --> CSV["CSV export
(Transportation → Parking Citations)"]
+ DS --> RES["Resource API
/resource/4f5p-udkv.json"]
+ DS --> QRY["Query API v3
/api/v3/views/4f5p-udkv/query"]
+
+ CSV --> BETA_LOAD["beta_pipeline load-csv"]
+ RES --> BETA_SYNC["beta_pipeline sync"]
+ QRY --> WEB["apps/web fetchParkingCitations"]
+```
+
+| Method | URL pattern | Consumer |
+|--------|-------------|----------|
+| CSV download | data.lacity.org export | `parking_db.py`, `parking_postgis.py` |
+| Resource API | `https://data.lacity.org/resource/4f5p-udkv.json` | `parking_db.py sync` |
+| Query API v3 | `https://data.lacity.org/api/v3/views/4f5p-udkv/query` | Next.js web app |
+
+Optional **`X-App-Token`** header improves rate limits (free registration at data.lacity.org).
+
+## Legacy dataset
+
+**[`wjz9-h9np`](https://data.lacity.org/)** — older parking citations dataset referenced by:
+
+- Legacy `make_dataset.py` and related ETL
+- OpenAPI external docs in `apps/api` (may be stale)
+
+New work should use **`4f5p-udkv`** unless explicitly migrating historical comparisons.
+
+## Canonical column set (23 fields)
+
+Both SQLite and PostGIS raw tables in `beta_pipeline` store these columns, defined as `COLUMNS` in `parking_db.py`:
+
+| Column | Typical type | Description |
+|--------|--------------|-------------|
+| `ticket_number` | string | Primary key — unique citation ID |
+| `issue_date` | datetime/text | Citation date (CSV: `"2025 Apr 26 12:00:00 AM"`) |
+| `issue_time` | string | Time as HHMM without leading zeros (e.g. `"904"`, `"1430"`) |
+| `meter_id` | string | Parking meter identifier |
+| `marked_time` | string | Marked time field from source |
+| `rp_state_plate` | string | Registered plate state |
+| `plate_expiry_date` | string | Plate expiration |
+| `vin` | string | Vehicle VIN |
+| `make` | string | Vehicle make |
+| `body_style` | string | Body style code |
+| `color` | string | Vehicle color code |
+| `location` | string | Street location description |
+| `route` | string | Route identifier |
+| `agency` | integer | Issuing agency code |
+| `violation_code` | string | Violation code |
+| `violation_description` | string | Human-readable violation |
+| `fine_amount` | float | Fine in dollars |
+| `agency_desc` | string | Agency description |
+| `color_desc` | string | Color description |
+| `body_style_desc` | string | Body style description |
+| `loc_lat` | float | Latitude |
+| `loc_long` | float | Longitude |
+| `geocodelocation` | string | WKT POINT geometry from CSV |
+
+### CSV parsing notes
+
+- Polars schema assigns explicit types per column
+- `null_values`: `""`, `"NA"`, `"N/A"`
+- `ignore_errors=True` — malformed rows skipped, not fatal
+- Batch size default: **100,000 rows**
+
+## PostGIS raw table (`citations`)
+
+All 23 columns plus:
+
+```sql
+geom geometry(Point, 4326)
+```
+
+**Geometry construction priority:**
+
+1. `ST_GeomFromText(geocodelocation)` when WKT present
+2. Else `ST_MakePoint(loc_long, loc_lat)`
+3. Else `NULL`
+
+**Indexes:** `issue_date`, `violation_code`, `make`, GiST on `geom`
+
+## PostGIS clean table (`citations_clean`)
+
+Slim analytics schema:
+
+| Column | Type | Notes |
+|--------|------|-------|
+| `ticket_number` | TEXT PK | |
+| `issue_datetime` | TIMESTAMPTZ NOT NULL | Combined date + parsed HHMM time |
+| `violation_code` | TEXT | Trimmed |
+| `violation_description` | TEXT | Trimmed |
+| `fine_amount` | DOUBLE PRECISION | |
+| `geom` | geometry(Point, 4326) | Copied from raw |
+
+**Datetime parsing:**
+
+```
+issue_date: 2025-04-26T00:00:00.000 (midnight from CSV)
+issue_time: "904" → pad → "0904" → 09:04 → 2025-04-26 09:04:00+TZ
+```
+
+Missing or non-numeric `issue_time` defaults to midnight on issue date.
+
+## SQLite schema
+
+### `citations`
+
+Same 23 columns, `ticket_number TEXT PRIMARY KEY`, `WITHOUT ROWID`.
+
+### `sync_log`
+
+| Column | Type | Purpose |
+|--------|------|---------|
+| `id` | INTEGER PK | Auto-increment |
+| `started_at` | TEXT | ISO UTC |
+| `finished_at` | TEXT | ISO UTC |
+| `source` | TEXT | `'csv'` or `'api'` |
+| `rows_inserted` | INTEGER | Rows processed |
+| `notes` | TEXT | e.g. `'caught_up'` |
+
+## API record shape (Socrata sync)
+
+Socrata returns GeoJSON for location; SQLite sync converts to WKT for `geocodelocation` consistency with CSV rows.
+
+Coerced fields: `agency`, `fine_amount`, `loc_lat`, `loc_long`.
+
+## MongoDB documents (legacy API)
+
+Expected shape (GeoJSON-centric, per OpenAPI):
+
+- `issue_date` — filterable datetime
+- `geometry` — GeoJSON Point or similar for `$geoWithin`
+
+Exact document schema is not fully documented in repo — inferred from `CitationService` query logic.
+
+## Auxiliary / boundary data
+
+| Data | Location | Used for |
+|------|----------|----------|
+| LA city boundary | `apps/web/src/data/los-angeles.json` | Map constraints |
+| LA county boundary | `apps/web/src/data/los-angeles-county.json` | Map bounds |
+| Neighborhood councils | Referenced in geocoder hook | Place search |
+| Mock citations / geocoder | `mock-*.json` | Development / testing |
+| Make / violation lookups | `data-science/references/` | Legacy normalization |
+| Zipcodes | Legacy upload scripts | Spatial joins |
+
+## Data quality quirks
+
+```mermaid
+flowchart TD
+ RAW[Raw citation row]
+ RAW --> Q1{Future issue_date?}
+ RAW --> Q2{Valid issue_time?}
+ RAW --> Q3{Has coordinates?}
+ RAW --> Q4{Duplicate ticket_number?}
+
+ Q1 -->|some rows| FILTER1[Filter in queries]
+ Q2 -->|missing / odd| DEFAULT[Default to midnight in clean]
+ Q3 -->|often no| NULLGEOM[geom IS NULL]
+ Q4 -->|re-load / sync| SKIP[INSERT OR IGNORE — no update]
+```
+
+| Quirk | Detail | Mitigation |
+|-------|--------|------------|
+| Future dates | Some `issue_date` values years ahead | Filter in analysis queries |
+| `issue_time` format | HHMM, no leading zeros; `"0"` → midnight | Handled in clean rebuild |
+| Missing geometry | Not all rows geocoded | `WHERE geom IS NOT NULL` for maps |
+| API vs CSV location | GeoJSON vs WKT | Normalized on API ingest |
+| No ticket updates | `INSERT OR IGNORE` | Switch to upsert if corrections needed |
+| Web app row cap | 50 rows per Socrata query | Pagination or DB backend |
+
+## Example queries
+
+**PostGIS — recent geocoded citations:**
+
+```sql
+SELECT ticket_number, issue_datetime, violation_description, ST_AsText(geom)
+FROM citations_clean
+WHERE geom IS NOT NULL
+ORDER BY issue_datetime DESC
+LIMIT 10;
+```
+
+**SQLite + Polars:**
+
+```python
+import sqlite3, polars as pl
+
+with sqlite3.connect("parking_citations.db") as conn:
+ df = pl.read_database("""
+ SELECT violation_description, COUNT(*) AS n,
+ ROUND(AVG(fine_amount), 2) AS avg_fine
+ FROM citations
+ WHERE violation_description IS NOT NULL
+ GROUP BY violation_description
+ ORDER BY n DESC
+ LIMIT 15
+ """, connection=conn)
+```
+
+## File artifacts (not in git)
+
+| Artifact | Typical path | Created by |
+|----------|--------------|------------|
+| Citation CSV | `Parking_Citations_*.csv` | Manual download |
+| SQLite DB | `parking_citations.db` | `parking_db.py` |
+| PostGIS volume | Docker `postgis_data` | `docker compose` |
+| Raw data dir | `raw_data/` (gitignored) | Various |
diff --git a/docs/08-roadmap-and-open-questions.md b/docs/08-roadmap-and-open-questions.md
new file mode 100644
index 00000000..3e72a25a
--- /dev/null
+++ b/docs/08-roadmap-and-open-questions.md
@@ -0,0 +1,205 @@
+# Roadmap and open questions
+
+Lucky Parking is an active Hack for LA project with **multiple partially overlapping systems**. This page catalogs what is unfinished, known bugs, and reasonable directions the project could take.
+
+**Status key:** ✅ Done · 🟡 Partial · 🔴 Not started / broken · 📋 Documented but missing
+
+## Integration map (today)
+
+```mermaid
+flowchart TB
+ subgraph done [Working today]
+ WEB[Web app map + filters]
+ SOC[Socrata live queries]
+ BETA_CSV[beta_pipeline CSV load]
+ BETA_SQL[beta_pipeline SQLite sync]
+ BETA_PG[beta_pipeline PostGIS + clean]
+ end
+
+ subgraph partial [Partial / limited]
+ WEB50[Web 50-row cap]
+ NB1[parking_db_explore notebook]
+ end
+
+ subgraph missing [Not connected / missing]
+ WEB_API[Web → Express API]
+ API_DB[API → MongoDB populated]
+ PG_SYNC[PostGIS API sync]
+ WEB_DB[Web → local DB]
+ VIO_NB[violation_analysis notebook]
+ ENV_EX[.env.example in beta_pipeline]
+ end
+
+ WEB --> SOC
+ WEB50 --> WEB
+ BETA_CSV --> BETA_PG
+ BETA_SQL --> BETA_CSV
+```
+
+## Unfinished by area
+
+### Web application
+
+| Item | Status | Location | Notes |
+|------|--------|----------|-------|
+| Socrata pagination | 🔴 | `apps/web/src/lib/socrata/parking-citations.ts:73` | Hard `pageSize: 50`; FIXME comment |
+| Map legend accuracy | 🔴 | `apps/web/src/components/data-legend.tsx` | TODO: use real violation categories |
+| Geocoder address/postcode handling | 🟡 | `apps/web/src/hooks/use-geocoder.tsx:65` | TODO for some result types |
+| Connect to backend API | 🔴 | architecture | Web bypasses `apps/api` entirely |
+| Connect to local PostGIS/SQLite | 🔴 | architecture | Pipeline output unused by web |
+| Automated tests | 🔴 | `apps/web` | No substantive test suite |
+
+### Backend API
+
+| Item | Status | Location | Notes |
+|------|--------|----------|-------|
+| Used by frontend | 🔴 | — | Express API is orphaned |
+| Env var naming bug | 🔴 | `citations.ts` vs `.env.schema` | Code uses `COL_CITATIONS`, schema says `COL_NAME_CITATIONS` |
+| GeoPolygon validation | 🟡 | `apps/api/src/utilities/schemas.ts` | TODO: proper Zod schema (`z.any()` today) |
+| Dataset ID alignment | 🔴 | OpenAPI external docs | May reference `wjz9-h9np` vs current `4f5p-udkv` |
+| Ingestion into MongoDB | 🔴 | — | No pipeline loads current dataset into API DB |
+| Automated tests | 🔴 | `apps/api` | Minimal coverage |
+
+### Beta data pipeline
+
+| Item | Status | Location | Notes |
+|------|--------|----------|-------|
+| SQLite CSV + API sync | ✅ | `parking_db.py` | Production-ready for local use |
+| PostGIS CSV + clean | ✅ | `parking_postgis.py`, `parking_clean.py` | Full pipeline works |
+| PostGIS API sync | 🔴 | ARCHITECTURE.md | Copy from SQLite, add geom handling |
+| Incremental clean rebuild | 🔴 | `parking_clean.py` | Full TRUNCATE each time |
+| Scheduled sync / load | 🔴 | — | No cron/CI automation |
+| `.env.example` file | 📋 | README references it | File not in repo |
+| `violation_analysis.ipynb` | 🔴 | stub imports only | Analysis not written |
+| Web/API consumption | 🔴 | — | Databases are offline analytics only |
+
+### Legacy data science
+
+| Item | Status | Notes |
+|------|--------|-------|
+| Makefile ETL | 🟡 | Works for old dataset/schema; not maintained for `4f5p-udkv` |
+| Sphinx docs | 🟡 | Some commands documented but not in Makefile |
+| S3 sync | 📋 | Referenced in docs, not implemented in Makefile |
+| Cookiecutter `src/models`, `src/features` | 🔴 | Never added; only `src/data/` exists |
+| Archived ML notebooks | 🟡 | Historical; not validated against current data |
+
+### Repository / ops
+
+| Item | Status | Notes |
+|------|--------|-------|
+| Root `pnpm test` | 🟡 | Task exists; apps lack meaningful tests |
+| Branch naming docs | 🟡 | CONTRIBUTING mentions `master`/`dev`; CI may use `main`/`stable` |
+| Monorepo Python tooling | 🔴 | Python lives outside pnpm; no unified root Python project |
+
+## Code-tracked TODOs / FIXMEs
+
+| File | Marker | Summary |
+|------|--------|---------|
+| `parking-citations.ts` | FIXME | Remove 50-row pagination limit |
+| `use-geocoder.tsx` | TODO | Handle address and postcode geocoder types |
+| `data-legend.tsx` | TODO | Refactor legend to real data |
+| `schemas.ts` (API) | TODO | Full GeoPolygon Zod schema |
+
+## Strategic forks (where the project could go)
+
+### Path A — "Live Socrata first" (minimal change)
+
+Improve the current web → Socrata path:
+
+- Implement client-side pagination in `fetchParkingCitations`
+- Add loading states and rate-limit handling
+- Fix legend and geocoder edge cases
+
+**Best if:** Team lacks infra for hosted DB; citations needed are small (recent + localized).
+
+**Risk:** Socrata rate limits and latency at scale; still no offline analysis parity with pipeline.
+
+### Path B — "PostGIS as source of truth"
+
+Make `beta_pipeline` + PostGIS the backend for everything:
+
+- Add PostGIS API sync
+- Expose citations via Next.js route handlers or revived Express API querying PostGIS
+- Point web app at internal API instead of Socrata
+- Deploy PostGIS (RDS, Supabase, etc.)
+
+**Best if:** Full-city or long date-range queries matter; team can run a database.
+
+**Risk:** Ops cost; need ingestion monitoring and auth for public API.
+
+### Path C — "Revive MongoDB API"
+
+Populate MongoDB from pipeline exports; web switches to existing OpenAPI contract:
+
+- ETL job: PostGIS/SQLite → GeoJSON documents → MongoDB
+- Fix env var bug and GeoPolygon schema
+- Update dataset to `4f5p-udkv`
+
+**Best if:** Team prefers document store + existing API spec.
+
+**Risk:** Two storage systems (PostGIS for analytics, Mongo for API) unless Mongo becomes sole store.
+
+### Path D — "Analytics focus"
+
+Prioritize data science deliverables over web integration:
+
+- Complete `violation_analysis.ipynb`
+- Build scheduled sync + clean jobs
+- Publish insights/reports; web remains demo with 50-row cap
+
+**Best if:** Primary stakeholders are researchers/policy analysts, not public map users.
+
+### Path E — "Consolidate and deprecate"
+
+Remove or archive legacy paths to reduce contributor confusion:
+
+- Mark `data-science/src/data/` and Mongo API as deprecated
+- Single Python package (`beta_pipeline`) + single dataset ID
+- Document one golden path in README
+
+**Best if:** Maintainer bandwidth is limited.
+
+## Suggested near-term priorities
+
+A pragmatic sequence many teams would follow:
+
+```mermaid
+gantt
+ title Possible near-term sequence
+ dateFormat YYYY-MM
+ section Quick wins
+ Add .env.example to beta_pipeline :a1, 2026-06, 1w
+ Fix API COL_CITATIONS env bug :a2, 2026-06, 1w
+ Document dataset ID in OpenAPI :a3, 2026-06, 1w
+ section Web UX
+ Socrata pagination OR raise limit :b1, 2026-07, 2w
+ Fix data legend :b2, 2026-07, 1w
+ section Data
+ violation_analysis notebook :c1, 2026-07, 3w
+ PostGIS API sync :c2, 2026-08, 3w
+ section Integration
+ Choose BFF strategy (B vs C) :d1, 2026-08, 2w
+ Wire web to internal API :d2, 2026-09, 4w
+```
+
+*Timeline is illustrative — not a committed project plan.*
+
+## Open questions for the team
+
+1. **Which dataset is canonical going forward?** Assume `4f5p-udkv` unless legacy comparisons require `wjz9-h9np`.
+2. **Should the Express API survive?** Or replace with Next.js server routes / PostgREST?
+3. **Is 50 rows acceptable temporarily?** Or is pagination/blocking for launch?
+4. **Who hosts PostGIS in production?** Docker locally only vs cloud managed service.
+5. **Are ticket corrections important?** If yes, move from `INSERT OR IGNORE` to upsert.
+6. **Should violation normalization (`references/`) feed into `citations_clean`?**
+7. **What is the public deployment target?** OpenAPI references `luckyparking.org` — document actual infra.
+
+## How to update this doc
+
+When closing a gap:
+
+1. Change status in the tables above
+2. Remove or resolve corresponding TODO/FIXME in code
+3. Link to PR or issue if tracked on GitHub
+
+When adding new scope, append to the strategic forks section with pros/cons so future contributors understand why a path was chosen or rejected.
diff --git a/docs/README.md b/docs/README.md
index 1738e47a..ba024a0c 100644
--- a/docs/README.md
+++ b/docs/README.md
@@ -103,6 +103,65 @@ Run these from the repository root.
| `pnpm verify` | Run type checks, linting, formatting, tests, and builds |
| `pnpm clean` | Remove generated workspace artifacts and dependencies |
+## Detailed documentation
+
+| Document | What you'll learn |
+|----------|-------------------|
+| [Overview](./01-overview.md) | Project mission, high-level architecture, and how the pieces relate |
+| [Monorepo structure](./02-monorepo-structure.md) | Turborepo layout, packages, tooling, and local dev workflow |
+| [Web application](./03-web-application.md) | Next.js map app, state, Socrata integration, components |
+| [Backend API](./04-backend-api.md) | Express/MongoDB API — current status and contract |
+| [Data pipeline (beta)](./05-data-pipeline.md) | Polars + SQLite/PostGIS ingestion, CLI, schemas |
+| [Legacy data science](./06-legacy-data-science.md) | Older ETL, notebooks, and normalized PostGIS schema |
+| [Data sources & schemas](./07-data-sources-and-schemas.md) | Dataset IDs, column dictionary, reference files |
+| [Roadmap & open questions](./08-roadmap-and-open-questions.md) | Unfinished work, known gaps, and possible directions |
+
+## Quick reference
+
+```mermaid
+flowchart TB
+ subgraph sources [External data]
+ CSV["Parking Citations CSV (~6 GB)"]
+ SOC["Socrata API (4f5p-udkv)"]
+ MAP["Mapbox Geocoding"]
+ end
+
+ subgraph monorepo [Lucky Parking monorepo]
+ WEB["apps/web — Next.js map"]
+ API["apps/api — Express + MongoDB"]
+ DS["data-science/beta_pipeline"]
+ LEG["data-science/ (legacy)"]
+ end
+
+ subgraph storage [Local / cloud storage]
+ SQLITE[(SQLite)]
+ PG[(PostGIS)]
+ MONGO[(MongoDB)]
+ end
+
+ CSV --> DS
+ SOC --> DS
+ SOC --> WEB
+ MAP --> WEB
+
+ DS --> SQLITE
+ DS --> PG
+ API --> MONGO
+
+ WEB -.->|"not connected today"| API
+ LEG -.->|"older dataset (wjz9-h9np)"| PG
+```
+
+## Related docs elsewhere in the repo
+
+- [`data-science/beta_pipeline/README.md`](../data-science/beta_pipeline/README.md) — quick start for the Python pipeline
+- [`data-science/beta_pipeline/ARCHITECTURE.md`](../data-science/beta_pipeline/ARCHITECTURE.md) — module-level code reference (complements [05-data-pipeline.md](./05-data-pipeline.md))
+- [`apps/api/src/docs/specs-v1.yaml`](../apps/api/src/docs/specs-v1.yaml) — OpenAPI spec for the citations API
+
+## Documentation status
+
+This documentation was written to reflect the repository as of mid-2026. Where behavior is uncertain or in flux, see [Roadmap & open questions](./08-roadmap-and-open-questions.md).
+
## Contributing
Contributions are welcome. Start with Hack for LA's [onboarding guide](https://www.hackforla.org/getting-started), then