This repository provides an ELT pipeline to transform the AmsterdamUMCdb v1.0.2 (legacy CSV format) into the Common Longitudinal ICU data Format (CLIF) v2.1.
AmsterdamUMCdb contains data from 23,106 ICU and high dependency unit admissions from 20,109 patients (2003-2016) at Amsterdam University Medical Centers.
Current status: 16 of the 20 tracked CLIF table pipelines are implemented (80%) and have been run against the full AmsterdamUMCdb dataset. The three microbiology tables and a standalone scores table are not yet implemented. See CLIF 2.1 Table Progress for coverage and limitations.
- Getting Started
- Prerequisites
- Usage
- Output
- Data Notes
- Adding New Tables
- Resources
- CLIF 2.1 Table Progress
To run the pipeline, ensure you have the prerequisites below, then follow the Usage instructions.
For mapping decisions, see the per-table YAML files in mappings/ which document every source column, value map, and English translation.
Before running the pipeline, ensure you have:
- AmsterdamUMCdb access: Request access at amsterdammedicaldatascience.nl
- AmsterdamUMCdb v1.0.2 CSV files (legacy format, not OMOP) downloaded to a local directory
- Python 3.12+ and Git installed
- Disk space:
- ~80 GB for extracted AmsterdamUMCdb CSV files, plus space for downloaded archives if retained
- ~1.1 GB for the current 16 CLIF output tables
git clone https://github.com/<your-org>/CLIF-AmsterdamUMCdb.git
cd CLIF-AmsterdamUMCdb
uv syncCopy or symlink the AmsterdamUMCdb v1.0.2 CSV files into AmesDB/. Extract numericitems.zip first because the pipeline reads numericitems.csv directly:
AmesDB/
admissions.csv
drugitems.csv
freetextitems.csv
listitems.csv
numericitems.csv
procedureorderitems.csv
processitems.csv
The patient_procedures transformer also requires the AMSTEL procedure mapping at ref/AMSTEL/data/mappings/procedureorderitems_item.usagi.csv.
Edit config.yaml to customize:
# Enable/disable specific CLIF tables (1=enabled, 0=disabled)
clif_tables:
patient: 1
hospitalization: 1
adt: 1
vitals: 1
labs: 1
# ...
# Paths
input_path: "./AmesDB"
output_path: "./output"
# Range handling strategy for binned columns (midpoint | lower_bound | upper_bound | null)
range_handling:
age:
strategy: "midpoint"# Run all enabled tables
uv run main.py --config config.yaml
# Run specific tables only
uv run main.py --config config.yaml --tables patient hospitalization
# Validate config and mappings without running transforms
uv run main.py --config config.yaml --dry-runGenerated Parquet files are written to output/:
| Table | Rows | Size | Description |
|---|---|---|---|
clif_patient.parquet |
20,109 | ~72 KB | Patient demographics and death time |
clif_hospitalization.parquet |
23,106 | ~228 KB | Admission/discharge records |
clif_adt.parquet |
42,298 | ~436 KB | ICU/MC locations and selected post-ICU destinations |
clif_vitals.parquet |
223,042,392 | ~628 MB | Weight, height, and continuous vital signs |
clif_labs.parquet |
10,248,660 | ~83 MB | Mapped numeric laboratory results |
clif_hospital_diagnosis.parquet |
185,854 | ~360 KB | APACHE/NICE diagnosis records |
clif_patient_procedures.parquet |
2,124,513 | ~11 MB | SNOMED-mapped procedure orders |
clif_respiratory_support.parquet |
16,464,183 | ~258 MB | Respiratory devices, modes, and parameters |
clif_medication_admin_continuous.parquet |
1,161,296 | ~12 MB | Curated continuous medications |
clif_medication_admin_intermittent.parquet |
663,161 | ~5.9 MB | Curated intermittent medications |
clif_patient_assessments.parquet |
1,415,874 | ~7.4 MB | GCS, RASS, VAS, CAM-ICU, and BPS assessments |
clif_position.parquet |
1,059,872 | ~4.4 MB | Prone/not-prone observations |
clif_crrt_therapy.parquet |
2,019,666 | ~15 MB | CVVH therapy parameters |
clif_ecmo_mcs.parquet |
749 | ~12 KB | ECMO and IABP observations |
clif_code_status.parquet |
15,762 | ~88 KB | Full code, DNAR, and AND status |
clif_output.parquet |
1,645,441 | ~9.5 MB | Urine output measurements |
These counts and sizes are from a full local pipeline run and may change as mappings evolve. Each Parquet file has a companion clif_<table>_summary.csv with descriptive counts, value frequencies, and numeric ranges. Source labels are retained where the corresponding CLIF table provides a name field, but naming and provenance vary by table.
The generated summaries are descriptive, not pass/fail validation. The repository does not yet include automated CLIF schema, category, referential-integrity, or clinical-range tests. Several implemented tables intentionally cover only source concepts currently mapped below.
All AmsterdamUMCdb timestamps across all tables are global per patient — milliseconds since the patient's first ICU admission (admittedat=0). The epoch is anchored by the first admission's admissionyeargroup:
| Year Group | Epoch |
|---|---|
2003-2009 |
2003-01-01 |
2010-2016 |
2010-01-01 |
Timestamps across different patients are not directly comparable. Negative timestamps are valid (pre-admission lab results).
AmsterdamUMCdb uses Dutch medical terminology. All mappings include English translations in the YAML files. Key terms:
| Dutch | English |
|---|---|
| Man / Vrouw | Male / Female |
| Overleden | Deceased |
| Eerste Hulp | Emergency Department |
| Verpleegafdeling | Nursing ward |
| Cardiochirurgie | Cardiac Surgery |
| Huis | Home |
Age, weight, and height are stored as binned ranges (e.g., "60-69", "80+", "59-"). The range_handling config controls extraction strategy (midpoint, lower bound, upper bound, or null).
Not collected in AmsterdamUMCdb (Dutch single-center hospital). Set to "Unknown" in CLIF output.
- Create
mappings/<table>.yamlwith source columns, value maps, and_namemaps - Create
clif_etl/transformers/<table>.pyextendingBaseTransformer - Register the transformer in
TRANSFORMER_REGISTRYinclif_etl/pipeline.py - Enable the table in
config.yamlunderclif_tables
See existing mapping YAMLs and transformers for the pattern.
- CLIF Website & Data Dictionary
- CLIF mCIDE Categories
- AmsterdamUMCdb Documentation
- AmsterdamUMCdb GitHub
- AMSTEL (OMOP ETL)
Implemented means that a transformer and mapping are registered and a full-dataset output has been generated. It does not imply complete source-domain coverage or formal CLIF validation.
| # | CLIF Table | Status | Amsterdam Source | Notes |
|---|---|---|---|---|
| 1 | patient |
Implemented | admissions.csv | Sex and death time; race, ethnicity, and language are unavailable and set to Unknown |
| 2 | hospitalization |
Implemented | admissions.csv | Admission/discharge times, binned age, admission type, and discharge category |
| 3 | adt |
Implemented | admissions.csv | ICU/MC locations and selected post-ICU destinations; not a complete movement feed |
| 4 | vitals |
Implemented | admissions.csv + numericitems.csv | Weight/height ranges plus heart rate, blood pressure, SpO2, respiratory rate, and temperature; no clinical-range filtering |
| 5 | labs |
Implemented | numericitems.csv | 69 numeric lab mappings with unit conversions; order, collection, and result times all use measuredat |
| 6 | hospital_diagnosis |
Implemented | listitems.csv | APACHE/NICE custom codes rather than ICD codes |
| 7 | patient_procedures |
Implemented | procedureorderitems.csv | SNOMED mappings from AMSTEL; unmapped items are excluded |
| 8 | respiratory_support |
Implemented | numericitems.csv + listitems.csv + processitems.csv | Devices, modes, tracheostomy, and respiratory parameters; some devices and metrics are unavailable |
| 9 | medication_admin_continuous |
Implemented | drugitems.csv | Curated continuous medication subset; dose units are not normalized |
| 10 | medication_admin_intermittent |
Implemented | drugitems.csv | Curated intermittent medication subset; dose units are not normalized |
| 11 | patient_assessments |
Implemented | listitems.csv | GCS, RASS, VAS, CAM-ICU, and BPS; Braden is not mapped |
| 12 | position |
Implemented | listitems.csv | Prone versus not-prone observations only |
| 13 | crrt_therapy |
Implemented | numericitems.csv | CVVH parameters; device ID and dialysate flow are unavailable |
| 14 | ecmo_mcs |
Implemented | numericitems.csv + listitems.csv + processitems.csv | Limited ECMO metrics and IABP start observations |
| 15 | code_status |
Implemented | listitems.csv | Patient-level Full, DNAR, and AND categories |
| 16 | output |
Implemented | numericitems.csv | Urine output only; other drains and intake are not mapped |
| 17 | microbiology_culture |
Not started | - | Culture-based microbiology results |
| 18 | microbiology_nonculture |
Not started | - | Non-culture microbiology results |
| 19 | microbiology_susceptibility |
Not started | - | Antibiotic susceptibility testing |
| 20 | scores |
Not started | - | No standalone scores transformer; GCS and BPS totals are in patient_assessments |