Skip to content

Repository files navigation

AmsterdamUMCdb to CLIF ETL Pipeline

Python 3.12+ uv CLIF 2.1 AmsterdamUMCdb 1.0.2

This repository provides an ELT pipeline to transform the AmsterdamUMCdb v1.0.2 (legacy CSV format) into the Common Longitudinal ICU data Format (CLIF) v2.1.

AmsterdamUMCdb contains data from 23,106 ICU and high dependency unit admissions from 20,109 patients (2003-2016) at Amsterdam University Medical Centers.

Current status: 16 of the 20 tracked CLIF table pipelines are implemented (80%) and have been run against the full AmsterdamUMCdb dataset. The three microbiology tables and a standalone scores table are not yet implemented. See CLIF 2.1 Table Progress for coverage and limitations.

Table of Contents

Getting Started {#getting-started}

To run the pipeline, ensure you have the prerequisites below, then follow the Usage instructions.

For mapping decisions, see the per-table YAML files in mappings/ which document every source column, value map, and English translation.

Prerequisites {#prerequisites}

Before running the pipeline, ensure you have:

  • AmsterdamUMCdb access: Request access at amsterdammedicaldatascience.nl
  • AmsterdamUMCdb v1.0.2 CSV files (legacy format, not OMOP) downloaded to a local directory
  • Python 3.12+ and Git installed
  • Disk space:
    • ~80 GB for extracted AmsterdamUMCdb CSV files, plus space for downloaded archives if retained
    • ~1.1 GB for the current 16 CLIF output tables

Usage {#usage}

1. Clone and install

git clone https://github.com/<your-org>/CLIF-AmsterdamUMCdb.git
cd CLIF-AmsterdamUMCdb
uv sync

2. Place source data

Copy or symlink the AmsterdamUMCdb v1.0.2 CSV files into AmesDB/. Extract numericitems.zip first because the pipeline reads numericitems.csv directly:

AmesDB/
  admissions.csv
  drugitems.csv
  freetextitems.csv
  listitems.csv
  numericitems.csv
  procedureorderitems.csv
  processitems.csv

The patient_procedures transformer also requires the AMSTEL procedure mapping at ref/AMSTEL/data/mappings/procedureorderitems_item.usagi.csv.

3. Configure

Edit config.yaml to customize:

# Enable/disable specific CLIF tables (1=enabled, 0=disabled)
clif_tables:
  patient: 1
  hospitalization: 1
  adt: 1
  vitals: 1
  labs: 1
  # ...

# Paths
input_path: "./AmesDB"
output_path: "./output"

# Range handling strategy for binned columns (midpoint | lower_bound | upper_bound | null)
range_handling:
  age:
    strategy: "midpoint"

4. Run the pipeline

# Run all enabled tables
uv run main.py --config config.yaml

# Run specific tables only
uv run main.py --config config.yaml --tables patient hospitalization

# Validate config and mappings without running transforms
uv run main.py --config config.yaml --dry-run

Output {#output}

Generated Parquet files are written to output/:

Table Rows Size Description
clif_patient.parquet 20,109 ~72 KB Patient demographics and death time
clif_hospitalization.parquet 23,106 ~228 KB Admission/discharge records
clif_adt.parquet 42,298 ~436 KB ICU/MC locations and selected post-ICU destinations
clif_vitals.parquet 223,042,392 ~628 MB Weight, height, and continuous vital signs
clif_labs.parquet 10,248,660 ~83 MB Mapped numeric laboratory results
clif_hospital_diagnosis.parquet 185,854 ~360 KB APACHE/NICE diagnosis records
clif_patient_procedures.parquet 2,124,513 ~11 MB SNOMED-mapped procedure orders
clif_respiratory_support.parquet 16,464,183 ~258 MB Respiratory devices, modes, and parameters
clif_medication_admin_continuous.parquet 1,161,296 ~12 MB Curated continuous medications
clif_medication_admin_intermittent.parquet 663,161 ~5.9 MB Curated intermittent medications
clif_patient_assessments.parquet 1,415,874 ~7.4 MB GCS, RASS, VAS, CAM-ICU, and BPS assessments
clif_position.parquet 1,059,872 ~4.4 MB Prone/not-prone observations
clif_crrt_therapy.parquet 2,019,666 ~15 MB CVVH therapy parameters
clif_ecmo_mcs.parquet 749 ~12 KB ECMO and IABP observations
clif_code_status.parquet 15,762 ~88 KB Full code, DNAR, and AND status
clif_output.parquet 1,645,441 ~9.5 MB Urine output measurements

These counts and sizes are from a full local pipeline run and may change as mappings evolve. Each Parquet file has a companion clif_<table>_summary.csv with descriptive counts, value frequencies, and numeric ranges. Source labels are retained where the corresponding CLIF table provides a name field, but naming and provenance vary by table.

Validation status

The generated summaries are descriptive, not pass/fail validation. The repository does not yet include automated CLIF schema, category, referential-integrity, or clinical-range tests. Several implemented tables intentionally cover only source concepts currently mapped below.

Data Notes {#data-notes}

Timestamps

All AmsterdamUMCdb timestamps across all tables are global per patient — milliseconds since the patient's first ICU admission (admittedat=0). The epoch is anchored by the first admission's admissionyeargroup:

Year Group Epoch
2003-2009 2003-01-01
2010-2016 2010-01-01

Timestamps across different patients are not directly comparable. Negative timestamps are valid (pre-admission lab results).

Dutch Terminology

AmsterdamUMCdb uses Dutch medical terminology. All mappings include English translations in the YAML files. Key terms:

Dutch English
Man / Vrouw Male / Female
Overleden Deceased
Eerste Hulp Emergency Department
Verpleegafdeling Nursing ward
Cardiochirurgie Cardiac Surgery
Huis Home

Range Columns

Age, weight, and height are stored as binned ranges (e.g., "60-69", "80+", "59-"). The range_handling config controls extraction strategy (midpoint, lower bound, upper bound, or null).

Race/Ethnicity/Language

Not collected in AmsterdamUMCdb (Dutch single-center hospital). Set to "Unknown" in CLIF output.

Adding New Tables {#adding-new-tables}

  1. Create mappings/<table>.yaml with source columns, value maps, and _name maps
  2. Create clif_etl/transformers/<table>.py extending BaseTransformer
  3. Register the transformer in TRANSFORMER_REGISTRY in clif_etl/pipeline.py
  4. Enable the table in config.yaml under clif_tables

See existing mapping YAMLs and transformers for the pattern.

Resources {#resources}

CLIF 2.1 Table Progress

Implemented means that a transformer and mapping are registered and a full-dataset output has been generated. It does not imply complete source-domain coverage or formal CLIF validation.

# CLIF Table Status Amsterdam Source Notes
1 patient Implemented admissions.csv Sex and death time; race, ethnicity, and language are unavailable and set to Unknown
2 hospitalization Implemented admissions.csv Admission/discharge times, binned age, admission type, and discharge category
3 adt Implemented admissions.csv ICU/MC locations and selected post-ICU destinations; not a complete movement feed
4 vitals Implemented admissions.csv + numericitems.csv Weight/height ranges plus heart rate, blood pressure, SpO2, respiratory rate, and temperature; no clinical-range filtering
5 labs Implemented numericitems.csv 69 numeric lab mappings with unit conversions; order, collection, and result times all use measuredat
6 hospital_diagnosis Implemented listitems.csv APACHE/NICE custom codes rather than ICD codes
7 patient_procedures Implemented procedureorderitems.csv SNOMED mappings from AMSTEL; unmapped items are excluded
8 respiratory_support Implemented numericitems.csv + listitems.csv + processitems.csv Devices, modes, tracheostomy, and respiratory parameters; some devices and metrics are unavailable
9 medication_admin_continuous Implemented drugitems.csv Curated continuous medication subset; dose units are not normalized
10 medication_admin_intermittent Implemented drugitems.csv Curated intermittent medication subset; dose units are not normalized
11 patient_assessments Implemented listitems.csv GCS, RASS, VAS, CAM-ICU, and BPS; Braden is not mapped
12 position Implemented listitems.csv Prone versus not-prone observations only
13 crrt_therapy Implemented numericitems.csv CVVH parameters; device ID and dialysate flow are unavailable
14 ecmo_mcs Implemented numericitems.csv + listitems.csv + processitems.csv Limited ECMO metrics and IABP start observations
15 code_status Implemented listitems.csv Patient-level Full, DNAR, and AND categories
16 output Implemented numericitems.csv Urine output only; other drains and intake are not mapped
17 microbiology_culture Not started - Culture-based microbiology results
18 microbiology_nonculture Not started - Non-culture microbiology results
19 microbiology_susceptibility Not started - Antibiotic susceptibility testing
20 scores Not started - No standalone scores transformer; GCS and BPS totals are in patient_assessments

About

This repository provides the workflow and code to convert the AmsterdamUMCdb dataset into the Common Longitudinal ICU data Format (CLIF)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages