This repository contains a machine learning pipeline for the Kaggle "Spaceship Titanic" binary classification competition.
The solution uses a heterogeneous ensemble learning architecture to maximize generalization:
- Tree-Based Models (Dominant): CatBoost and LightGBM handle non-linear relationships and categorical features.
- Distance-Based Models (Complementary): SVM and KNN provide alternative decision boundaries to reduce overfitting.
- Ensemble Strategy: A weighted average blending approach combined with 5-Fold Stratified Cross-Validation.
- Dynamic Thresholding: A post-processing step aligns the test set prediction distribution with the training set's true target distribution.
- Data Preprocessing: Handles missing values via logic-based imputation and performs feature engineering (e.g., extracting Cabin details, binning ages, calculating family group sizes).
- Feature Selection: Uses a baseline LightGBM model to evaluate feature importance, dropping redundant features automatically.
- Model Training: Trains the four baseline models (CatBoost, LightGBM, SVM, KNN) using 5-Fold Stratified CV.
- Post-Processing: Dynamically searches for the optimal probability threshold.
- Automated Logging: Generates timestamped
.txtlog files for every run.
- Python Version: Python 3.8+
- Required Libraries:
pandasnumpyscikit-learncatboostlightgbmoptuna
Install the dependencies using pip:
pip install pandas numpy scikit-learn catboost lightgbm optuna- Ensure
train.csvandtest.csvare located in the path specified in theload_data()function insidemodel.py. - Open your terminal or command prompt.
- Navigate to the project directory.
- Execute the Python script:
python model.py
- The script will output the final predictions to
submission.csvand save a detailed log file (e.g.,training_log_YYYYMMDD_HHMMSS.txt).