Skip to content

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Spaceship Titanic Prediction Pipeline

This repository contains a machine learning pipeline for the Kaggle "Spaceship Titanic" binary classification competition.

Architecture & Design

The solution uses a heterogeneous ensemble learning architecture to maximize generalization:

  • Tree-Based Models (Dominant): CatBoost and LightGBM handle non-linear relationships and categorical features.
  • Distance-Based Models (Complementary): SVM and KNN provide alternative decision boundaries to reduce overfitting.
  • Ensemble Strategy: A weighted average blending approach combined with 5-Fold Stratified Cross-Validation.
  • Dynamic Thresholding: A post-processing step aligns the test set prediction distribution with the training set's true target distribution.

Project Structure

  • Data Preprocessing: Handles missing values via logic-based imputation and performs feature engineering (e.g., extracting Cabin details, binning ages, calculating family group sizes).
  • Feature Selection: Uses a baseline LightGBM model to evaluate feature importance, dropping redundant features automatically.
  • Model Training: Trains the four baseline models (CatBoost, LightGBM, SVM, KNN) using 5-Fold Stratified CV.
  • Post-Processing: Dynamically searches for the optimal probability threshold.
  • Automated Logging: Generates timestamped .txt log files for every run.

Environment & Dependencies

  • Python Version: Python 3.8+
  • Required Libraries:
    • pandas
    • numpy
    • scikit-learn
    • catboost
    • lightgbm
    • optuna

Install the dependencies using pip:

pip install pandas numpy scikit-learn catboost lightgbm optuna

How to Run

  1. Ensure train.csv and test.csv are located in the path specified in the load_data() function inside model.py.
  2. Open your terminal or command prompt.
  3. Navigate to the project directory.
  4. Execute the Python script:
    python model.py
  5. The script will output the final predictions to submission.csv and save a detailed log file (e.g., training_log_YYYYMMDD_HHMMSS.txt).

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages