Skip to content
All projects

Machine learning · 2023

Kidney Stone Prediction

Kaggle Playground Series S3E12 — binary classification

Kaggle competition page for binary classification with a tabular kidney stone prediction dataset

End-to-end ML pipeline predicting kidney stones from urine physiology, using feature engineering and a CatBoost ensemble validated with stratified k-fold.

0.756
private ROC AUC
CatBoost · XGB
models compared

#Project overview

I used a Python-based Kaggle kernel to run a full data analysis and predictive modelling workflow for a machine learning competition. The goal was to predict a binary target from a set of physiological features using an ensemble algorithm, CatBoost.

#Tools and libraries

  • Python for data manipulation and modelling
  • Pandas & NumPy for data processing and linear algebra
  • scikit-learn for scaling and model evaluation
  • Seaborn & Matplotlib for visual insight into the data
  • CatBoost, an ensemble algorithm that is robust on categorical data

#Data handling

  • Loading: imported the train and test sets within the Kaggle environment and explored their size and shape.
  • Exploration: studied distributions and relationships using correlation matrices and plots.
  • Visualisation: heatmaps to explore correlations between features, and scatter plots to compare key feature pairs across conditions.

#Feature engineering

I created ratios and products of existing features to uncover interaction effects, giving the model more nuanced information than the raw features alone.

#Modelling

  • Preprocessing: normalised the feature set so every feature carried equal weight.
  • Training: CatBoostRegressor with stratified k-fold cross-validation to make predictions more robust and reduce overfitting.
  • Evaluation: ROC AUC, a single metric that captures how well the model separates the classes.

#Results

Training was refined iteratively to optimise ROC AUC across the validation folds, and I compared CatBoost against an XGBoost classifier. Final predictions on the test set were formatted into a competition submission, scoring 0.756 ROC AUC on the private leaderboard.

#What it shows

  • Managing and preprocessing data efficiently with industry-standard Python libraries
  • Using ensemble learning for robust predictive modelling
  • Analysing and visualising data to uncover underlying patterns
  • Optimising and evaluating models with cross-validation and the right metrics