01
Versioned, validated data
Race data is schema-checked before training and versioned with DVC, so any model can be traced back to the exact data it saw.
Personal projectCode and write-up
A machine learning pipeline that predicts Formula 1 lap times and retrains itself every race week. A new model replaces the current one only when it scores better.
Lap times shift between practice, qualifying and the race, so a model trained once goes stale within a weekend. A notebook can't keep up with that.
The requirement: retrain on each new race without a person in the loop, and never ship a model that is worse than the one already serving.
01
Race data is schema-checked before training and versioned with DVC, so any model can be traced back to the exact data it saw.
02
Airflow pulls the latest race, retrains XGBoost and LightGBM models tuned with Optuna, and logs each run to MLflow. A candidate is promoted only if it beats the current model.
03
A FastAPI service returns each prediction with an uncertainty estimate to a React dashboard. Prometheus, Grafana and Loki track accuracy and drift.
Airflow drives the loop from race data to the dashboard, and Grafana watches every step.
Automatic retraining is only safe if a worse model can never replace a better one. Every candidate is scored against the current model before promotion.
Trade-off: a sudden regime change can leave an older model serving until a new one wins.
A lap time without a range invites false precision. Each prediction comes with an interval so a reader knows how far to trust it.
Trade-off: more compute per request and a more complex response.
The dashboard is on Vercel, the API on Render and the database on Supabase, which keeps the whole system free to run.
Trade-off: cold starts on the API and three dashboards to watch.