Files
ObsidianVault/Running Start/AD450 - Data Science Development/Discussion - Sharing Experiences on Data Cleaning.md
T
2026-05-17 12:19:19 -07:00

3.3 KiB

#rs/class/ad450 #rs/discussion


Description

For another one of my classes I recently completed a project where I had to analyze a dataset including cleaning the data before using it with machine learning models to predict one of the features. The dataset that I chose to work with was a collection of data from individual laps of Formula 1 races over the last 4 years. This was a very large dataset with over 100k instances and there were several things that I had do do to preprocess and clean the data.

Challenges

  • Missing Values: One of the features in the dataset had some instances with missing values. There were 66 of the rows of the feature describing the tire compound used by the team that were missing values. Without doing something to remove the missing data later steps in the process would fail since they don't have ways to deal with null values.
  • Outliers: A few of the features in the dataset (especially the LapTime feature) had some extreme outliers. Most of the values of the LapTime feature were between 0 and 200 seconds while there were a few that were over 2000 seconds. This could cause models later on to be biased and metrics such as average could also not properly represent the data.
  • Redundant Features: There were some features that were redundant and contained effectively the same data in the dataset. One of these was RaceProgress and LapNumber which are highly correlated and contain the same information in slightly different forms. This introduces duplicate data that could harm model performance later.

Strategies Used

  • Missing Values: Since there were only 66 missing values for the Compound feature I decided to drop the rows with the missing values. With over 100k rows 66 is a very small percentage of the data and dropping the rows is easier than imputing the data. This was effective since I didn't lose much of the data and didn't introduce the possible inaccuracy of imputation.
  • Outliers: To remove the outliers from the dataset I created a function that looked for values outside of a +/- 3 standard deviation range. This range contains 99.7% of the data so it will only remove values that are very extreme and biasing the data. This removed the outliers and made the data much more consistent. This was mostly effective but one of the features that I used the function on removed almost 1k instances. I decided to keep this but that probably means that the data was more spread rather than having outliers.
  • Redundant Features: For the redundant features I decided to remove one of them to get rid of the duplicate data. For RaceProgress and LapNumber I removed the LapNumber feature since RaceProgress was more granular and was there for likely more accurate. This was pretty effective although I could have performed some more analysis to better decide which feature was better to keep.

Reflections

Overall I learned a lot about data cleaning with this project and this was one of the first times that I went through the full data pipeline. I especially learned about different considerations you need to keep in mind when preprocessing data for machine learning pipelines such as looking for outliers and removing or imputing missing data. In the future I don't think I would approach data cleaning very differently but I do think I have a better understanding and will likely be able to go more in depth in the future.