#rs/notes #rs/class/csb320 - - - - Data cleaning - incomplete - noisy (negative, wrong) - inconsistent - incorrect age vs. birthday - intentional - ex. not actually adding values, same for everything - Missing data - missing completely at random - missing data completely at random - missing at random - missing based on variables but not missing values - ex. no garage space data in specific location - missing not at random - missing due to unobserved factors or systematic reasons - handling missing values - constant value - attribute mean - most probable value - can use regression models to predict missing features - KNN imputation - estimates based on k most similar instances - handling noisy data - smooth data by partitioning into bins - smooth data by fitting into regression - clustering to detect and remove outliers - pairwise deletion: only delete missing/incorrect values - listwise deletion: delete entire row if missing value - some features need to be removed - duplicates - high proportion of missing values - all same - no pattern - Feature transformation - continuous -> discrete (simplification) - reduce data noise - decision trees can perform better with discrete data - equal frequency or equal interval binning - normalization - scales data to range between 0 and 1 - standardization - scales to have a mean of 0 and standard deviation of 1 - encoding categorical features - binary encoding - categories as binary vectors (00, 01, 10) - label encoding - unique integers for each category (1, 2, 3) - one-hot encoding - convert categories into binary vectors - decision trees and random forests care about data split so they work better with label or binary encoding - logistic or linear regression work better with one-hot encoding since they care about distance between data points - imbalanced data - evaluation metrics - different ones work better with imbalanced data - data-level methods - undersampling - reduce majority class to balance - the minority class needs to be large enough to still have enough data - random undersampling - Tomek links - pair from majority and minority class, are each others closest neighbors - removing pairs can help to create clearer separation between classes - can make decision boundary more clear - oversampling - increase minority class through SMOTE - ![[SMOTE.excalidraw]] - synthetic minority oversampling technique - increases size of minority and variety - find k nearest minority neighbors, select j of them - add new synthetic data along line between these data points - can overgeneralize data - algorithm-level methods - can weight classes differently - assigns misclassification costs - algorithm selection (random forest and adaboost handle imbalanced data better) - threshold adjustment - dimensionality - too many features creates sparse data - can cause overfitting - distance becomes less meaningful - reduction - simplifies representation, improves performance - can preserve essential information - feature selection - ranked with statistical tests - iterative model building and testing - linear -> straight line - monotonic -> moving in same direction - want a linear/monotonic relationship between target class and each feature - do not want linear/monotonic relationship between feature classes - different metrics depending on if inputs/outputs are numerical or categorical - PCA (principle component analysis) - creates new uncorrelated features (principle components) - maintains the maximum amount of variance in the data -