97 lines
3.6 KiB
Markdown
97 lines
3.6 KiB
Markdown
#rs/notes #rs/class/csb320
|
|
- - -
|
|
- Data cleaning
|
|
- incomplete
|
|
- noisy (negative, wrong)
|
|
- inconsistent
|
|
- incorrect age vs. birthday
|
|
- intentional
|
|
- ex. not actually adding values, same for everything
|
|
- Missing data
|
|
- missing completely at random
|
|
- missing data completely at random
|
|
- missing at random
|
|
- missing based on variables but not missing values
|
|
- ex. no garage space data in specific location
|
|
- missing not at random
|
|
- missing due to unobserved factors or systematic reasons
|
|
- handling missing values
|
|
- constant value
|
|
- attribute mean
|
|
- most probable value
|
|
- can use regression models to predict missing features
|
|
- KNN imputation
|
|
- estimates based on k most similar instances
|
|
- handling noisy data
|
|
- smooth data by partitioning into bins
|
|
- smooth data by fitting into regression
|
|
- clustering to detect and remove outliers
|
|
- pairwise deletion: only delete missing/incorrect values
|
|
- listwise deletion: delete entire row if missing value
|
|
- some features need to be removed
|
|
- duplicates
|
|
- high proportion of missing values
|
|
- all same
|
|
- no pattern
|
|
- Feature transformation
|
|
- continuous -> discrete (simplification)
|
|
- reduce data noise
|
|
- decision trees can perform better with discrete data
|
|
- equal frequency or equal interval binning
|
|
- normalization
|
|
- scales data to range between 0 and 1
|
|
- standardization
|
|
- scales to have a mean of 0 and standard deviation of 1
|
|
- encoding categorical features
|
|
- binary encoding
|
|
- categories as binary vectors (00, 01, 10)
|
|
- label encoding
|
|
- unique integers for each category (1, 2, 3)
|
|
- one-hot encoding
|
|
- convert categories into binary vectors
|
|
- decision trees and random forests care about data split so they work better with label or binary encoding
|
|
- logistic or linear regression work better with one-hot encoding since they care about distance between data points
|
|
- imbalanced data
|
|
- evaluation metrics
|
|
- different ones work better with imbalanced data
|
|
- data-level methods
|
|
- undersampling
|
|
- reduce majority class to balance
|
|
- the minority class needs to be large enough to still have enough data
|
|
- random undersampling
|
|
- Tomek links
|
|
- pair from majority and minority class, are each others closest neighbors
|
|
- removing pairs can help to create clearer separation between classes
|
|
- can make decision boundary more clear
|
|
- oversampling
|
|
- increase minority class through SMOTE
|
|
- ![[SMOTE.excalidraw]]
|
|
- synthetic minority oversampling technique
|
|
- increases size of minority and variety
|
|
- find k nearest minority neighbors, select j of them
|
|
- add new synthetic data along line between these data points
|
|
- can overgeneralize data
|
|
- algorithm-level methods
|
|
- can weight classes differently
|
|
- assigns misclassification costs
|
|
- algorithm selection (random forest and adaboost handle imbalanced data better)
|
|
- threshold adjustment
|
|
- dimensionality
|
|
- too many features creates sparse data
|
|
- can cause overfitting
|
|
- distance becomes less meaningful
|
|
- reduction
|
|
- simplifies representation, improves performance
|
|
- can preserve essential information
|
|
- feature selection
|
|
- ranked with statistical tests
|
|
- iterative model building and testing
|
|
- linear -> straight line
|
|
- monotonic -> moving in same direction
|
|
- want a linear/monotonic relationship between target class and each feature
|
|
- do not want linear/monotonic relationship between feature classes
|
|
- different metrics depending on if inputs/outputs are numerical or categorical
|
|
- PCA (principle component analysis)
|
|
- creates new uncorrelated features (principle components)
|
|
- maintains the maximum amount of variance in the data
|
|
- |