Files
ObsidianVault/Running Start/CSB320 - Machine Learning Concepts/Class 4-30 (Model Improvement & Dimensionality Reduction).md
2026-05-17 12:19:19 -07:00

97 lines
3.6 KiB
Markdown

#rs/notes #rs/class/csb320
- - -
- Data cleaning
- incomplete
- noisy (negative, wrong)
- inconsistent
- incorrect age vs. birthday
- intentional
- ex. not actually adding values, same for everything
- Missing data
- missing completely at random
- missing data completely at random
- missing at random
- missing based on variables but not missing values
- ex. no garage space data in specific location
- missing not at random
- missing due to unobserved factors or systematic reasons
- handling missing values
- constant value
- attribute mean
- most probable value
- can use regression models to predict missing features
- KNN imputation
- estimates based on k most similar instances
- handling noisy data
- smooth data by partitioning into bins
- smooth data by fitting into regression
- clustering to detect and remove outliers
- pairwise deletion: only delete missing/incorrect values
- listwise deletion: delete entire row if missing value
- some features need to be removed
- duplicates
- high proportion of missing values
- all same
- no pattern
- Feature transformation
- continuous -> discrete (simplification)
- reduce data noise
- decision trees can perform better with discrete data
- equal frequency or equal interval binning
- normalization
- scales data to range between 0 and 1
- standardization
- scales to have a mean of 0 and standard deviation of 1
- encoding categorical features
- binary encoding
- categories as binary vectors (00, 01, 10)
- label encoding
- unique integers for each category (1, 2, 3)
- one-hot encoding
- convert categories into binary vectors
- decision trees and random forests care about data split so they work better with label or binary encoding
- logistic or linear regression work better with one-hot encoding since they care about distance between data points
- imbalanced data
- evaluation metrics
- different ones work better with imbalanced data
- data-level methods
- undersampling
- reduce majority class to balance
- the minority class needs to be large enough to still have enough data
- random undersampling
- Tomek links
- pair from majority and minority class, are each others closest neighbors
- removing pairs can help to create clearer separation between classes
- can make decision boundary more clear
- oversampling
- increase minority class through SMOTE
- ![[SMOTE.excalidraw]]
- synthetic minority oversampling technique
- increases size of minority and variety
- find k nearest minority neighbors, select j of them
- add new synthetic data along line between these data points
- can overgeneralize data
- algorithm-level methods
- can weight classes differently
- assigns misclassification costs
- algorithm selection (random forest and adaboost handle imbalanced data better)
- threshold adjustment
- dimensionality
- too many features creates sparse data
- can cause overfitting
- distance becomes less meaningful
- reduction
- simplifies representation, improves performance
- can preserve essential information
- feature selection
- ranked with statistical tests
- iterative model building and testing
- linear -> straight line
- monotonic -> moving in same direction
- want a linear/monotonic relationship between target class and each feature
- do not want linear/monotonic relationship between feature classes
- different metrics depending on if inputs/outputs are numerical or categorical
- PCA (principle component analysis)
- creates new uncorrelated features (principle components)
- maintains the maximum amount of variance in the data
-