Files
ObsidianVault/Running Start/CSB320 - Machine Learning Concepts/Class 4-30 (Model Improvement & Dimensionality Reduction).md
T
2026-05-17 12:19:19 -07:00

3.6 KiB

#rs/notes #rs/class/csb320


  • Data cleaning
    • incomplete
    • noisy (negative, wrong)
    • inconsistent
      • incorrect age vs. birthday
    • intentional
      • ex. not actually adding values, same for everything
  • Missing data
    • missing completely at random
      • missing data completely at random
    • missing at random
      • missing based on variables but not missing values
      • ex. no garage space data in specific location
    • missing not at random
      • missing due to unobserved factors or systematic reasons
    • handling missing values
      • constant value
      • attribute mean
      • most probable value
      • can use regression models to predict missing features
      • KNN imputation
        • estimates based on k most similar instances
  • handling noisy data
    • smooth data by partitioning into bins
    • smooth data by fitting into regression
    • clustering to detect and remove outliers
  • pairwise deletion: only delete missing/incorrect values
  • listwise deletion: delete entire row if missing value
  • some features need to be removed
    • duplicates
    • high proportion of missing values
    • all same
    • no pattern
  • Feature transformation
    • continuous -> discrete (simplification)
      • reduce data noise
      • decision trees can perform better with discrete data
      • equal frequency or equal interval binning
    • normalization
      • scales data to range between 0 and 1
    • standardization
      • scales to have a mean of 0 and standard deviation of 1
    • encoding categorical features
      • binary encoding
        • categories as binary vectors (00, 01, 10)
      • label encoding
        • unique integers for each category (1, 2, 3)
      • one-hot encoding
        • convert categories into binary vectors
      • decision trees and random forests care about data split so they work better with label or binary encoding
      • logistic or linear regression work better with one-hot encoding since they care about distance between data points
  • imbalanced data
    • evaluation metrics
      • different ones work better with imbalanced data
    • data-level methods
      • undersampling
        • reduce majority class to balance
        • the minority class needs to be large enough to still have enough data
        • random undersampling
        • Tomek links
          • pair from majority and minority class, are each others closest neighbors
          • removing pairs can help to create clearer separation between classes
          • can make decision boundary more clear
      • oversampling
        • increase minority class through SMOTE
          • !SMOTE.excalidraw
          • synthetic minority oversampling technique
          • increases size of minority and variety
          • find k nearest minority neighbors, select j of them
          • add new synthetic data along line between these data points
          • can overgeneralize data
    • algorithm-level methods
      • can weight classes differently
      • assigns misclassification costs
      • algorithm selection (random forest and adaboost handle imbalanced data better)
      • threshold adjustment
  • dimensionality
    • too many features creates sparse data
      • can cause overfitting
      • distance becomes less meaningful
    • reduction
      • simplifies representation, improves performance
      • can preserve essential information
    • feature selection
      • ranked with statistical tests
      • iterative model building and testing
    • linear -> straight line
    • monotonic -> moving in same direction
    • want a linear/monotonic relationship between target class and each feature
    • do not want linear/monotonic relationship between feature classes
    • different metrics depending on if inputs/outputs are numerical or categorical
    • PCA (principle component analysis)
      • creates new uncorrelated features (principle components)
      • maintains the maximum amount of variance in the data