Files
2026-05-22 12:17:26 -07:00

2.6 KiB

#rs/notes #rs/class/csb320


  • Classification models
    • supervised
    • classify instances into class (category)
    • can predict class directly or probabilities
  • Each column in data table is feature
    • Each row is sample/tuple/instance
  • K-Nearest neighbors
    • lazy learner
      • memorizes data, doesn't create model during training
    • if k=6 it looks at 6 closes neighbors in dataset
      • predicts based on what these are classified as
      • there can be very different predictions based on k
    • gets very expensive as the dataset grows
    • selecting value of k
      • k is a hyperparameter (user set value )
      • typical values of 3-15
      • lower value can be sensitive to noise
      • higher value risks underfitting
    • distance measures
      • euclidean distance
        • good when data is compact and continuous
      • manhattan distance
        • sum of absolute differences in coordinates
        • good when data is discrete or with large distances
      • Minkowski distance
        • includes euclidean and manhattan
        • parameter allows to interpolate between the two
        • can be used for model tuning since distance function can be changed between euclidean and manhattan
      • features should be standardized to ensure fair distance measures
  • Logistic regression
    • log-odds: natural log of the probability ratio
    • uses a logistic regression to predict the chances of something being categorized in certain way
    • logistic regression \hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}
    • output of logistic regression is compared to threshold T
    • default T = 0.5
    • linear regression will underfit for classifying data, logistic regression is better
    • multiple input features: \hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}
  • Gaussian Naive Bayes
    • normal distributions for each outcome
    • Baye's rule: P(A|B) = \frac{P(B|A) * P(A)}{P(B)}
      • P(A|B): Posterior probability
      • Assumptions
        • all input features are independent
        • all input features contribute equally to classification
    • requires that each probability is 0
      • if there is one option that is 0, can add 1 to each to ensure it works
  • Linear Discriminant Analysis
    • supervised
    • dimensional reduction
    • maximize the distance between groups
      • within-class variance is minimized
      • maximize the distances between the means of the two categories on the new axis
    • trying to maximize: \frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}
      • Top of equation is the distance between the averages of the data projected onto the new line
      • bottom is minimizing the scatter within each category
    • discriminant analysis determines decision boundary between classes