Files
2026-05-22 12:17:26 -07:00

60 lines
2.6 KiB
Markdown

#rs/notes #rs/class/csb320
- - -
- Classification models
- supervised
- classify instances into class (category)
- can predict class directly or probabilities
- Each column in data table is feature
- Each row is sample/tuple/instance
- K-Nearest neighbors
- lazy learner
- memorizes data, doesn't create model during training
- if k=6 it looks at 6 closes neighbors in dataset
- predicts based on what these are classified as
- there can be very different predictions based on k
- gets very expensive as the dataset grows
- selecting value of k
- k is a hyperparameter (user set value )
- typical values of 3-15
- lower value can be sensitive to noise
- higher value risks underfitting
- distance measures
- euclidean distance
- good when data is compact and continuous
- manhattan distance
- sum of absolute differences in coordinates
- good when data is discrete or with large distances
- Minkowski distance
- includes euclidean and manhattan
- parameter allows to interpolate between the two
- can be used for model tuning since distance function can be changed between euclidean and manhattan
- features should be standardized to ensure fair distance measures
- Logistic regression
- log-odds: natural log of the probability ratio
- uses a logistic regression to predict the chances of something being categorized in certain way
- logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
- output of logistic regression is compared to threshold T
- default T = 0.5
- linear regression will underfit for classifying data, logistic regression is better
- multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
- Gaussian Naive Bayes
- normal distributions for each outcome
- Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
- $P(A|B$): Posterior probability
- Assumptions
- all input features are independent
- all input features contribute equally to classification
- requires that each probability is 0
- if there is one option that is 0, can add 1 to each to ensure it works
- Linear Discriminant Analysis
- supervised
- dimensional reduction
- maximize the distance between groups
- within-class variance is minimized
- maximize the distances between the means of the two categories on the new axis
- trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$
- Top of equation is the distance between the averages of the data projected onto the new line
- bottom is minimizing the scatter within each category
- discriminant analysis determines decision boundary between classes