#rs/notes #rs/class/csb320 - - - > [!NOTE]- Bullet Notes > - Classification models > - supervised > - classify instances into class (category) > - can predict class directly or probabilities > - Each column in data table is feature > - Each row is sample/tuple/instance > - K-Nearest neighbors > - lazy learner > - memorizes data, doesn't create model during training > - if k=6 it looks at 6 closes neighbors in dataset > - predicts based on what these are classified as > - there can be very different predictions based on k > - gets very expensive as the dataset grows > - selecting value of k > - k is a hyperparameter (user set value ) > - typical values of 3-15 > - lower value can be sensitive to noise > - higher value risks underfitting > - distance measures > - euclidean distance > - good when data is compact and continuous > - manhattan distance > - sum of absolute differences in coordinates > - good when data is discrete or with large distances > - Minkowski distance > - includes euclidean and manhattan > - parameter allows to interpolate between the two > - can be used for model tuning since distance function can be changed between euclidean and manhattan > - features should be standardized to ensure fair distance measures > - Logistic regression > - log-odds: natural log of the probability ratio > - uses a logistic regression to predict the chances of something being categorized in certain way > - logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$ > - output of logistic regression is compared to threshold T > - default T = 0.5 > - linear regression will underfit for classifying data, logistic regression is better > - multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$ > - Gaussian Naive Bayes > - normal distributions for each outcome > - Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$ > - $P(A|B$): Posterior probability > - Assumptions > - all input features are independent > - all input features contribute equally to classification > - requires that each probability is 0 > - if there is one option that is 0, can add 1 to each to ensure it works > - Linear Discriminant Analysis > - supervised > - dimensional reduction > - maximize the distance between groups > - within-class variance is minimized > - maximize the distances between the means of the two categories on the new axis > - trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$ > - Top of equation is the distance between the averages of the data projected onto the new line > - bottom is minimizing the scatter within each category > - discriminant analysis determines decision boundary between classes