vault backup: 2026-05-22 12:17:26

This commit is contained in:
ben committed 2026-05-22 12:17:26 -07:00
1 parent 0972ee377a
commit d7103664a0
3 files changed
+142 -143

No files matched your search

@@ -1,60 +1,59 @@
#rs/notes #rs/class/csb320 #rs/notes #rs/class/csb320
- - - - - -
> [!NOTE]- Bullet Notes - Classification models
> - Classification models - supervised
> - supervised - classify instances into class (category)
> - classify instances into class (category) - can predict class directly or probabilities
> - can predict class directly or probabilities - Each column in data table is feature
> - Each column in data table is feature - Each row is sample/tuple/instance
> - Each row is sample/tuple/instance - K-Nearest neighbors
> - K-Nearest neighbors - lazy learner
> - lazy learner - memorizes data, doesn't create model during training
> - memorizes data, doesn't create model during training - if k=6 it looks at 6 closes neighbors in dataset
> - if k=6 it looks at 6 closes neighbors in dataset - predicts based on what these are classified as
> - predicts based on what these are classified as - there can be very different predictions based on k
> - there can be very different predictions based on k - gets very expensive as the dataset grows
> - gets very expensive as the dataset grows - selecting value of k
> - selecting value of k - k is a hyperparameter (user set value )
> - k is a hyperparameter (user set value ) - typical values of 3-15
> - typical values of 3-15 - lower value can be sensitive to noise
> - lower value can be sensitive to noise - higher value risks underfitting
> - higher value risks underfitting - distance measures
> - distance measures - euclidean distance
> - euclidean distance - good when data is compact and continuous
> - good when data is compact and continuous - manhattan distance
> - manhattan distance - sum of absolute differences in coordinates
> - sum of absolute differences in coordinates - good when data is discrete or with large distances
> - good when data is discrete or with large distances - Minkowski distance
> - Minkowski distance - includes euclidean and manhattan
> - includes euclidean and manhattan - parameter allows to interpolate between the two
> - parameter allows to interpolate between the two - can be used for model tuning since distance function can be changed between euclidean and manhattan
> - can be used for model tuning since distance function can be changed between euclidean and manhattan - features should be standardized to ensure fair distance measures
> - features should be standardized to ensure fair distance measures - Logistic regression
> - Logistic regression - log-odds: natural log of the probability ratio
> - log-odds: natural log of the probability ratio - uses a logistic regression to predict the chances of something being categorized in certain way
> - uses a logistic regression to predict the chances of something being categorized in certain way - logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
> - logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$ - output of logistic regression is compared to threshold T
> - output of logistic regression is compared to threshold T - default T = 0.5
> - default T = 0.5 - linear regression will underfit for classifying data, logistic regression is better
> - linear regression will underfit for classifying data, logistic regression is better - multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
> - multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$ - Gaussian Naive Bayes
> - Gaussian Naive Bayes - normal distributions for each outcome
> - normal distributions for each outcome - Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
> - Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$ - $P(A|B$): Posterior probability
> - $P(A|B$): Posterior probability - Assumptions
> - Assumptions - all input features are independent
> - all input features are independent - all input features contribute equally to classification
> - all input features contribute equally to classification - requires that each probability is 0
> - requires that each probability is 0 - if there is one option that is 0, can add 1 to each to ensure it works
> - if there is one option that is 0, can add 1 to each to ensure it works - Linear Discriminant Analysis
> - Linear Discriminant Analysis - supervised
> - supervised - dimensional reduction
> - dimensional reduction - maximize the distance between groups
> - maximize the distance between groups - within-class variance is minimized
> - within-class variance is minimized - maximize the distances between the means of the two categories on the new axis
> - maximize the distances between the means of the two categories on the new axis - trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$
> - trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$ - Top of equation is the distance between the averages of the data projected onto the new line
> - Top of equation is the distance between the averages of the data projected onto the new line - bottom is minimizing the scatter within each category
> - bottom is minimizing the scatter within each category - discriminant analysis determines decision boundary between classes
> - discriminant analysis determines decision boundary between classes
@@ -1,89 +1,61 @@
#rs/notes #rs/class/csb320 #rs/notes #rs/class/csb320
- - - - - -
> [!NOTE]- Bullet Notes - loss functions
> - loss functions - quantifies differences between predictions and observed values
> - quantifies differences between predictions and observed values - should minimize loss
> - should minimize loss - absolute loss
> - absolute loss - prediction of instance
> - prediction of instance - log loss
> - log loss - penalizes wrong predictions more harshly when they are more confident $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
> - penalizes wrong predictions more harshly when they are more confident $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$ - cross-entropy loss
> - cross-entropy loss - used for multi-class classification
> - used for multi-class classification - measures difference between observed and predicted distributions
> - measures difference between observed and predicted distributions - type I error: false positive
> - type I error: false positive - type II error: false negative
> - type II error: false negative - evaluation metrics
> - evaluation metrics - accuracy $$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$
> - accuracy $$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$ - precision $$\frac{TP}{TP + FP}$$
> - precision $$\frac{TP}{TP + FP}$$ - F-measures
> - F-measures - harmonic mean of precision and recall
> - harmonic mean of precision and recall - beta allows for emphasizing precision or recall in metric: $$F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}$$
> - beta allows for emphasizing precision or recall in metric: $$F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}$$ - $\beta>1$ emphasizes recall
> - $\beta>1$ emphasizes recall - $\beta<1$ emphasizes precision
> - $\beta<1$ emphasizes precision - kappa
> - kappa - overall proportion of correct predictions
> - overall proportion of correct predictions - evaluates performance when compared to random classifier
> - evaluates performance when compared to random classifier - 1: perfect classifier
> - 1: perfect classifier - 0: same accuracy as random chance
> - 0: same accuracy as random chance - <0: worse than random chance
> - <0: worse than random chance - accuracy fails for imbalanced datasets
> - accuracy fails for imbalanced datasets - ex. when most people don't have cancer
> - ex. when most people don't have cancer - AUC-ROC curve
> - AUC-ROC curve - Receiver operating curve
> - Receiver operating curve - AUC represents the degree of seperability
> - AUC represents the degree of seperability - ROC curves
> - ROC curves - shows the trade off between true positive rate and false positive rate
> - shows the trade off between true positive rate and false positive rate - ![[rocCurve.excalidraw]]
> - ![[rocCurve.excalidraw]] - parametric graphs, false positive rate is on the x-axis and true positive rate is on the y-axis
> - parametric graphs, false positive rate is on the x-axis and true positive rate is on the y-axis - parameter is threshold for positive classification
> - parameter is threshold for positive classification - when threshold is raised there are less false positives but also less true positives
> - when threshold is raised there are less false positives but also less true positives - vice versa
> - vice versa - area underneath the curve is a measure of the accuracy
> - area underneath the curve is a measure of the accuracy - AUC (Area under the curve)
> - AUC (Area under the curve) - When to use metrics
> - When to use metrics - accuracy is very bad when dataset is not balanced
> - accuracy is very bad when dataset is not balanced - precision vs. recall
> - precision vs. recall - depends on the situation (don't want to have false negatives for cancer)
> - depends on the situation (don't want to have false negatives for cancer) - F1 can maximize precision and recall and gives balance
> - F1 can maximize precision and recall and gives balance - when balanced classes, maximize accuracy
> - when balanced classes, maximize accuracy - unbalanced classes can prioritize F1 on only one class if one is more important
> - unbalanced classes can prioritize F1 on only one class if one is more important - Holdout method
> - Holdout method - data is randomly partitioned into two independent sets
> - data is randomly partitioned into two independent sets - validation set (often half of testing set) is used to decide optimal hyperparameter values
> - validation set (often half of testing set) is used to decide optimal hyperparameter values - Random subsampling
> - Random subsampling - holdout is repeated k times and accuracy is average
> - holdout is repeated k times and accuracy is average - Cross validation
> - Cross validation - separate into sections, iterate through and use different section as the test set each time
> - separate into sections, iterate through and use different section as the test set each time - average the error
> - average the error - if the average accuracy goes down with cross validation compared to holdout then model is likely overfitting
> - if the average accuracy goes down with cross validation compared to holdout then model is likely overfitting - stratified cross-validaiton
> - stratified cross-validaiton - ensure that the distribution of data is the same in each section as in the general data set
> - ensure that the distribution of data is the same in each section as in the general data set
# Model Validation and Evaluation
There are many ways to evaluate the performance of models. This is necessary to ensure that they are performing at an acceptable level and to help to tune them with different hyperparameters, preprocessing, data, and other pipeline changes.
## Loss Functions
Loss functions are meant to quantify the difference between observed values and the predictions of a model. They can be used to minimize loss and improve model performance.
### Absolute Loss
$$|Y_{i} - \hat{Y}_{i}|$$
Used to measure the difference from the observed value and predicted value by the model.
### Mean Absolute Error
$$MAE=\frac{1}{n}\sum|\hat{Y}_{i} - Y_{i}|$$
Calculates the average absolute loss of a model based on a training set of data. Is a metric that can be used to improve the accuracy of a model and train it. Should be minimized.
Mean squared error is also used to penalize predictions that are further from observed values more heavily.
$$MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2$$
Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model.
### Log Loss
$$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
Log loss is used to optimize the weights of a logistic regression. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used.
Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal.
The log loss function is used to penalize wrong predictions more harshly when they are more confident.
![[LogLoss.excalidraw]]
Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value.
### Cross Entropy Loss
$$CE = -\sum{Y_{i} * \log(\hat{Y}_{i})}$$
@@ -0,0 +1,28 @@
#machine_learning
- - -
There are many ways to evaluate the performance of models. This is necessary to ensure that they are performing at an acceptable level and to help to tune them with different hyperparameters, preprocessing, data, and other pipeline changes.
## Loss Functions
Loss functions are meant to quantify the difference between observed values and the predictions of a model. They can be used to minimize loss and improve model performance.
### Absolute Loss
$$|Y_{i} - \hat{Y}_{i}|$$
Used to measure the difference from the observed value and predicted value by the model.
### Mean Absolute Error
$$MAE=\frac{1}{n}\sum|\hat{Y}_{i} - Y_{i}|$$
Calculates the average absolute loss of a model based on a training set of data. Is a metric that can be used to improve the accuracy of a model and train it. Should be minimized.
Mean squared error is also used to penalize predictions that are further from observed values more heavily.
$$MSE = \frac{1}{n}\sum(\hat{Y}_{i} - Y_{i})^2$$
Since MSE produces a convex curve in the error metric it also allows the use of algorithms like gradient descent optimization to tune the weights of a model.
### Log Loss
$$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
Log loss is used to optimize the weights of a logistic regression. MSE cannot be used since a logistic regression is not linear and the error with respect to weights is not a convex curve. This makes it challenging to find the weights since algorithms like gradient descent cannot be used.
Logistic regressions also only predict between 0 - 1 (probabilities) so any error value will be between 0 - 1 using MSE which is not ideal.
The log loss function is used to penalize wrong predictions more harshly when they are more confident.
![[LogLoss.excalidraw]]
Only one of the terms inside of the parentheses will be non-zero based on if the observed class is 0 or 1. The log loss function then penalizes the incorrect prediction much more heavily the closer it is to the incorrect value.
### Cross Entropy Loss
$$CE = -\sum{Y_{i} * \log(\hat{Y}_{i})}$$