initial commit

This commit is contained in:
ben committed 2026-05-17 12:19:19 -07:00
commit fbd12d4a8e
320 files changed
+124298

No files matched your search

@@ -0,0 +1,118 @@
#rs/notes #rs/class/csb320
- - -
> [!NOTE]- Bullet Notes
> - Classification models
> - supervised
> - classify instances into class (category)
> - can predict class directly or probabilities
> - Each column in data table is feature
> - Each row is sample/tuple/instance
> - K-Nearest neighbors
> - lazy learner
> - memorizes data, doesn't create model during training
> - if k=6 it looks at 6 closes neighbors in dataset
> - predicts based on what these are classified as
> - there can be very different predictions based on k
> - gets very expensive as the dataset grows
> - selecting value of k
> - k is a hyperparameter (user set value )
> - typical values of 3-15
> - lower value can be sensitive to noise
> - higher value risks underfitting
> - distance measures
> - euclidean distance
> - good when data is compact and continuous
> - manhattan distance
> - sum of absolute differences in coordinates
> - good when data is discrete or with large distances
> - Minkowski distance
> - includes euclidean and manhattan
> - parameter allows to interpolate between the two
> - can be used for model tuning since distance function can be changed between euclidean and manhattan
> - features should be standardized to ensure fair distance measures
> - Logistic regression
> - log-odds: natural log of the probability ratio
> - uses a logistic regression to predict the chances of something being categorized in certain way
> - logistic regression $$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
> - output of logistic regression is compared to threshold T
> - default T = 0.5
> - linear regression will underfit for classifying data, logistic regression is better
> - multiple input features: $$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
> - Gaussian Naive Bayes
> - normal distributions for each outcome
> - Baye's rule: $$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
> - $P(A|B$): Posterior probability
> - Assumptions
> - all input features are independent
> - all input features contribute equally to classification
> - requires that each probability is 0
> - if there is one option that is 0, can add 1 to each to ensure it works
> - Linear Discriminant Analysis
> - supervised
> - dimensional reduction
> - maximize the distance between groups
> - within-class variance is minimized
> - maximize the distances between the means of the two categories on the new axis
> - trying to maximize: $$\frac{(\mu_{1}-\mu_{2})^2}{s_{1}^2-s_{2}^2}$$
> - Top of equation is the distance between the averages of the data projected onto the new line
> - bottom is minimizing the scatter within each category
> - discriminant analysis determines decision boundary between classes
# Classification Models
Classification models are meant to categorize instances into new classes based on data and training. They can be supervised (using labeled data) and unsupervised (using unlabeled data). They can also either predict classes directly or the probabilities of belonging to certain classes.
### Dataset Terminology
| | Feature | Feature | Feature |
| -------- | ------- | ------- | ------- |
| Sample | | | |
| Tuple | | | |
| Instance | | | |
## Models
### K Nearest Neighbors
Attributes
- Lazy learner (memorizes data, doesn't create model during training)
- Gets very expensive as the dataset grows
- Requires normalization since it is based on distance measures
Prediction process
![[KNN.excalidraw]]
A KNN model selects the k closest points in the dataset and then predicts based on the most common class in this set. Ties are often broken by the class of the closest point.
Several different distance metrics can be used:
![[DistanceMetrics.excalidraw]]
- Euclidean Distance
- Good for when data is compact and continuous $$d(x,y) = (\sum_{i=1}^n {|x_{i} - y_{i}|}^2)^{ \frac{1}{2} }$$
- Manhattan Distance
- Good for when data is discrete or has large distances $$d(x,y) = \sum_{i=1}^n |x_{i} - y_{i}|$$
- Minkowski Distance
- Allows for interpolation between Euclidean and Manhattan distance $$d(x,y) = (\sum_{i=1}^n {|x_{i} - y_{i}|}^p)^{ \frac{1}{p} }$$
- if p = 1 it is the same as Manhattan distance
- if p = 2 it is the same as Euclidean distance
### Logistic Regression
A logistic regression is used to predict the probability of an instance belonging to a certain class. Linear regressions will often underfit data so a logistic regression can be a better choice.
![[LogisticVsLinear.excalidraw]]
The equation for a logistic regression is:
$$\hat{p}=\frac{\exp(w_0 + w_1x_i)}{1 + \exp(w_{0} + w_{1}x_{i})}$$
The output of this regression is compared to the threshold, T. This is usually set to 0.5 but can be changed to bias towards one class.
Logistic regressions can also be used with multiple input features rather than just one with the equation:
$$\hat{p}= \frac{\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}{1+\exp(w_{0} + w_{1}x_{1i}+\dots+w_{p}x_{p i})}$$
### Gaussian Naive Bayes
A Gaussian NB model assumes that classes follow a Gaussian distribution and uses that to calculate the probability an instance will belong to each class.
$$P(x_{i}|y) = \frac{1}{\sigma \sqrt{ 2\pi }}e^{-\frac{(x-\mu)^2}{2\sigma^2}}$$
$x_{i}$ is the feature value
$\mu$ is the mean of the feature values for a given class $y_{k}$
$\sigma$ is the standard deviation of the feature values for the class
It also uses Bayes rule to calculate posterior probabilities:
$$P(A|B) = \frac{P(B|A) * P(A)}{P(B)}$$
The posterior probability ( $P(A|B)$ ) is the probability that $A$ happens given that $B$ has happened
The algorithm is referred to as "naive" because it makes a few assumptions:
- There is no correlation between the features in the dataset, they are all independent
- Each feature has an equal importance when predicting the output class
@@ -0,0 +1,60 @@
#rs/notes #rs/class/csb320
- - -
- loss functions
- quantifies differences between predictions and observed values
- should minimize loss
- absolute loss
- prediction of instance
- log loss
- penalizes wrong predictions more harshly when they are more confident $$L_{\log}(y_{i}, \hat{p}_{i}) = -(y_{i}\ln(\hat{p}_{i}) + (1-y_{i})\ln(1-\hat{p}_{i}))$$
- cross-entropy loss
- used for multi-class classification
- measures difference between observed and predicted distributions
- type I error: false positive
- type II error: false negative
- evaluation metrics
- accuracy $$\frac{\text{num correct predictions}}{\text{num incorrect predictions}}$$
- precision $$\frac{TP}{TP + FP}$$
- F-measures
- harmonic mean of precision and recall
- beta allows for emphasizing precision or recall in metric: $$F_{\beta} = (1 + \beta^2)\frac{{\text{precision} * \text{recall}}}{\beta^2 * \text{precision} + \text{recall}}$$
- $\beta>1$ emphasizes recall
- $\beta<1$ emphasizes precision
- kappa
- overall proportion of correct predictions
- evaluates performance when compared to random classifier
- 1: perfect classifier
- 0: same accuracy as random chance
- <0: worse than random chance
- accuracy fails for imbalanced datasets
- ex. when most people don't have cancer
- AUC-ROC curve
- Receiver operating curve
- AUC represents the degree of seperability
- ROC curves
- shows the trade off between true positive rate and false positive rate
- ![[rocCurve.excalidraw]]
- parametric graphs, false positive rate is on the x-axis and true positive rate is on the y-axis
- parameter is threshold for positive classification
- when threshold is raised there are less false positives but also less true positives
- vice versa
- area underneath the curve is a measure of the accuracy
- AUC (Area under the curve)
- When to use metrics
- accuracy is very bad when dataset is not balanced
- precision vs. recall
- depends on the situation (don't want to have false negatives for cancer)
- F1 can maximize precision and recall and gives balance
- when balanced classes, maximize accuracy
- unbalanced classes can prioritize F1 on only one class if one is more important
- Holdout method
- data is randomly partitioned into two independent sets
- validation set (often half of testing set) is used to decide optimal hyperparameter values
- Random subsampling
- holdout is repeated k times and accuracy is average
- Cross validation
- separate into sections, iterate through and use different section as the test set each time
- average the error
- if the average accuracy goes down with cross validation compared to holdout then model is likely overfitting
- stratified cross-validaiton
- ensure that the distribution of data is the same in each section as in the general data set
@@ -0,0 +1,97 @@
#rs/notes #rs/class/csb320
- - -
- Data cleaning
- incomplete
- noisy (negative, wrong)
- inconsistent
- incorrect age vs. birthday
- intentional
- ex. not actually adding values, same for everything
- Missing data
- missing completely at random
- missing data completely at random
- missing at random
- missing based on variables but not missing values
- ex. no garage space data in specific location
- missing not at random
- missing due to unobserved factors or systematic reasons
- handling missing values
- constant value
- attribute mean
- most probable value
- can use regression models to predict missing features
- KNN imputation
- estimates based on k most similar instances
- handling noisy data
- smooth data by partitioning into bins
- smooth data by fitting into regression
- clustering to detect and remove outliers
- pairwise deletion: only delete missing/incorrect values
- listwise deletion: delete entire row if missing value
- some features need to be removed
- duplicates
- high proportion of missing values
- all same
- no pattern
- Feature transformation
- continuous -> discrete (simplification)
- reduce data noise
- decision trees can perform better with discrete data
- equal frequency or equal interval binning
- normalization
- scales data to range between 0 and 1
- standardization
- scales to have a mean of 0 and standard deviation of 1
- encoding categorical features
- binary encoding
- categories as binary vectors (00, 01, 10)
- label encoding
- unique integers for each category (1, 2, 3)
- one-hot encoding
- convert categories into binary vectors
- decision trees and random forests care about data split so they work better with label or binary encoding
- logistic or linear regression work better with one-hot encoding since they care about distance between data points
- imbalanced data
- evaluation metrics
- different ones work better with imbalanced data
- data-level methods
- undersampling
- reduce majority class to balance
- the minority class needs to be large enough to still have enough data
- random undersampling
- Tomek links
- pair from majority and minority class, are each others closest neighbors
- removing pairs can help to create clearer separation between classes
- can make decision boundary more clear
- oversampling
- increase minority class through SMOTE
- ![[SMOTE.excalidraw]]
- synthetic minority oversampling technique
- increases size of minority and variety
- find k nearest minority neighbors, select j of them
- add new synthetic data along line between these data points
- can overgeneralize data
- algorithm-level methods
- can weight classes differently
- assigns misclassification costs
- algorithm selection (random forest and adaboost handle imbalanced data better)
- threshold adjustment
- dimensionality
- too many features creates sparse data
- can cause overfitting
- distance becomes less meaningful
- reduction
- simplifies representation, improves performance
- can preserve essential information
- feature selection
- ranked with statistical tests
- iterative model building and testing
- linear -> straight line
- monotonic -> moving in same direction
- want a linear/monotonic relationship between target class and each feature
- do not want linear/monotonic relationship between feature classes
- different metrics depending on if inputs/outputs are numerical or categorical
- PCA (principle component analysis)
- creates new uncorrelated features (principle components)
- maintains the maximum amount of variance in the data
-
@@ -0,0 +1,65 @@
- Natural language processing
- concerned with the interactions between computers and human languages
- identify the structure and meaning of words, sentences and text
- Applications
- sentiment analysis
- text summarization
- named entity recognition
- data can be structured or unstructured (posts/reviews or address and date formats)
- Regular expressions
- text strings that are used to find patterns in text
- [Pythex](https://pythex.org) is very helpful for testing regex
- regex can be used to prepare data for models
- remove punctuation and replace with spaces
- remove numbers
- remove any extra spaces
- can also convert to lowercase if you are not trying to find named entities
- if not trying to find named entities should convert all to lower case so ex. October = october
- tokenization is converting a string to a list of words
- each token is evaluated separately
- some words are not helpful and should be removed ("a", "the", "in" etc.)
- stemming
- converting words into their most root version
- running -> run
- desperately -> desper
- studies -> studi
- words in sentence and roots have essentially the same meaning
- lemmatization
- ensures that stem version of the word is a real word
- desperately -> desper -> desperate
- bag of words representation
- tokenization -> map words to indexes -> count word occurrences per document.
- first tokenize sentence
- then build vocabulary overall documents in corpus by tokenizing them
- each phrase transformed into vector based on each words frequency in that phrase
- vector dimensions is size of vocabulary
- creates sparse matrix
- can use this to categorize phrases by meaning, compare similarity or sentiment
- cosine similarity matrix
- can be used to compare similarity of phrases
- 1.0 means that it is the same phrase
- Term Frequency - Inverse Document Frequency
- weights words based on importance across documents
- high weight to word that appears often in one document but not in many documents in the corpus
- avoids words like "the" and "in"
- words like these are likely to be very descriptive of the contents of the document
- $$tfidf(w,d) = tf * \log\left( \frac{N+1}{N_{w} + 1} \right) + 1$$
- $N_{w}$: number of documents in training set that word $w$ appears in
- $N$: number of documents in training set
- $tf$ (term frequency): number of times that the word $w$ appears in the query document $d$
- $tf(w, d) = \frac{{\text{number of w in d}}}{\text{total words in d}}$
- importance of word in specific document in comparison to the entire corpus
- N-grams
- splitting sentences into groups of n words instead of individual words when tokenizing
- can help to analyze relationships between words
- Topic modeling
- unsupervised learning to group them into topics
- Latent Dirichlet Allocation (LDA) is a popular technique
- Logistic regression, Naive Bayes and SVMs can be used for text classification
- deep learning is best way to do this
- Word embeddings
- vector representations of words
- captures meaning, relationships to other words and context in a continuous vector space
- ex. king - man + woman = queen
- models can dynamically create vector embeddings based on context for words with different meanings
-
@@ -0,0 +1,44 @@
#rs/notes #rs/class/csb320
- - -
- Support vector machines
- works for linear and nonlinear data
- if nonlinear will map data into higher dimension
- attempts to find optimal linear separating hyperplane
- training can be slow but model is accurate
- Margins expand as much as they can past the decision boundary until hitting the closest points
- ![[SupportVectorMachine.excalidraw]]
- algorithm attempts to maximize margins in order to have most distance between classes
- "support vectors" are the closest points to the decision boundary
- margins: perpendicular distance from the hyperplane to closest instance
- no probabilities are given
- mapping functions are used to map data into higher dimensional space
- inner product: function that combines two vectors to one scalar value (dot product)
- different kernels can be used
- polynomial kernel: good when data is not linearly separable but has regular curved boundary
- RBF: default when boundary is complex or unknown
- Sigmoid: good when modeling data similar to neural network behavior.
- Decision trees
- greedy, continues forward and does not backtrack
- features must be categorical, discretize continuous features beforehand
- conditions for stopping partitioning
- all samples belong to same class for certain node
- no remaining attributes for partitioning
- no samples left
- each leaf node represents a predicted class
- decisions
- numerical uses inequalities
- categorical uses equality
- decision trees divide feature space with hyperplanes perpendicular to decision feature's axis
- measure of fit
- node is completely pure if all instances belong to same class
- impurity measures include gini coefficient, entropy
- gini coefficient
- imputiry reaches a max at 0.5 (classes are evenly split)
- more of one class or another means that data is less split
- entropy or log loss
- negative ensures positive purity value
- not used quite as much
- overfitting can occur if tree gets too deep
- should stop tree early
- can also prune leaves
@@ -0,0 +1,3 @@
#rs/discussion #rs/class/csb320
- - -
I've been enjoying the course so far and learning about new machine learning ideas that I haven't spent much time on before. At first it was a little strange that the lectures, ZyBooks and Kaggle cover a lot of the same content but now I like how it has helped to reinforce the concepts for me and give me practice. I feel more confident with all the different functions and processes and I think the repetition has helped. I'm looking forward to learning more about machine learning and getting better at using the tools!