initial commit
This commit is contained in:
commit
fbd12d4a8e
320 files changed
+124298
No files matched your search
@@ -0,0 +1,44 @@
|
||||
#rs/notes #rs/class/csb320
|
||||
- - -
|
||||
- Support vector machines
|
||||
- works for linear and nonlinear data
|
||||
- if nonlinear will map data into higher dimension
|
||||
- attempts to find optimal linear separating hyperplane
|
||||
- training can be slow but model is accurate
|
||||
- Margins expand as much as they can past the decision boundary until hitting the closest points
|
||||
- ![[SupportVectorMachine.excalidraw]]
|
||||
- algorithm attempts to maximize margins in order to have most distance between classes
|
||||
- "support vectors" are the closest points to the decision boundary
|
||||
- margins: perpendicular distance from the hyperplane to closest instance
|
||||
- no probabilities are given
|
||||
- mapping functions are used to map data into higher dimensional space
|
||||
- inner product: function that combines two vectors to one scalar value (dot product)
|
||||
- different kernels can be used
|
||||
- polynomial kernel: good when data is not linearly separable but has regular curved boundary
|
||||
- RBF: default when boundary is complex or unknown
|
||||
- Sigmoid: good when modeling data similar to neural network behavior.
|
||||
- Decision trees
|
||||
- greedy, continues forward and does not backtrack
|
||||
- features must be categorical, discretize continuous features beforehand
|
||||
- conditions for stopping partitioning
|
||||
- all samples belong to same class for certain node
|
||||
- no remaining attributes for partitioning
|
||||
- no samples left
|
||||
- each leaf node represents a predicted class
|
||||
- decisions
|
||||
- numerical uses inequalities
|
||||
- categorical uses equality
|
||||
- decision trees divide feature space with hyperplanes perpendicular to decision feature's axis
|
||||
- measure of fit
|
||||
- node is completely pure if all instances belong to same class
|
||||
- impurity measures include gini coefficient, entropy
|
||||
- gini coefficient
|
||||
- imputiry reaches a max at 0.5 (classes are evenly split)
|
||||
- more of one class or another means that data is less split
|
||||
- entropy or log loss
|
||||
- negative ensures positive purity value
|
||||
- not used quite as much
|
||||
- overfitting can occur if tree gets too deep
|
||||
- should stop tree early
|
||||
- can also prune leaves
|
||||
|
||||
Reference in new issue
Block a user