#rs/notes #rs/class/csb320 - - - - Support vector machines - works for linear and nonlinear data - if nonlinear will map data into higher dimension - attempts to find optimal linear separating hyperplane - training can be slow but model is accurate - Margins expand as much as they can past the decision boundary until hitting the closest points - ![[SupportVectorMachine.excalidraw]] - algorithm attempts to maximize margins in order to have most distance between classes - "support vectors" are the closest points to the decision boundary - margins: perpendicular distance from the hyperplane to closest instance - no probabilities are given - mapping functions are used to map data into higher dimensional space - inner product: function that combines two vectors to one scalar value (dot product) - different kernels can be used - polynomial kernel: good when data is not linearly separable but has regular curved boundary - RBF: default when boundary is complex or unknown - Sigmoid: good when modeling data similar to neural network behavior. - Decision trees - greedy, continues forward and does not backtrack - features must be categorical, discretize continuous features beforehand - conditions for stopping partitioning - all samples belong to same class for certain node - no remaining attributes for partitioning - no samples left - each leaf node represents a predicted class - decisions - numerical uses inequalities - categorical uses equality - decision trees divide feature space with hyperplanes perpendicular to decision feature's axis - measure of fit - node is completely pure if all instances belong to same class - impurity measures include gini coefficient, entropy - gini coefficient - imputiry reaches a max at 0.5 (classes are evenly split) - more of one class or another means that data is less split - entropy or log loss - negative ensures positive purity value - not used quite as much - overfitting can occur if tree gets too deep - should stop tree early - can also prune leaves