- Natural language processing - concerned with the interactions between computers and human languages - identify the structure and meaning of words, sentences and text - Applications - sentiment analysis - text summarization - named entity recognition - data can be structured or unstructured (posts/reviews or address and date formats) - Regular expressions - text strings that are used to find patterns in text - [Pythex](https://pythex.org) is very helpful for testing regex - regex can be used to prepare data for models - remove punctuation and replace with spaces - remove numbers - remove any extra spaces - can also convert to lowercase if you are not trying to find named entities - if not trying to find named entities should convert all to lower case so ex. October = october - tokenization is converting a string to a list of words - each token is evaluated separately - some words are not helpful and should be removed ("a", "the", "in" etc.) - stemming - converting words into their most root version - running -> run - desperately -> desper - studies -> studi - words in sentence and roots have essentially the same meaning - lemmatization - ensures that stem version of the word is a real word - desperately -> desper -> desperate - bag of words representation - tokenization -> map words to indexes -> count word occurrences per document. - first tokenize sentence - then build vocabulary overall documents in corpus by tokenizing them - each phrase transformed into vector based on each words frequency in that phrase - vector dimensions is size of vocabulary - creates sparse matrix - can use this to categorize phrases by meaning, compare similarity or sentiment - cosine similarity matrix - can be used to compare similarity of phrases - 1.0 means that it is the same phrase - Term Frequency - Inverse Document Frequency - weights words based on importance across documents - high weight to word that appears often in one document but not in many documents in the corpus - avoids words like "the" and "in" - words like these are likely to be very descriptive of the contents of the document - $$tfidf(w,d) = tf * \log\left( \frac{N+1}{N_{w} + 1} \right) + 1$$ - $N_{w}$: number of documents in training set that word $w$ appears in - $N$: number of documents in training set - $tf$ (term frequency): number of times that the word $w$ appears in the query document $d$ - $tf(w, d) = \frac{{\text{number of w in d}}}{\text{total words in d}}$ - importance of word in specific document in comparison to the entire corpus - N-grams - splitting sentences into groups of n words instead of individual words when tokenizing - can help to analyze relationships between words - Topic modeling - unsupervised learning to group them into topics - Latent Dirichlet Allocation (LDA) is a popular technique - Logistic regression, Naive Bayes and SVMs can be used for text classification - deep learning is best way to do this - Word embeddings - vector representations of words - captures meaning, relationships to other words and context in a continuous vector space - ex. king - man + woman = queen - models can dynamically create vector embeddings based on context for words with different meanings -