67 lines
3.3 KiB
Markdown
67 lines
3.3 KiB
Markdown
#rs/notes #rs/class/csb320
|
|
- - -
|
|
- Natural language processing
|
|
- concerned with the interactions between computers and human languages
|
|
- identify the structure and meaning of words, sentences and text
|
|
- Applications
|
|
- sentiment analysis
|
|
- text summarization
|
|
- named entity recognition
|
|
- data can be structured or unstructured (posts/reviews or address and date formats)
|
|
- Regular expressions
|
|
- text strings that are used to find patterns in text
|
|
- [Pythex](https://pythex.org) is very helpful for testing regex
|
|
- regex can be used to prepare data for models
|
|
- remove punctuation and replace with spaces
|
|
- remove numbers
|
|
- remove any extra spaces
|
|
- can also convert to lowercase if you are not trying to find named entities
|
|
- if not trying to find named entities should convert all to lower case so ex. October = october
|
|
- tokenization is converting a string to a list of words
|
|
- each token is evaluated separately
|
|
- some words are not helpful and should be removed ("a", "the", "in" etc.)
|
|
- stemming
|
|
- converting words into their most root version
|
|
- running -> run
|
|
- desperately -> desper
|
|
- studies -> studi
|
|
- words in sentence and roots have essentially the same meaning
|
|
- lemmatization
|
|
- ensures that stem version of the word is a real word
|
|
- desperately -> desper -> desperate
|
|
- bag of words representation
|
|
- tokenization -> map words to indexes -> count word occurrences per document.
|
|
- first tokenize sentence
|
|
- then build vocabulary overall documents in corpus by tokenizing them
|
|
- each phrase transformed into vector based on each words frequency in that phrase
|
|
- vector dimensions is size of vocabulary
|
|
- creates sparse matrix
|
|
- can use this to categorize phrases by meaning, compare similarity or sentiment
|
|
- cosine similarity matrix
|
|
- can be used to compare similarity of phrases
|
|
- 1.0 means that it is the same phrase
|
|
- Term Frequency - Inverse Document Frequency
|
|
- weights words based on importance across documents
|
|
- high weight to word that appears often in one document but not in many documents in the corpus
|
|
- avoids words like "the" and "in"
|
|
- words like these are likely to be very descriptive of the contents of the document
|
|
- $$tfidf(w,d) = tf * \log\left( \frac{N+1}{N_{w} + 1} \right) + 1$$
|
|
- $N_{w}$: number of documents in training set that word $w$ appears in
|
|
- $N$: number of documents in training set
|
|
- $tf$ (term frequency): number of times that the word $w$ appears in the query document $d$
|
|
- $tf(w, d) = \frac{{\text{number of w in d}}}{\text{total words in d}}$
|
|
- importance of word in specific document in comparison to the entire corpus
|
|
- N-grams
|
|
- splitting sentences into groups of n words instead of individual words when tokenizing
|
|
- can help to analyze relationships between words
|
|
- Topic modeling
|
|
- unsupervised learning to group them into topics
|
|
- Latent Dirichlet Allocation (LDA) is a popular technique
|
|
- Logistic regression, Naive Bayes and SVMs can be used for text classification
|
|
- deep learning is best way to do this
|
|
- Word embeddings
|
|
- vector representations of words
|
|
- captures meaning, relationships to other words and context in a continuous vector space
|
|
- ex. king - man + woman = queen
|
|
- models can dynamically create vector embeddings based on context for words with different meanings
|
|
- |