initial commit
This commit is contained in:
commit
fbd12d4a8e
320 files changed
+124298
No files matched your search
+65
@@ -0,0 +1,65 @@
|
||||
- Natural language processing
|
||||
- concerned with the interactions between computers and human languages
|
||||
- identify the structure and meaning of words, sentences and text
|
||||
- Applications
|
||||
- sentiment analysis
|
||||
- text summarization
|
||||
- named entity recognition
|
||||
- data can be structured or unstructured (posts/reviews or address and date formats)
|
||||
- Regular expressions
|
||||
- text strings that are used to find patterns in text
|
||||
- [Pythex](https://pythex.org) is very helpful for testing regex
|
||||
- regex can be used to prepare data for models
|
||||
- remove punctuation and replace with spaces
|
||||
- remove numbers
|
||||
- remove any extra spaces
|
||||
- can also convert to lowercase if you are not trying to find named entities
|
||||
- if not trying to find named entities should convert all to lower case so ex. October = october
|
||||
- tokenization is converting a string to a list of words
|
||||
- each token is evaluated separately
|
||||
- some words are not helpful and should be removed ("a", "the", "in" etc.)
|
||||
- stemming
|
||||
- converting words into their most root version
|
||||
- running -> run
|
||||
- desperately -> desper
|
||||
- studies -> studi
|
||||
- words in sentence and roots have essentially the same meaning
|
||||
- lemmatization
|
||||
- ensures that stem version of the word is a real word
|
||||
- desperately -> desper -> desperate
|
||||
- bag of words representation
|
||||
- tokenization -> map words to indexes -> count word occurrences per document.
|
||||
- first tokenize sentence
|
||||
- then build vocabulary overall documents in corpus by tokenizing them
|
||||
- each phrase transformed into vector based on each words frequency in that phrase
|
||||
- vector dimensions is size of vocabulary
|
||||
- creates sparse matrix
|
||||
- can use this to categorize phrases by meaning, compare similarity or sentiment
|
||||
- cosine similarity matrix
|
||||
- can be used to compare similarity of phrases
|
||||
- 1.0 means that it is the same phrase
|
||||
- Term Frequency - Inverse Document Frequency
|
||||
- weights words based on importance across documents
|
||||
- high weight to word that appears often in one document but not in many documents in the corpus
|
||||
- avoids words like "the" and "in"
|
||||
- words like these are likely to be very descriptive of the contents of the document
|
||||
- $$tfidf(w,d) = tf * \log\left( \frac{N+1}{N_{w} + 1} \right) + 1$$
|
||||
- $N_{w}$: number of documents in training set that word $w$ appears in
|
||||
- $N$: number of documents in training set
|
||||
- $tf$ (term frequency): number of times that the word $w$ appears in the query document $d$
|
||||
- $tf(w, d) = \frac{{\text{number of w in d}}}{\text{total words in d}}$
|
||||
- importance of word in specific document in comparison to the entire corpus
|
||||
- N-grams
|
||||
- splitting sentences into groups of n words instead of individual words when tokenizing
|
||||
- can help to analyze relationships between words
|
||||
- Topic modeling
|
||||
- unsupervised learning to group them into topics
|
||||
- Latent Dirichlet Allocation (LDA) is a popular technique
|
||||
- Logistic regression, Naive Bayes and SVMs can be used for text classification
|
||||
- deep learning is best way to do this
|
||||
- Word embeddings
|
||||
- vector representations of words
|
||||
- captures meaning, relationships to other words and context in a continuous vector space
|
||||
- ex. king - man + woman = queen
|
||||
- models can dynamically create vector embeddings based on context for words with different meanings
|
||||
-
|
||||
Reference in new issue
Block a user