initial commit

This commit is contained in:
ben committed 2026-05-17 12:19:19 -07:00
commit fbd12d4a8e
320 files changed
+124298

No files matched your search

@@ -0,0 +1,65 @@
- Natural language processing
- concerned with the interactions between computers and human languages
- identify the structure and meaning of words, sentences and text
- Applications
- sentiment analysis
- text summarization
- named entity recognition
- data can be structured or unstructured (posts/reviews or address and date formats)
- Regular expressions
- text strings that are used to find patterns in text
- [Pythex](https://pythex.org) is very helpful for testing regex
- regex can be used to prepare data for models
- remove punctuation and replace with spaces
- remove numbers
- remove any extra spaces
- can also convert to lowercase if you are not trying to find named entities
- if not trying to find named entities should convert all to lower case so ex. October = october
- tokenization is converting a string to a list of words
- each token is evaluated separately
- some words are not helpful and should be removed ("a", "the", "in" etc.)
- stemming
- converting words into their most root version
- running -> run
- desperately -> desper
- studies -> studi
- words in sentence and roots have essentially the same meaning
- lemmatization
- ensures that stem version of the word is a real word
- desperately -> desper -> desperate
- bag of words representation
- tokenization -> map words to indexes -> count word occurrences per document.
- first tokenize sentence
- then build vocabulary overall documents in corpus by tokenizing them
- each phrase transformed into vector based on each words frequency in that phrase
- vector dimensions is size of vocabulary
- creates sparse matrix
- can use this to categorize phrases by meaning, compare similarity or sentiment
- cosine similarity matrix
- can be used to compare similarity of phrases
- 1.0 means that it is the same phrase
- Term Frequency - Inverse Document Frequency
- weights words based on importance across documents
- high weight to word that appears often in one document but not in many documents in the corpus
- avoids words like "the" and "in"
- words like these are likely to be very descriptive of the contents of the document
- $$tfidf(w,d) = tf * \log\left( \frac{N+1}{N_{w} + 1} \right) + 1$$
- $N_{w}$: number of documents in training set that word $w$ appears in
- $N$: number of documents in training set
- $tf$ (term frequency): number of times that the word $w$ appears in the query document $d$
- $tf(w, d) = \frac{{\text{number of w in d}}}{\text{total words in d}}$
- importance of word in specific document in comparison to the entire corpus
- N-grams
- splitting sentences into groups of n words instead of individual words when tokenizing
- can help to analyze relationships between words
- Topic modeling
- unsupervised learning to group them into topics
- Latent Dirichlet Allocation (LDA) is a popular technique
- Logistic regression, Naive Bayes and SVMs can be used for text classification
- deep learning is best way to do this
- Word embeddings
- vector representations of words
- captures meaning, relationships to other words and context in a continuous vector space
- ex. king - man + woman = queen
- models can dynamically create vector embeddings based on context for words with different meanings
-