Files
2026-05-24 20:21:50 -07:00

3.3 KiB

#rs/notes #rs/class/csb320


  • Natural language processing
    • concerned with the interactions between computers and human languages
    • identify the structure and meaning of words, sentences and text
    • Applications
      • sentiment analysis
      • text summarization
      • named entity recognition
    • data can be structured or unstructured (posts/reviews or address and date formats)
    • Regular expressions
      • text strings that are used to find patterns in text
      • Pythex is very helpful for testing regex
      • regex can be used to prepare data for models
        • remove punctuation and replace with spaces
        • remove numbers
        • remove any extra spaces
        • can also convert to lowercase if you are not trying to find named entities
          • if not trying to find named entities should convert all to lower case so ex. October = october
    • tokenization is converting a string to a list of words
      • each token is evaluated separately
      • some words are not helpful and should be removed ("a", "the", "in" etc.)
    • stemming
      • converting words into their most root version
      • running -> run
      • desperately -> desper
      • studies -> studi
      • words in sentence and roots have essentially the same meaning
    • lemmatization
      • ensures that stem version of the word is a real word
      • desperately -> desper -> desperate
  • bag of words representation
    • tokenization -> map words to indexes -> count word occurrences per document.
    • first tokenize sentence
    • then build vocabulary overall documents in corpus by tokenizing them
    • each phrase transformed into vector based on each words frequency in that phrase
      • vector dimensions is size of vocabulary
      • creates sparse matrix
    • can use this to categorize phrases by meaning, compare similarity or sentiment
    • cosine similarity matrix
      • can be used to compare similarity of phrases
      • 1.0 means that it is the same phrase
  • Term Frequency - Inverse Document Frequency
    • weights words based on importance across documents
    • high weight to word that appears often in one document but not in many documents in the corpus
      • avoids words like "the" and "in"
      • words like these are likely to be very descriptive of the contents of the document
    • tfidf(w,d) = tf * \log\left( \frac{N+1}{N_{w} + 1} \right) + 1
    • N_{w}: number of documents in training set that word w appears in
    • N: number of documents in training set
    • tf (term frequency): number of times that the word w appears in the query document d
      • tf(w, d) = \frac{{\text{number of w in d}}}{\text{total words in d}}
    • importance of word in specific document in comparison to the entire corpus
  • N-grams
    • splitting sentences into groups of n words instead of individual words when tokenizing
    • can help to analyze relationships between words
  • Topic modeling
    • unsupervised learning to group them into topics
    • Latent Dirichlet Allocation (LDA) is a popular technique
    • Logistic regression, Naive Bayes and SVMs can be used for text classification
      • deep learning is best way to do this
  • Word embeddings
    • vector representations of words
    • captures meaning, relationships to other words and context in a continuous vector space
    • ex. king - man + woman = queen
    • models can dynamically create vector embeddings based on context for words with different meanings