Get in Touch

Course Outline

Detailed training outline

  1. Introduction to NLP
    • Fundamentals of NLP
    • NLP Frameworks
    • Commercial use cases for NLP
    • Web data scraping techniques
    • Utilizing various APIs to acquire text data
    • Managing text corpora, including content storage and metadata handling
    • Benefits of using Python and an NLTK crash course
  2. Practical Understanding of a Corpus and Dataset
    • The necessity of a corpus
    • Corpus Analysis techniques
    • Categorization of data attributes
    • Various file formats for corpora
    • Dataset preparation for NLP applications
  3. Understanding the Structure of a Sentences
    • Core components of NLP
    • Natural language understanding
    • Morphological analysis: stems, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Addressing ambiguity
  4. Text data preprocessing
    • Raw text corpus handling
      • Sentence tokenization
      • Stemming raw text
      • Lemmatization of raw text
      • Stop word elimination
    • Raw sentence corpus handling
      • Word tokenization
      • Word lemmatization
    • Handling Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Customized and practical preprocessing strategies
  5. Analyzing Text data
    • Fundamental NLP features
      • Parsers and parsing mechanisms
      • Part-of-speech (POS) tagging and taggers
      • Named entity recognition
      • N-grams
      • Bag of words approach
    • Statistical features in NLP
      • Linear algebra concepts applied to NLP
      • Probabilistic theory in NLP
      • TF-IDF
      • Vectorization
      • Encoders and Decoders
      • Normalization
      • Probabilistic Models
    • Advanced feature engineering and NLP
      • Foundations of word2vec
      • Internal components of the word2vec model
      • Operational logic of the word2vec model
      • Extensions of the word2vec concept
      • Practical application of the word2vec model
    • Case study: Applying bag of words for automatic text summarization using simplified and original Luhn's algorithms
  6. Document Clustering, Classification and Topic Modeling
    • Document clustering and pattern mining (including hierarchical clustering and k-means)
    • Document comparison and classification using TF-IDF, Jaccard, and cosine distance metrics
    • Document classification via Naïve Bayes and Maximum Entropy
  7. Identifying Important Text Elements
    • Dimensionality reduction: Principal Component Analysis, Singular Value Decomposition, and Non-negative Matrix Factorization
    • Topic modeling and information retrieval utilizing Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis and Advanced Topic Modeling
    • Differentiating positive vs. negative sentiment degrees
    • Item Response Theory
    • Part-of-speech tagging applications: extracting people, places, and organizations from text
    • Advanced topic modeling: Latent Dirichlet Allocation
  9. Case studies
    • Extracting insights from unstructured user reviews
    • Sentiment classification and visualization of Product Review Data
    • Mining search logs to identify usage patterns
    • Text classification
    • Topic modelling

Requirements

Familiarity with NLP principles and an understanding of AI applications within business contexts

 21 Hours

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories