Get in Touch
 Duration 21 hours (3 days)

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief overview of Python and Scala

Foundational Theory:

  • Architecture
  • RDD (Resilient Distributed Datasets)
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Practical Exploration via Databricks (Hands-on Workshop):

  • Exercises utilizing the RDD API
  • Core action and transformation functions
  • Working with PairRDD
  • Join operations
  • Caching strategies
  • Exercises utilizing the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • UDF (User Defined Functions)
  • Overview of the DataSet API
  • Streaming concepts

Deployment Mastery via AWS (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Distinguishing between AWS EMR and AWS Glue
  • Implementing example jobs in both environments
  • Evaluating the advantages and limitations of each

Additional Topics:

  • Introduction to Apache Airflow for orchestration

Requirements

Programming skills (preferably in Python or Scala)

Foundational knowledge of SQL

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories