Overview

This course provides an in-depth exploration of Apache Spark and Delta Lake on Databricks, focusing on the core architectural components of Spark, the DataFrame API, and Structured Streaming. Participants will learn how to efficiently read, transform, and aggregate data using SparkSQL and the DataFrame API. The course also covers user-defined functions (UDFs), query optimization, partitioning strategies, and the advantages of Delta Lake for improving data pipelines. By the end of the course, learners will be able to execute streaming queries and understand how Delta Lake enhances real-time data processing.

Read more +

Prerequisites

Participants should have:

  • Familiarity with Python and fundamental programming concepts, including data types, lists, dictionaries, variables, functions, loops, conditional statements, exception handling, accessing classes, and using third-party libraries.
  • Basic knowledge of SQL, including writing queries using SELECT, WHERE, GROUP BY, ORDER BY, LIMIT, and JOIN.

If you do not have one or more of the pre-requisites QA recommends:

Target Audience

This course is designed for:

  • Data engineers and data scientists looking to enhance their Spark programming skills.
  • Developers who want to leverage Apache Spark and Delta Lake on Databricks.
  • Professionals working with large-scale data processing and real-time analytics.
Read more +

Delegates will learn how to

By the end of this course, learners will be able to:

  • Describe the architecture and core components of Apache Spark.
  • Implement data transformations using the DataFrame API.
  • Optimise Spark queries for performance improvements.
  • Apply partitioning strategies to manage large datasets efficiently.
  • Use Structured Streaming to process real-time data.
  • Implement Delta Lake to enhance data reliability and performance.
Read more +

Outline

Introduction to Spark and Databricks

  • Overview of Apache Spark and its role in big data processing.
  • Introduction to the Databricks platform.

Working with SparkSQL and DataFrames

  • Understanding SparkSQL and its use cases.
  • DataFrame operations: reading, writing, transformations, and aggregations.
  • Working with complex data types and datetime functions.

Optimisation and Performance Tuning

  • Introduction to Spark internals and execution plans.
  • Query optimization techniques and best practices.
  • Implementing partitioning strategies to improve performance.

User-Defined Functions and Advanced APIs

  • Creating and using user-defined functions (UDFs).
  • Vectorized UDFs for efficient data processing.

Streaming and Real-Time Processing

  • Introduction to Spark’s Structured Streaming API.
  • Executing and managing real-time streaming queries.

Delta Lake and Data Reliability

  • Understanding the advantages of Delta Lake.
  • Implementing Delta Lake for scalable and reliable data pipelines.

Exams and Assessments

This course does not include formal assessments.

Hands-On Learning

This course includes:

  • Practical exercises using Apache Spark on Databricks.
  • Hands-on labs to implement and optimise Spark queries.
  • Guided projects focusing on real-time data processing with Structured Streaming and Delta Lake.
Read more +

Databricks training partner

Maximize your data and AI potential with Databricks certified training. Bridge skills gaps across your organization to accelerate data-driven innovation, by enabling teams to scale insights and deploy AI for business growth.

Why choose QA

  • Expert instructors: Learn from experienced professionals with expertise in Apache Spark and Databricks.
  • Hands-on learning: Engage in practical exercises, labs, and real-world use cases.
  • End-to-end learning solutions: From foundational concepts to certification preparation, ensuring continuous professional development.
  • Recognized Databricks training partner: Trusted by leading enterprises for AI/ML upskilling.

Dates & Locations

Need to know

Frequently asked questions

How can I create an account on myQA.com?

There are a number of ways to create an account. If you are a self-funder, simply select the "Create account" option on the login page.

If you have been booked onto a course by your company, you will receive a confirmation email. From this email, select "Sign into myQA" and you will be taken to the "Create account" page. Complete all of the details and select "Create account".

If you have the booking number you can also go here and select the "I have a booking number" option. Enter the booking reference and your surname. If the details match, you will be taken to the "Create account" page from where you can enter your details and confirm your account.

Find more answers to frequently asked questions in our FAQs: Bookings & Cancellations page.

How do QA’s virtual classroom courses work?

Our virtual classroom courses allow you to access award-winning classroom training, without leaving your home or office. Our learning professionals are specially trained on how to interact with remote attendees and our remote labs ensure all participants can take part in hands-on exercises wherever they are.

We use the WebEx video conferencing platform by Cisco. Before you book, check that you meet the WebEx system requirements and run a test meeting to ensure the software is compatible with your firewall settings. If it doesn’t work, try adjusting your settings or contact your IT department about permitting the website.

How do QA’s online courses work?

QA online courses, also commonly known as distance learning courses or elearning courses, take the form of interactive software designed for individual learning, but you will also have access to full support from our subject-matter experts for the duration of your course. When you book a QA online learning course you will receive immediate access to it through our e-learning platform and you can start to learn straight away, from any compatible device. Access to the online learning platform is valid for one year from the booking date.

All courses are built around case studies and presented in an engaging format, which includes storytelling elements, video, audio and humour. Every case study is supported by sample documents and a collection of Knowledge Nuggets that provide more in-depth detail on the wider processes.

When will I receive my joining instructions?

Joining instructions for QA courses are sent two weeks prior to the course start date, or immediately if the booking is confirmed within this timeframe. For course bookings made via QA but delivered by a third-party supplier, joining instructions are sent to attendees prior to the training course, but timescales vary depending on each supplier’s terms. Read more FAQs.

When will I receive my certificate?

Certificates of Achievement are issued at the end the course, either as a hard copy or via email. Read more here.

Let's talk

By submitting this form, you agree to QA processing your data in accordance with our Privacy Policy and Terms & Conditions. You can unsubscribe at any time by clicking the link in our emails or contacting us directly.