Learn Apache Spark and Databricks by seeing what happens under the hood.
Free Apache Spark tutorial and Databricks tutorial. Learn PySpark, Spark architecture, shuffle, joins, Catalyst, Delta Lake, and Unity Catalog through a simulator you can step through — then practice interview questions.
Your progress is saved automatically.
What you will actually see
Educational simulation · not live cluster data
Now showing: Driver
How it works
Visualization first, text second — so the mental model comes before the jargon.
Step 1
Watch it execute
Step through Driver → Job → Stages → Tasks → Shuffle at your own pace. Every step says what Spark is doing and why.
Step 2
Inspect any piece
Tap a task to see its partition, executor, and host. The three ideas people most often confuse, kept clearly separate.
Step 3
Then go deeper
Switch to Learn or Interview when you want internals, production failure cases, and the questions you will be asked.
Partitions, tasks, executors
See why one partition means one task, and how executors pick that work up.
What groupBy really costs
Watch records fly across the cluster so each key lands on one reducer.
Simulation
41 interactive Apache Spark and Databricks sims
- Spark
Spark Execution Simulator
Five labs from lazy code → plan → job → stages → tasks → shuffle → result.
- Spark
Cluster Architecture
Driver, cluster manager, and executors — who does what.
- Spark
Lazy Evaluation
Watch transformations pile up until an action finally runs.
- Spark
Partitions & Tasks
How a table splits into partitions and one task per partition.
- Spark
Shuffle Simulator
Watch keys hash to reducers across the network.
- Spark
Join Simulator
Broadcast vs shuffle join — when the big table should stay put.
What happens when you run a Spark job?
Five chained labs — from your code to the shuffle result.
Apache Spark and Databricks curriculum
43 lessons · 41 simulations · Spark and Databricks interview questions
- Fundamentals
Spark Overview
What Spark is, why it exists, and the mental model of driver, executors, and lazy plans.
- Fundamentals
Spark Architecture
Cluster manager, driver JVM, executor JVMs, and how a SparkSession talks to a cluster.
- Fundamentals
Driver & Executors
The driver plans. Executors run tasks. Never ship a huge collect() back to the driver.
- Fundamentals
Execution Flow
Trace groupBy + show() from lazy code through the plan, job, stages, tasks, shuffle, and result — one picture.
- Fundamentals
DataFrame
Structured data with a schema. Catalyst can optimize DataFrames in ways RDDs cannot.
- Fundamentals
Spark SQL
SQL strings and the DataFrame API compile to the same plan. Temp views, the catalog, and when names actually resolve.
- Fundamentals
Transformations
Narrow vs wide transformations, and why nothing runs until an action arrives.
- Fundamentals
Actions
show, count, collect, write — the calls that force Spark to produce a result.
- Fundamentals
Lazy Evaluation
Spark builds a DAG of transformations and waits. That wait is a feature, not a bug.
