1 · Foundations · lesson 2 of 20

Installing PySpark & the SparkSession

Your first program — and every piece of the setup explained.

Install PySpark with pip, then create a SparkSession — the single entry point to every Spark feature (DataFrames, SQL, streaming, ML).

shell
pip install pyspark
Python
from pyspark.sql import SparkSession

spark = (
    SparkSession.builder
        .appName("HelloSpark")
        .master("local[*]")    # use all local cores
        .getOrCreate()
)

print(spark.version)
spark.range(5).show()
  • ▸appName — label your app shows in the Spark UI.
  • ▸master — local[*] for laptop, or yarn / k8s:// / spark://host in a cluster.
  • ▸getOrCreate — reuse an existing session if one already exists.
  • ▸spark.range(5) — smallest DataFrame you can build, useful for demos.
NOTE
In notebooks like Databricks or Colab, a SparkSession named 'spark' is usually pre-created for you.
Key takeaways
  • ✓SparkSession is the single entry point to DataFrames, SQL, streaming and ML.
  • ✓'local[*]' runs Spark in one JVM using every CPU core.
  • ✓Always call getOrCreate() so you don't build two sessions by accident.