2 · DataFrames · lesson 6 of 20

Schemas & Types

Inferred vs explicit schemas, and why explicit almost always wins.

Python
from pyspark.sql.types import (
    StructType, StructField, IntegerType, StringType, DoubleType, TimestampType,
)

schema = StructType([
    StructField("id",        IntegerType(),   nullable=False),
    StructField("name",      StringType(),    nullable=True),
    StructField("amount",    DoubleType(),    nullable=True),
    StructField("created_at", TimestampType(), nullable=True),
])

df = spark.read.schema(schema).csv("/data/orders.csv", header=True)
df.printSchema()
  • ▸Inferring the schema scans the file — slow and sometimes wrong.
  • ▸An explicit schema is faster and rejects bad rows early.
  • ▸Common types: IntegerType, LongType, DoubleType, StringType, BooleanType, TimestampType, DateType, ArrayType, MapType, StructType.
Key takeaways
  • ✓Explicit schemas skip a full-file scan and guarantee stable types.
  • ✓StructType composes nested structs; ArrayType and MapType handle collections.