2 · DataFrames · lesson 6 of 20
Schemas & Types
Inferred vs explicit schemas, and why explicit almost always wins.
Python
from pyspark.sql.types import (
StructType, StructField, IntegerType, StringType, DoubleType, TimestampType,
)
schema = StructType([
StructField("id", IntegerType(), nullable=False),
StructField("name", StringType(), nullable=True),
StructField("amount", DoubleType(), nullable=True),
StructField("created_at", TimestampType(), nullable=True),
])
df = spark.read.schema(schema).csv("/data/orders.csv", header=True)
df.printSchema()- ▸Inferring the schema scans the file — slow and sometimes wrong.
- ▸An explicit schema is faster and rejects bad rows early.
- ▸Common types: IntegerType, LongType, DoubleType, StringType, BooleanType, TimestampType, DateType, ArrayType, MapType, StructType.
Key takeaways
- ✓Explicit schemas skip a full-file scan and guarantee stable types.
- ✓StructType composes nested structs; ArrayType and MapType handle collections.