Items related to The Modern PySpark & Databricks Data Engineer:...

The Modern PySpark & Databricks Data Engineer: Build Scalable Data Pipelines with Apache Spark, Delta Lake, Streaming, Workflows, and Unity Catalog - Softcover

RAJ, RISHU

 
9798193827795: The Modern PySpark & Databricks Data Engineer: Build Scalable Data Pipelines with Apache Spark, Delta Lake, Streaming, Workflows, and Unity Catalog

Synopsis

Build Real-World Data Engineering Skills with PySpark and Databricks

Modern data engineering is no longer just about writing SQL or moving files from one system to another. Today's data engineers need to build pipelines that are scalable, reliable, optimized, secure, observable, and ready for production.

PySpark & Databricks Data Engineering provides a practical, structured path from foundational data engineering concepts to advanced enterprise-level data platforms.

Starting with Python, Spark, and PySpark fundamentals, the book progressively takes you into advanced transformations, performance optimization, Delta Lake, incremental processing, CDC, Structured Streaming, Auto Loader, Databricks Workflows, Unity Catalog, security, governance, system design, and production troubleshooting.

Inside this book, you will learn:
  • Data engineering fundamentals and architecture
  • Python concepts required for data engineering
  • PySpark DataFrames and Spark SQL
  • Transformations, actions, joins, aggregations, and window functions
  • Complex and nested data processing
  • Spark partitioning and shuffle
  • Broadcast joins and data-skew handling
  • Spark performance optimization
  • Databricks Lakehouse architecture
  • Delta Lake and ACID transactions
  • MERGE, UPDATE, DELETE, and time travel
  • Incremental data processing
  • Change Data Capture (CDC)
  • Structured Streaming
  • Auto Loader
  • Checkpointing and watermarks
  • Databricks Workflows and orchestration
  • Data quality and production reliability
  • Unity Catalog
  • Security and access control
  • Data lineage and governance
  • Data products and enterprise architecture
  • Monitoring, observability, and cost optimization
  • CI/CD and testing
  • Disaster recovery concepts
  • End-to-end data engineering projects
  • Production troubleshooting
  • Databricks system-design interviews
  • PySpark coding challenges

"synopsis" may belong to another edition of this title.