Star schema explained from memory. SQL drill set completed under 5 minutes.
Skewed-join story told end to end with real before/after numbers.
Every lab: README, Guided → Applied → Stretch exercises, self-check, answer key, honest time estimate.
| Module | What you will do | Time |
|---|---|---|
| A0 Python for Data Engineering | Comprehensions, dicts, pathlib, error handling | 3–4 h |
| A1 SQL Mastery | Window functions, CTEs, ranking, aggregation pitfalls | 2–3 h |
| A2 Dimensional Modelling | Star schema, grain, SCD Type 2, additivity bug | 2–3 h |
| A3 Why Distributed? | Pandas vs Spark crossover; four reasons NOT to use Spark | 1–2 h |
| A4 Git and Shipping | Branching, PRs, conflict resolution, secrets hygiene | 2–3 h |
| A5 Cloud Storage and Warehouses | S3/ADLS semantics, object store vs RDBMS, warehouse options | 2–3 h |
| Module | What you will do | Time |
|---|---|---|
| B1 Spark Architecture | Driver, executor, task hierarchy; Spark UI tour | 60–75 min |
| B2 DataFrames API | Schema, transformations, clean_loans pipeline | 60–75 min |
| B3 Aggregations and Joins | groupBy, window, join types, broadcast hint | 60–75 min |
| B4 Window Functions | Rank, lead/lag, forward-fill as window | 45–60 min |
| B5 File Formats and I/O | Parquet vs CSV: size, read time, schema evolution | 45–60 min |
| B6 Spark UI and explain() | Query plan reading, stage/task analysis | 60–75 min |
| B7 Shuffles | Wide vs narrow transforms, shuffle cost measurement | 45–60 min |
| B8 Joins, Skew, AQE | Skewed join: create → diagnose → fix with salting | 60–75 min |
| B9 Caching and OOM | Deliberate Java heap OOM, cache cost/benefit | 45–60 min |
| B10 Spark SQL and Catalog | SQL vs DataFrame API equivalence | 45–60 min |
| B11 UDFs | Write a UDF, then replace it with a built-in | 45–60 min |
| B12 Delta Lake | ACID, time travel, MERGE, compaction | 60–75 min |
| B13 Structured Streaming | File source, checkpointing, watermarks, foreachBatch | 90–120 min |
| B14 Testing PySpark | pytest fixtures, parametrize, CI-ready | 90–120 min |
| B15 Production PySpark | Config externalisation, logging, idempotent overwrite | 90–120 min |
| B16 Data Quality and Observability | Z-score anomaly, quarantine, row-count expectations | 60–75 min |
| B17 GenAI-Aware Pipelines | Chunking, embeddings, retrieval, agentic tie-in | 75–90 min |
| Module | What you will do | Time |
|---|---|---|
| C1 What dbt Is | DAG concept, compile vs run, warehouse-agnostic | 30 min |
| C2 Models, ref, DAG | source(), ref(), build order, additivity fix | 60–75 min |
| C3 Materialisations | Table vs view vs incremental — choosing with a reason | 45–60 min |
| C4 Project Structure | staging / intermediate / marts layering | 45–60 min |
| C5 Testing in dbt | not_null, unique, accepted-values, custom tests | 45–60 min |
| C6 Docs and Lineage | dbt docs generate, lineage graph, exposures | 45–60 min |
| C7 Jinja and Macros | Macro authoring, safe_divide, risk_tier reuse | 45–60 min |
| C8 Incremental Models | Backfill, late-arriving data, on_schema_change | 60–75 min |
| C9 Snapshots and SCD2 | check strategy, unique_key, tier-flip capture | 45–60 min |
| C10 Seeds, Hooks, Ops | Seeds, post-hook audit, run-operation | 30–45 min |
| C11 Environments | profiles.yml, dev vs prod targets, var() | 30–45 min |
| C12 CI/CD and Slim CI | GitHub Actions, state:modified+ selection | 45–60 min |
| C13 Modern dbt | Contracts, unit tests, MetricFlow awareness | 45–60 min |
| C14 Performance and Cost | Materialisation choice, partition pruning, tuning | 45–60 min |
| C15 dbt with Airflow | BashOperator, DbtTaskGroup, dependency ordering | 60–75 min |
| C16 Orchestration Fundamentals | Scheduling, retries, idempotency, backfill | 45–60 min |
| Module | What you will do | Time |
|---|---|---|
| D1 Online Assessments | 9 SQL patterns: median, pct-of-total, recursive CTE | 60+ min habit |
| D2 System Design | 7-step framework: requirements → trade-offs → weakness | 60–90 min |
| D3 Troubleshooting | 5 broken pipelines to diagnose + post-mortem | 90–120 min |
| D4 Take-Home Practice | Full 4-hour timed submission, auto-checked | 4–5 h |
| D5 Question Bank | 150+ interview questions mapped to every module | Ongoing |
| D6 Cheat Sheets | Quick-reference: Spark UI, dbt CLI, Git flow | Reference |
| Module | What you will do | Time |
|---|---|---|
| E1 Production-Grade BFSI Platform | MEASUREMENTS.md + INCIDENTS.md + 3-layer dbt project + pytest suite + screen-share walkthrough | Ongoing |
The same five small tables flow through every lab — SQL, Spark, dbt, and production, all on one schema.
MEASUREMENTS.md — six before/after tables with real numbers: Parquet vs CSV, shuffle tuning, skew fix, broadcast vs sort-merge, dbt materialisation cost, partition pruning.
INCIDENTS.md — eight deliberate failures you caused and diagnosed: OOM, row multiplication, silent null loss, late-arriving data, and more.
transform_spark.py — a production PySpark job: config externalised, logging in place, idempotent overwrite proven.
23+ pytest tests, CI-ready, across Guided/Applied tiers.
A three-layer dbt project — staging / intermediate / marts — with snapshots, exposures, and a proven CI slim-select.
A portfolio repo — architecture diagram, a 12-week DECISIONS.md interview story bank, and an honest README.
A RAG mini-pipeline — chunk / embed / retrieve / answer over real BFSI policy notices, with real GenAI/agentic vocabulary you can name cold.
Comment or message us now — we'll notify you the moment enrollment opens, with early-bird pricing for the first cohort.
Get Early Access on WhatsApp →