Projify
Home Courses Project Catalog Get Early Access
Coming Soon — Now Enrolling

Data Engineer Bootcamp

SQL → PySpark → dbt → Production. One real BFSI dataset flows through 43 hands-on labs, ending in a working pipeline and a GitHub repo you can screen-share and defend line by line.

12 Weeks + Capstone 43 Labs Intermediate+ Certificate
Get Early Access on WhatsApp → See the Full Curriculum ↓
Duration
12 weeks + capstone
Domain
BFSI — district loan & NPA data (India)
Core Stack
Python · PySpark · dbt-core + DuckDB · SQL · Git
Format
Self-paced + weekly 1:1 tutor sessions
Lab Runtime
Colab (PySpark) · Browser HTML (SQL) · Terminal (dbt, Git)
Prerequisites
Basic Python & SQL — no cloud/Java on day 1

Two Checkpoints Gate Your Progress

Checkpoint 1 — End of Week 1

Star schema explained from memory. SQL drill set completed under 5 minutes.

Checkpoint 2 — End of Week 8

Skewed-join story told end to end with real before/after numbers.

Full Curriculum — 43 Labs

Every lab: README, Guided → Applied → Stretch exercises, self-check, answer key, honest time estimate.

Part A — Foundations · Weeks 0–1

ModuleWhat you will doTime
A0 Python for Data EngineeringComprehensions, dicts, pathlib, error handling3–4 h
A1 SQL MasteryWindow functions, CTEs, ranking, aggregation pitfalls2–3 h
A2 Dimensional ModellingStar schema, grain, SCD Type 2, additivity bug2–3 h
A3 Why Distributed?Pandas vs Spark crossover; four reasons NOT to use Spark1–2 h
A4 Git and ShippingBranching, PRs, conflict resolution, secrets hygiene2–3 h
A5 Cloud Storage and WarehousesS3/ADLS semantics, object store vs RDBMS, warehouse options2–3 h

Part B — PySpark · Weeks 2–8

ModuleWhat you will doTime
B1 Spark ArchitectureDriver, executor, task hierarchy; Spark UI tour60–75 min
B2 DataFrames APISchema, transformations, clean_loans pipeline60–75 min
B3 Aggregations and JoinsgroupBy, window, join types, broadcast hint60–75 min
B4 Window FunctionsRank, lead/lag, forward-fill as window45–60 min
B5 File Formats and I/OParquet vs CSV: size, read time, schema evolution45–60 min
B6 Spark UI and explain()Query plan reading, stage/task analysis60–75 min
B7 ShufflesWide vs narrow transforms, shuffle cost measurement45–60 min
B8 Joins, Skew, AQESkewed join: create → diagnose → fix with salting60–75 min
B9 Caching and OOMDeliberate Java heap OOM, cache cost/benefit45–60 min
B10 Spark SQL and CatalogSQL vs DataFrame API equivalence45–60 min
B11 UDFsWrite a UDF, then replace it with a built-in45–60 min
B12 Delta LakeACID, time travel, MERGE, compaction60–75 min
B13 Structured StreamingFile source, checkpointing, watermarks, foreachBatch90–120 min
B14 Testing PySparkpytest fixtures, parametrize, CI-ready90–120 min
B15 Production PySparkConfig externalisation, logging, idempotent overwrite90–120 min
B16 Data Quality and ObservabilityZ-score anomaly, quarantine, row-count expectations60–75 min
B17 GenAI-Aware PipelinesChunking, embeddings, retrieval, agentic tie-in75–90 min

Part C — dbt · Weeks 9–11

ModuleWhat you will doTime
C1 What dbt IsDAG concept, compile vs run, warehouse-agnostic30 min
C2 Models, ref, DAGsource(), ref(), build order, additivity fix60–75 min
C3 MaterialisationsTable vs view vs incremental — choosing with a reason45–60 min
C4 Project Structurestaging / intermediate / marts layering45–60 min
C5 Testing in dbtnot_null, unique, accepted-values, custom tests45–60 min
C6 Docs and Lineagedbt docs generate, lineage graph, exposures45–60 min
C7 Jinja and MacrosMacro authoring, safe_divide, risk_tier reuse45–60 min
C8 Incremental ModelsBackfill, late-arriving data, on_schema_change60–75 min
C9 Snapshots and SCD2check strategy, unique_key, tier-flip capture45–60 min
C10 Seeds, Hooks, OpsSeeds, post-hook audit, run-operation30–45 min
C11 Environmentsprofiles.yml, dev vs prod targets, var()30–45 min
C12 CI/CD and Slim CIGitHub Actions, state:modified+ selection45–60 min
C13 Modern dbtContracts, unit tests, MetricFlow awareness45–60 min
C14 Performance and CostMaterialisation choice, partition pruning, tuning45–60 min
C15 dbt with AirflowBashOperator, DbtTaskGroup, dependency ordering60–75 min
C16 Orchestration FundamentalsScheduling, retries, idempotency, backfill45–60 min

Part D — Interview Machinery · Week 12

ModuleWhat you will doTime
D1 Online Assessments9 SQL patterns: median, pct-of-total, recursive CTE60+ min habit
D2 System Design7-step framework: requirements → trade-offs → weakness60–90 min
D3 Troubleshooting5 broken pipelines to diagnose + post-mortem90–120 min
D4 Take-Home PracticeFull 4-hour timed submission, auto-checked4–5 h
D5 Question Bank150+ interview questions mapped to every moduleOngoing
D6 Cheat SheetsQuick-reference: Spark UI, dbt CLI, Git flowReference

Part E — Capstone · Week 12+

ModuleWhat you will doTime
E1 Production-Grade BFSI PlatformMEASUREMENTS.md + INCIDENTS.md + 3-layer dbt project + pytest suite + screen-share walkthroughOngoing

One Schema, Twelve Weeks

The same five small tables flow through every lab — SQL, Spark, dbt, and production, all on one schema.

FileSizeContents
district_loans.csv96 rows8 districts × 12 months of loan and NPA data for 2024
district_master.csv11 rowsSCD Type 2 source — Delhi, Pune, Lucknow change risk tier mid-year
rbi_bank_groups.csv20 rowsGross NPA % by RBI bank group and year
macro_indicators.csv5 rowsGDP, inflation, repo rate by year
district_loans_dirty.csv29 rowsNulls, duplicates, wrong month, negative amount — used in B16 and C5

What You Leave With

MEASUREMENTS.md — six before/after tables with real numbers: Parquet vs CSV, shuffle tuning, skew fix, broadcast vs sort-merge, dbt materialisation cost, partition pruning.

INCIDENTS.md — eight deliberate failures you caused and diagnosed: OOM, row multiplication, silent null loss, late-arriving data, and more.

transform_spark.py — a production PySpark job: config externalised, logging in place, idempotent overwrite proven.

23+ pytest tests, CI-ready, across Guided/Applied tiers.

A three-layer dbt project — staging / intermediate / marts — with snapshots, exposures, and a proven CI slim-select.

A portfolio repo — architecture diagram, a 12-week DECISIONS.md interview story bank, and an honest README.

A RAG mini-pipeline — chunk / embed / retrieve / answer over real BFSI policy notices, with real GenAI/agentic vocabulary you can name cold.

Ready to start building?

Comment or message us now — we'll notify you the moment enrollment opens, with early-bird pricing for the first cohort.

Get Early Access on WhatsApp →