Cost-Optimize a Large-Scale Spark Job for an Ad-Tech Platform
Overview
What this challenge is about.
Profile a PySpark job for an ad-tech platform, find its top 3 cost drivers, and prototype fixes to earn a verifiable certificate.
The scenario
The ad-tech platform (around 140 staff) processes around 12B impressions/day; a 40 percent reduction is approximately USD 45k annual savings on this single job and a template for 6 sibling jobs.
The Brief
What you'll do, and what you'll demonstrate.
Find and prove a 40 percent cost reduction on a 4TB nightly Spark job without breaking the 4-hour SLA.
Earning criteria — what you'll demonstrate
- Profile a real Spark job with the Spark UI and cloud-platform metrics
- Apply standard Spark optimizations (broadcast joins, partition tuning, instance mix)
- Build a defensible cost extrapolation from subset to full data
- Communicate cost trade-offs to finance + engineering stakeholders
Program Fit
Where this fits in your program.
Sharpens the same skills your degree expects you to demonstrate.
Cloud Computing for Data and ML
Master · Data Engineering
Strong alignment
This challenge maps to Cloud Computing for Data and ML at the Master level. It sharpens the same practical skills your coursework expects — but in a real industry context with actual constraints and deliverables.
Skills
Skills you'll demonstrate.
Each one shows up on your verified credential.
- Spark Optimization
Apply spark optimization to solve real industry problems and demonstrate production-level capability.
- Cloud Services
Apply cloud services to solve real industry problems and demonstrate production-level capability.
- Cost Engineering
Apply cost engineering to solve real industry problems and demonstrate production-level capability.
- Profiling
Apply profiling to solve real industry problems and demonstrate production-level capability.
- Etl Pipelines
Apply etl pipelines to solve real industry problems and demonstrate production-level capability.
- Benchmarking
Apply benchmarking to solve real industry problems and demonstrate production-level capability.
Careers
Career paths this challenge builds toward
Completing this challenge demonstrates skills that transfer directly to these roles:
Data Engineer
Spark cost optimization on a real EMR workload is the kind of project a data engineer ships in the first quarter at any ad-tech or large-data company.
This challenge sharpens
- spark-optimization
- cost-engineering
- etl-pipelines
MLOps Engineer
Profiling and cost-optimizing large compute workloads is the same skillset MLOps engineers use to tame training-cluster bills.
This challenge sharpens
- profiling
- cloud-services
- benchmarking
AI Solutions Architect
Translating profiling + optimization into a finance-team-defensible recommendation is core AI solutions architect work.
This challenge sharpens
- cost-engineering
- cloud-services
- spark-optimization