Data Engineer — Modelling & Pipeline Quality
Warehouse modelling, data-quality contracts and a query-performance investigation on a very large table.
- Modelling
- dbt
- SQL performance
- Data quality
- Duration
- 40 min
- Questions
- 4 questions
- Difficulty
Design a near-real-time pipeline, argue when streaming is the wrong answer, and make silent failures impossible.
1. Take me through how you'd model a warehouse layer for an e-commerce business whose analysts constantly ask new questions.
~3 minChecks modelling fundamentals and consumer empathy.
2. A daily pipeline silently produced 30% fewer rows for a week before anyone noticed. How do you prevent the next one?
~3 minData-quality and observability thinking.
3. When would you choose streaming over batch, and when is streaming the wrong answer?
~3 minTests honest trade-off reasoning instead of hype.
4. A key query on a 4-billion-row table takes nine minutes. What do you look at?
~3 minWarehouse performance depth.
Warm-up & context · 5 min
Interviewer intro, your background, how the session runs.
Core questions · 27 min
Role-specific questions with live follow-ups based on your answers.
Deep dive · 9 min
One topic explored to the edge of your experience.
Your questions & wrap · 5 min
Close the loop and hand over to feedback generation.
Warehouse modelling, data-quality contracts and a query-performance investigation on a very large table.
Design a multi-region delivery and rollback strategy, then defend the operational cost of your choices.
Decompose a product requirement into services, data ownership and contracts, then defend the boundaries you drew.