The plate you see, the kitchen you do not
Eve's neighbor is a data engineer, and every explanation so far has leaned on the word pipeline. Surya is asked to explain the job without it.
He starts with an ordinary Tuesday. A finance person opens a three-year revenue report and the correct figure is already waiting. Nobody notices, because there is nothing to notice. That invisibility, he says, is the work: not producing the report, but making sure its number is right and on time.
The image that carries the series is a large restaurant chain. Diners see the finished plate, never the early delivery, the cold store, the hours of prep, or the schedule board that lands every dish hot at once. Data engineering is that unseen back half. The plate is a dashboard, a recommendation on a phone, or a fraud model judging a card purchase. When the kitchen slips, the diner gets a plate that is late, wrong, or contains something it never should have.
Why the till cannot answer the question
Eve raises the objection most people are too polite to voice: the company already holds this data, so why can't an analyst query it where it sits?
Surya's answer is about purpose. The till at the front records each sale immediately, correctly, and completely, or not at all. Ask it, mid lunch rush, to hand-count three years of receipts by region and cuisine, and the queue stops. This is mechanics, not metaphor: he cites Microsoft's published architecture guidance that aggregating across millions of transactions is heavy work for a transactional system, runs slowly, and can block the sales it exists to record.
Shape is a second issue. A transactional store splits records into small pieces to avoid duplication, so a large question means reassembling them with a query few people can write unassisted. The response is a copy: lift the data out of the system where it was born and put it somewhere built for big reads. A lighter first step, keeping only the current year in the operational system, buys time rather than a fix.
The skill is not copying. It is not breaking
Copying a file, Eve notes, is something a phone does on its own. Surya draws the line: moving the same collection of tables every night, from systems never built to cooperate, some of which change structure silently, and having the figure be correct before the first morning meeting. That is the craft.
Most of his week is not new construction. It is what happens when a nightly copy fails while nobody is awake. A strong team has a check, written months earlier, that quarantines the bad batch before anyone sees it. A weak team learns from an executive already looking at the wrong number. Nearly every practice discussed later in the series is the industry trying to be the first kind of team.
Two acronyms that only work as a pair
Eve permits one acronym; Surya takes two. Online Transaction Processing, OLTP, is the till side: recording business events as they happen in databases such as SQL Server or Postgres, the kind behind a phone app. Online Analytical Processing, OLAP, is the back office: a separate store tuned for heavy reading and light writing that deliberately keeps history so large queries never disturb the lunch rush.
The transactional systems hold the valuable facts but were not designed to be analyzed, so extracting from them is slow and awkward. Hence the second building. Surya casts himself as the truck between the two, plus the loading dock, plus the person confirming that what arrived matches the manifest. He is honest about freshness: the analytical copy is rarely live, refreshes on a slower rhythm, and needs deliberate cleaning and scheduling. The distance between what the till knows and what the report knows absorbs much of a data engineer's time.
The stations, picture first and name second
Ingredients originate on farms and fisheries: the source systems, meaning the order database, the customer database, the CRM. Delivery trucks on a fixed overnight route are scheduled bulk movement; the conveyor that never stops is the event-by-event alternative, best known as Apache Kafka, with Kinesis or managed Kafka on Amazon, Event Hubs on Azure, and Pub/Sub on Google Cloud.
The walk-in fridge is the data lake, cheap and mostly raw: S3, Blob Storage, and Cloud Storage are three logos on one door. The prep station is Apache Spark, run on all three clouds. The labeled pantry is the warehouse: Redshift, BigQuery, Synapse and now Fabric, or Snowflake on any of them. Recipes are dbt, plain SQL select statements that gain version control, tests, and documentation. The schedule board is Apache Airflow, with jobs written in Python. The plate is dashboards, reports, and models.
One word gets special care. In the kitchen picture, warehouse is the pantry. Inside Snowflake, a virtual warehouse is a compute cluster, which is the prep station instead. Same word, two rooms.
One rule to soften, one sentence to keep
Asked what he would take back, Surya softens a slogan. Keeping analytics off the operational database is a guideline, not a physical law; a named category, hybrid transactional and analytical processing, exists for the overlap, and in places the two buildings already blur. Yet for most teams the failure mode still decides: get the split wrong and the person who suffers is the customer waiting at the till.
Surya: “I move data from where it's created to where it's useful — without breaking it, losing it, or leaking it.”
Each verb has a face. Broken is a wrong figure in the board pack. Lost is an order that never lands anywhere. Leaked is the story in the news with the company's name on it. Everything discussed later in the series serves that one copy job. Next: two companies, one fork.