Cloudadorn Academy · Podcast

Modern Data Engineering

How real systems move data from where it is created to where it is useful.

Hosted by Eve and Surya — original AI radio personas, not impersonations of real people.

12 episodes 2h 48m Season 1 Free

Where to listen

Every episode is free. Nothing plays on this site — pick a service:

Season 1 · 12 episodes · 2h 48m

  1. Episode 01 Runtime 13:17 AI-narrated

    What a Data Engineer Actually Does

    Your neighbour says she is a data engineer and then says pipeline three different ways. This episode is the sentence you can say back: data engineers copy data from where it is created to where it is useful, without breaking it, losing it, or leaking it.

    Eve and Surya start in a restaurant kitchen you can picture. The cashier is brilliant at one order at a time and terrible at counting three years of receipts in the lunch rush. That is why companies copy data out of transactional systems into analytical ones. You will hear OLTP and OLAP only after that picture lands, plus the stations on the line: sources, scheduled trucks or a live conveyor, the cheap fridge (the data lake), Spark, the warehouse pantry, dbt recipes, and the Airflow board. Amazon, Microsoft, and Google each get a door. Snowflake's virtual warehouse is not the pantry.

    Takeaway: the split is a rule of thumb, not a law of physics. The failure mode still decides it for most teams.

    Read the written companion

  2. Episode 02 Runtime 14:35 AI-narrated

    Late vs Wrong: When to Stream and When to Batch

    A concert lets out and last night's prices are already useless. Friday night, the same company pays drivers, and a messy number is a lawsuit. This episode teaches the fork that decides most architectures: what happens if the number is twenty minutes late, versus wrong by a little.

    Eve and Surya use Uber's public H3 work and a DoorDash-shaped Tuesday without inventing anyone's production diagram. You will learn why a grown-up platform runs a conveyor and a truck, how an ordered log is like a ship's log, what partitioning and replay actually do, and why a merge on a unique event ID turns duplicate delivery into a boring failure. Payroll does not belong on the live belt because it feels modern.

    Takeaway: technology preference is the last reason, not the first.

    Read the written companion

  3. Episode 03 Runtime 14:59 AI-narrated

    Netflix and Telecom: Same Data, Different Fear

    Two businesses drown in events. One sells attention. One sells minutes and a map of your Tuesday. Both use streams, lakes, and warehouses. They are terrified of different mistakes.

    From the public record, Eve and Surya walk Netflix-style playback: a progress mutter becomes Continue Watching, the same events land in the lake for slow questions, and Apache Iceberg is the gift Netflix gave the rest of us. Then the phone company: mediation, rating, revenue assurance, live tower health versus a bill that cannot be sloppy. Location data is not a marketing field. It is a life.

    Takeaway: ask what event is money, what error makes the newspaper, and what the regulator will want to see. Tools come fourth.

    Read the written companion

  4. Episode 04 Runtime 14:20 AI-narrated

    Hospital and Bank Data: Privacy, Fraud, Regulators

    There is no single patient file. The ER, the lab, the pharmacy, and the insurer each speak a dialect. At the card terminal, the bank has already pre-computed the clues, and an auditor later wants the movie of every no.

    Eve and Surya teach hospital interoperability (the older messaging standard and FHIR, said fire), why a master patient index must be conservative, and why analytics runs on de-identified data wherever it can. On the bank side: tens-of-milliseconds public talk, not an invented budget; train-serve skew in kitchen words; and the Basel Committee's principles, BCBS 239, which pushed big banks toward lineage, completeness, and named owners. Card numbers get tokenized at the door.

    Takeaway: same bricks as Netflix. Different fear.

    Read the written companion

  5. Episode 05 Runtime 13:36 AI-narrated

    The Data Stack Explained as a Kitchen Line

    Enough industries. Walk the line. What actually sits on the counter, and why the logo is furniture.

    Eve and Surya tour the stations: cashier-style sources, copy-by-following-the-diary (change data capture), the cheap object-store fridge, Spark as the big knife, the labelled warehouse pantry, dbt recipes in git, and Airflow as the board that yells. Each station gets an Amazon door, a Microsoft door, and a Google door, including Dataplex rather than the retired Data Catalog. Decoupled storage and compute is just this: the fridge and the cook are separate bills.

    Takeaway: brands rotate. The line does not.

    Read the written companion

  6. Episode 06 Runtime 14:14 AI-narrated

    Batch vs Streaming: Why Nightly Is Usually Right

    Every step toward real time is a new way to stay up at night. Buy it only when a business decision actually changes inside the window.

    Eve and Surya walk a Tuesday at 1 a.m.: files land, a sensor checks they exist, something washes the ugly stuff, the warehouse copies them in, dbt builds silver then gold, and tests shut the gold door if they fail. Then the live pipe: phones, a log, a schema-registry bouncer, a labelled junk drawer, a windowed count, and a checkpoint, which is just where the job wrote down how far it got. One checkout-dying minute versus a chart that can wait for coffee.

    Takeaway: truck, then shuttle, then belt. Streams will duplicate unless you paid for exactly-once, and often even then.

    Read the written companion

  7. Episode 07 Runtime 14:04 AI-narrated

    Lakehouse Explained: Iceberg, Contracts, and Mesh

    People say lakehouse like it is a smoothie. Here is what broke, and what the word actually means on air: lake prices, warehouse manners, one copy of the data instead of two.

    Eve and Surya teach open table formats (Iceberg, Delta, Hudi) as metadata brains on cheap files, why you can update and time-travel without rewriting the ocean, and why your petabytes should not live in one vendor's private drawer. Data contracts are a versioned promise that fails the producer's deploy, not the consumer's 3 a.m. Mesh is medicine you can steal (owners, service levels, a paved road) without renaming the company around a blog post. Lambda is two kitchens. Kappa says everything is a stream. Most grown-ups are hybrids.

    Takeaway: steal the medicine. Skip the religion.

    Read the written companion

  8. Episode 08 Runtime 13:42 AI-narrated

    Metrics, Feature Stores, and Sending Data Back

    Gold is not a mythic city. It is tables humans can think with, and anywhere a decision happens with a clean name.

    Eve and Surya draw a star schema without a whiteboard: tickets in the middle, customer, product, and date around them. When a customer moves cities, the trade calls that a slowly changing dimension. Finance, marketing, and the chief executive do not match because revenue is three different SQL queries; a metrics layer writes the formula once. Reverse ETL pushes gold back into Salesforce or the ward chart, with the same masking as any exit door. The feature-store Tuesday: swipes in the last five minutes must be the same number at 3 a.m. training and at the card terminal.

    Takeaway: if it is not named, it is not gold. It is leftover soup.

    Read the written companion

  9. Episode 09 Runtime 13:45 AI-narrated

    DataOps and FinOps: Tests, Alerts, and the Bill

    Great models and a Slack channel named why is Tuesday empty is not DataOps. That is hope, scheduled.

    Eve and Surya treat data like software: branches, pull requests, tests, and a staging warehouse that is not production with a funny hat. Airflow job graphs, the dags, live in the same git. Freshness is whether the truck arrived. Observability is whether the right cow arrived. Data-aware scheduling means the dashboard job waits until yesterday's orders exist and pass tests, not until seven o'clock. Zero-ETL is fine for a prototype. You still do not beat up the cashier computer for the annual report. Then the bill: tag everything, and kill idle compute, the Snowflake-style virtual warehouses, when nobody is querying.

    Takeaway: the platform that cannot show cost per team will be given a number by someone who does not like you.

    Read the written companion

  10. Episode 10 Runtime 13:55 AI-narrated

    Data Quality Gates: Bronze, Silver, and Gold

    Bad data is worse than no data. That is a poster. The engineering is gates. Food does not move to the next station until something signs.

    Eve and Surya teach bronze as the crate still sealed, silver as washed and labelled, gold as plated for the meeting. On the live belt, a schema registry stops bad events at the door and a dead-letter drawer holds the rest, with a label. A drawer that suddenly fills is often your first news of someone else's bad deploy. On the truck: schema, nulls, duplicates, volume versus last Mondays, freshness. Gold asks if the number is true. You may add the cashier computer for yesterday, cheaply, a sum, not three years.

    Takeaway: quality is not a cleanup crew at the end.

    Read the written companion

  11. Episode 11 Runtime 14:07 AI-narrated

    Data Governance: Owners, Catalogs, and Lineage

    Someone shipped a dashboard. Nobody knows who owns the table. Do not start with a tool. Start with a name. If nobody gets the page, you do not have governance. You have folklore.

    Eve and Surya teach the phone book (Glue Data Catalog, Microsoft Purview, Google Dataplex) and why it is useless if the owner field is empty. Lineage means walking a wrong chart back to the join that blew up. The hard security problem is people with legitimate access doing illegitimate curiosity. Master data is one customer, three spellings, and a steward. Retention is long enough for the lawyer and short enough for the privacy law. Time travel is cute until you cannot delete a person. A small council writes a few global rules. The platform enforces them in automated build checks, not in a meeting everyone skips.

    Monday: put an owner and a freshness promise on the ten tables the business actually uses.

    Read the written companion

  12. Episode 12 Runtime 13:16 AI-narrated

    How Generative AI Changes Data Engineering

    Does this job still exist, or do we all become prompt people? The job gets weirder, not smaller. Models eat data. They also emit data. Someone still has to move it without breaking it, losing it, or leaking it, including leaking it into a prompt.

    Eve and Surya recap the season, then get specific: drafting SQL and tests gets faster; judgment does not. Confident-wrong is the new bad data. A golden set is a folder of the ugliest cases you have ever seen, run on every change. Lineage now includes prompts, retrieved chunks, and tool calls. Secrets in a prompt are the new password checked into the code repo. The old gates still apply: mask first, and a human before money or medical.

    Takeaway: tell your neighbour you finally see the kitchen. The plate arriving on time is not an accident.

    Read the written companion

Masterclasses (read)

Prefer to read? Four long-form write-ups of how we build these systems.

In production Certification video tracks — not published. No launch date until a track is ready to watch. The list