Building reliable ETL pipelines with built-in observability - Data Engineering with Databricks

Databricks · Advanced ·🔄 Data Engineering ·11mo ago

Key Takeaways

This video demonstrates how to build reliable ETL pipelines with native observability using Databricks LakeFlow, a unified intelligence solution for data engineering. It showcases how to monitor, debug, and optimize data pipelines, and highlights the importance of observability in ensuring trustworthy downstream analytics and machine learning.

Full Transcript

[Music] Hi, I'm Teresa. I'm on the product team here at Data Bricks and welcome to another installment of data engineering with data bricks. I'm going to discuss and demo how to build efficient and reliable data pipelines with native observability in LakeFlow. As a data engineer, you better be a responsibility of ensuring that your data pipelines are healthy, reliable, and efficient so that you can enable trustworthy downstream analytics and machine learning. This means that you're also accountable for monitoring their performance and resolving any issues should they arise. However, monitoring and keeping everything afloat isn't easy. In fact, most data engineers have the following challenges. Firstly, complex data pipelines, which makes it incredibly difficult to effectively manage and quickly access related jobs and pipelines. As a result, data engineers are not able to properly view the workloads and keep an eye on their performance and making sure everything is running as it should. Secondly, we have error detection and root cause analysis. When issues arise, it can be difficult to identify that there is a problem, pinpoint the source and also assess its level of severity. This can lead to ineffective monitoring as well as time consuming root cause analysis and troubleshooting. Finally, we have optimization at monitoring at scale. With more data comes more complexity. Ensuring that your data pipelines are consistently efficient and running well becomes more challenging at scale. Without the right observability capabilities, you risk impacting the quality of your data and your costs. These challenges can be resolved with observability natively integrated in your data engineering tool. something that data bricks can solve today with leg flow. Data bricks lakeflow is a unified intelligence solution for data engineering with streamlined ETL development and operations built on the data intelligence platform. With LegFlow, data engineers can enjoy a rich set of integrated observability features across data ingestion, transformation, and orchestration. This allows them to diagnose bottlenecks, optimize performance, and better manage resource usage and costs in one single place. When you have access to endtoend monitoring, you stay in control of your data and your pipelines. Observability in native lakeflow consists of different use cases. To start off, we have discoverability. When we talk about discoverability, we mean making it really easy to organize and access your jobs and pipelines. This is important for proactively monitoring your workloads, managing them, and checking on their status. This can be especially key when you're developing new pipelines or are productionizing them. Then we have monitoring and alerts. Once your workloads are up and running, it is common to rely on a more reactive approach of monitoring. Regardless of your entry point, this step will help you inform the decision if you need to drill deeper. If you've identified a problem and really need to understand what's happening, we have root cause analysis. This is really about understanding why something went wrong. Whether this is about debugging an error or resolving performance bottlenecks. This will equip you with the right insights to move on to debugging and optimization. Let's take a look at what this looks like in practice. Let's assume I'm a data engineer and I'm working on data pipelines to ingest and transform sales data for my business. To achieve this, I use LakeFlow Connect to ingest my data from Salesforce and other sources. Lakefully declarative pipelines to aggregate the data and perform data quality checks and lakeflow jobs to orchestrate to make sure the data is always fresh and reliable and business user can rely on the dashboards and reports we produce. Now I want to proactively monitor my project because we've made some recent changes. In order to do this jobs and pipelines is my central entry point from here I can filter for my tag. My project is called SE sales data and this will allow me to see all of the different related jobs and pipelines that are part of my project at one glance as well as key information about them and the recent runs. Here I can immediately see that my sales data processing job had an issue. So I'm going to jump into this and see what went wrong. In order to do this I'm going to jump into the latest failed job and go to the timeline view. I really like to use this view because it adds a dimension of duration to my regular DAG view. This helps me to both immediately spot when something went wrong as well as identifying performance bottlenecks. In my case, I can see my aggregate sales data pipeline went wrong. So, I'm going to click into this to understand what happened here. Looking into my pipeline that has failed, I can immediately see that there was an issue with a data quality check failing. If I want to understand this in more detail, I can look at the DAG view to understand which table was impacted. And if I wanted to directly jump into the code and fix the problem, let's assume I've now fixed a problem and I want to make sure I can automate this going forward so I don't have to proactively keep checking on my jobs and pipelines. I can do this by setting up a notification. To do that, I'm going to jump back into my job and I'm going to find the notification section. Once I have it, I'm going to click add notification. And in this case, failure is already selected. That's exactly what I want to be notified about. Now, the only thing I need to do to complete it is to select a destination here. I can select from an email address or a system destination. I'm going to select my email address here. And once I'm happy, I'm going to save this. Perfect. What this means is that now I don't have to proactively check on this job, but I will be automatically notified if there's a problem. So this was just one example of showcasing an end toend flow made possible with LakeFlow observability. But there's even more. So we've just seen an example of setting up an alert for a failure, but I could also set up more sophisticated alerts for duration thresholds or streaming backlog. Other than that, I can also build dashboards to visualize observability at scale across my data bricks jobs and pipelines and across workspaces. Finally, LakeFlow allows me to track data lineage whether upstream or downstream with Unity catalog, data bricks unified governance solution. And that's about it. I hope this video helped you learn about observability in LakeFlow. Hear about the latest tools we offer for pipeline monitoring, debugging, and optimization. and see how easy it is to use Lake to help you solve your toughest data engineering challenges. For more information, check out our website or dive into the technical documentation which includes demos and other resources. In the meantime, thank you for watching data engineering with data bricks.

Original Description

As a data engineer, you bear the heavy responsibility of ensuring that the data pipelines you are building are healthy, reliable and efficient, enabling trustworthy downstream analytics and ML. With Databricks Lakeflow, data engineers can enjoy a rich set of integrated observability features across data ingestion, transformation and orchestration so they can diagnose bottlenecks, optimize performance, and better manage resource usage and costs, in one single place. When you have access to end-to-end monitoring, you stay in control of your data and your pipelines. Watch this video to learn about observability in Lakeflow Jobs and the latest tools we offer for pipeline monitoring, debugging and optimization. Additional resources: * Watch 2025 Data + AI Summit session, Lakeflow Observability: From UI Monitoring to Deep Analytics - https://youtu.be/rbzQYdRhbOo?feature=shared * Learn more about Lakeflow Jobs -https://www.databricks.com/product/data-engineering/lakeflow-jobs * Blog: What’s New: Lakeflow Jobs Provides More Efficient Data Orchestration - https://www.databricks.com/blog/whats-new-lakeflow-jobs-provides-more-efficient-data-orchestration
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Databricks · Databricks · 0 of 60

← Previous Next →
1 Building AI Agent Systems with Databricks
Building AI Agent Systems with Databricks
Databricks
2 Databricks Workflows
Databricks Workflows
Databricks
3 Automate Unity Catalog Upgrade with UCX Part 1: Overview
Automate Unity Catalog Upgrade with UCX Part 1: Overview
Databricks
4 Automate Unity Catalog Upgrade with UCX Part 2: Installation
Automate Unity Catalog Upgrade with UCX Part 2: Installation
Databricks
5 Automate Unity Catalog Upgrade with UCX Part 3 - Assessment
Automate Unity Catalog Upgrade with UCX Part 3 - Assessment
Databricks
6 Automate Unity Catalog Upgrade with UCX  Part 4 - Group Migration
Automate Unity Catalog Upgrade with UCX Part 4 - Group Migration
Databricks
7 Table Migration and Catalog Design with UCX | Part 5
Table Migration and Catalog Design with UCX | Part 5
Databricks
8 Setting Up Azure Access for UCX Table Migration | Part 6
Setting Up Azure Access for UCX Table Migration | Part 6
Databricks
9 UCX Table Migration: Creating Catalogs and Schemas | Part 7
UCX Table Migration: Creating Catalogs and Schemas | Part 7
Databricks
10 Automate Unity Catalog Upgrade with UCX  Part 8: Code Migration
Automate Unity Catalog Upgrade with UCX Part 8: Code Migration
Databricks
11 Streaming to Kafka Just Got Easier with DLT Pipelines
Streaming to Kafka Just Got Easier with DLT Pipelines
Databricks
12 Data Engineering From Data to Dashboards with DABs: Crunching the Cookies Dataset
Data Engineering From Data to Dashboards with DABs: Crunching the Cookies Dataset
Databricks
13 Epsilon helps businesses connect with their consumers using Databricks Data Intelligence Platform
Epsilon helps businesses connect with their consumers using Databricks Data Intelligence Platform
Databricks
14 Unilever transforms operations with GenAI using the Databricks Data Intelligence Platform
Unilever transforms operations with GenAI using the Databricks Data Intelligence Platform
Databricks
15 ActionIQ enables businesses to unlock customer data with the Databricks Data Intelligence Platform
ActionIQ enables businesses to unlock customer data with the Databricks Data Intelligence Platform
Databricks
16 Mixed Attention & LLM Context | Data Brew | Episode 35
Mixed Attention & LLM Context | Data Brew | Episode 35
Databricks
17 Inside Databricks SQL: Engineering innovation with Hans
Inside Databricks SQL: Engineering innovation with Hans
Databricks
18 Inside Databricks: Engineering innovation with Michael Armbrust
Inside Databricks: Engineering innovation with Michael Armbrust
Databricks
19 The Money Team at Databricks: driving revenue and customer growth
The Money Team at Databricks: driving revenue and customer growth
Databricks
20 Unity Catalog unveiled: engineering data governance at scale
Unity Catalog unveiled: engineering data governance at scale
Databricks
21 Create a view in Databricks and share it with Power BI using Delta Sharing
Create a view in Databricks and share it with Power BI using Delta Sharing
Databricks
22 NDUS leverages Databricks Data Intelligence Platform to revolutionize higher education management
NDUS leverages Databricks Data Intelligence Platform to revolutionize higher education management
Databricks
23 Démo Databricks de AI/BI
Démo Databricks de AI/BI
Databricks
24 EMEA Data + AI World Tour 2024
EMEA Data + AI World Tour 2024
Databricks
25 GenAI: The Shift to Data Intelligence - Customer Panel on Industry Use Cases
GenAI: The Shift to Data Intelligence - Customer Panel on Industry Use Cases
Databricks
26 GenAI: The Shift to Data Intelligence - Ft. Ash Jhaveri, VP of Reality Labs Partnerships at Meta
GenAI: The Shift to Data Intelligence - Ft. Ash Jhaveri, VP of Reality Labs Partnerships at Meta
Databricks
27 Virtue Foundation leverages the Databricks Data Intelligence Platform to advance global health
Virtue Foundation leverages the Databricks Data Intelligence Platform to advance global health
Databricks
28 Announcing Synthetic Data Generation in Mosaic AI Agent Evaluation
Announcing Synthetic Data Generation in Mosaic AI Agent Evaluation
Databricks
29 AI/BI Dashboards Embedding - A tutorial
AI/BI Dashboards Embedding - A tutorial
Databricks
30 Bayer transforms global data management with the Databricks Data Intelligence Platform
Bayer transforms global data management with the Databricks Data Intelligence Platform
Databricks
31 Databricks at AWS re:Invent 2024
Databricks at AWS re:Invent 2024
Databricks
32 Hive Metastore and AWS Glue Federation in Unity Catalog
Hive Metastore and AWS Glue Federation in Unity Catalog
Databricks
33 Data + AI World Tour Paris 2024
Data + AI World Tour Paris 2024
Databricks
34 Retail reimagined: Currys data-first strategy to driving growth and improving operations
Retail reimagined: Currys data-first strategy to driving growth and improving operations
Databricks
35 Mixture of Memory Experts (MoME) | Data Brew | Episode 36
Mixture of Memory Experts (MoME) | Data Brew | Episode 36
Databricks
36 Verana Health Data Curation and Innovation with Databricks and AWS
Verana Health Data Curation and Innovation with Databricks and AWS
Databricks
37 Securing SaaS Applications: Obsidian Security on Their Journey with Databricks and AWS
Securing SaaS Applications: Obsidian Security on Their Journey with Databricks and AWS
Databricks
38 Twilio Eng VP on Data Intelligence & AI at AWS re:Invent 2024
Twilio Eng VP on Data Intelligence & AI at AWS re:Invent 2024
Databricks
39 Chegg Eng SVP on Data-Driven Approach to Student Success with Databricks and AWS
Chegg Eng SVP on Data-Driven Approach to Student Success with Databricks and AWS
Databricks
40 Ibotta Personalized Rewards Innovation with Databricks and AWS
Ibotta Personalized Rewards Innovation with Databricks and AWS
Databricks
41 Simplify AI governance with #databricks AI Gateway
Simplify AI governance with #databricks AI Gateway
Databricks
42 Databricks SQL and Power BI Integration
Databricks SQL and Power BI Integration
Databricks
43 Databricks Serverless SQL Warehouses
Databricks Serverless SQL Warehouses
Databricks
44 7 West powers audience growth with the Databricks Data Intelligence Platform
7 West powers audience growth with the Databricks Data Intelligence Platform
Databricks
45 Secret to Production AI: Tools & Infrastructure | Data Brew | Episode 37
Secret to Production AI: Tools & Infrastructure | Data Brew | Episode 37
Databricks
46 Skyflow CEO on Data Privacy with Databricks at AWS re:Invent
Skyflow CEO on Data Privacy with Databricks at AWS re:Invent
Databricks
47 Databricks Clean Rooms Product Demo
Databricks Clean Rooms Product Demo
Databricks
48 Dun & Bradstreet Enrichment & Monitoring, powered by Delta Sharing & Databricks Marketplace
Dun & Bradstreet Enrichment & Monitoring, powered by Delta Sharing & Databricks Marketplace
Databricks
49 Unpacking Libraries in Databricks
Unpacking Libraries in Databricks
Databricks
50 Providence uses an AI agent system from Databricks to help doctors improve their communication
Providence uses an AI agent system from Databricks to help doctors improve their communication
Databricks
51 How State Street Uses AI to Transform Millions of Trades Daily
How State Street Uses AI to Transform Millions of Trades Daily
Databricks
52 Vevo Therapeutics CEO on Curing Disease with Data at AWS re:Invent
Vevo Therapeutics CEO on Curing Disease with Data at AWS re:Invent
Databricks
53 Over Architected with Nick & Holly: Databricks updates for Feb 2025
Over Architected with Nick & Holly: Databricks updates for Feb 2025
Databricks
54 The Power of Synthetic Data | Data Brew | Episode 38
The Power of Synthetic Data | Data Brew | Episode 38
Databricks
55 Use Databricks Lakehouse Federation to break down data silos
Use Databricks Lakehouse Federation to break down data silos
Databricks
56 AI's rugby score: National Rugby League rallies fans with analytics and unified data
AI's rugby score: National Rugby League rallies fans with analytics and unified data
Databricks
57 Open Variant Data Type in Delta Lake and Apache Spark
Open Variant Data Type in Delta Lake and Apache Spark
Databricks
58 How would you sort Ætheldred in the alphabet using Databricks?
How would you sort Ætheldred in the alphabet using Databricks?
Databricks
59 A guide on how to operationalize the Databricks AI Security Framework (DASF)
A guide on how to operationalize the Databricks AI Security Framework (DASF)
Databricks
60 Future-Proof Your Asset Performance Management with Generative AI - Field Assistant Live Demo
Future-Proof Your Asset Performance Management with Generative AI - Field Assistant Live Demo
Databricks

This video teaches how to build reliable ETL pipelines with native observability using Databricks LakeFlow, and demonstrates how to monitor, debug, and optimize data pipelines. It highlights the importance of observability in ensuring trustworthy downstream analytics and machine learning. By watching this video, viewers can learn how to use LakeFlow to streamline ETL development and operations, and improve the reliability and efficiency of their data pipelines.

Key Takeaways
  1. Use LakeFlow to ingest data from various sources
  2. Create declarative pipelines to aggregate and transform data
  3. Set up jobs to orchestrate data pipelines and ensure data freshness and reliability
  4. Monitor data pipelines using LakeFlow's observability features
  5. Debug and optimize data pipelines using LakeFlow's monitoring and alerting capabilities
  6. Set up notifications for pipeline failures or performance issues
  7. Build dashboards to visualize observability at scale
  8. Track data lineage using Unity catalog
💡 Native observability is crucial for ensuring the reliability and efficiency of ETL pipelines, and Databricks LakeFlow provides a rich set of integrated observability features to support data engineers in monitoring, debugging, and optimizing their data pipelines.

Related Reads

📰
I Built My Second ETL Pipeline. This Time, I Started Thinking Like a Data Engineer
Learn how to build a production-ready ETL pipeline with Python, Docker, PostgreSQL, and Kestra by thinking like a data engineer
Towards Data Science
📰
JuiceFS Sync for PB-Scale Data Transfers: Resumable Sync, Encryption, and Bandwidth Control
Learn how to efficiently transfer large volumes of data using JuiceFS Sync, which offers resumable sync, encryption, and bandwidth control, ideal for PB-scale data transfers.
Dev.to AI
📰
How Airflow is using AI to make data engineering more resilient, not more complex
Airflow uses AI to make data engineering more resilient by detecting data drift, resuming failed pipelines, and fixing issues automatically, reducing complexity and improving reliability.
Medium · AI
📰
What Can We Do When Memory Becomes the New Bottleneck in Data Engineering?
Learn how to overcome memory bottlenecks in data engineering using Pandas chunking, Dask, and Polars, and why it matters for processing large datasets
Towards Data Science
Up next
A Moment Frozen in Time | Arnav Iyengar | TEDxJenks Youth
TEDx Talks
Watch →