✕ Clear all filters
237 articles
▶ Videos →

Blog Posts

237 articles · Updated every 3 hours · View all reads

All Articles 144,468Blog Posts 147,112Tech Tutorials 37,587Research Papers 28,007News 19,904 ⚡ AI Lessons
OPTIMIZE TABLE ... FINAL in ClickHouse: when to use it, when to avoid it, and how merges work
Dev.to · Aman Puri 🔄 Data Engineering ⚡ AI Lesson 1mo ago
OPTIMIZE TABLE ... FINAL in ClickHouse: when to use it, when to avoid it, and how merges work
If you manage a ClickHouse cluster in production, you may have hit duplicate rows or the "too many...
Understanding ETL: A Chaotic Introduction
Dev.to · Abdi Omari 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Understanding ETL: A Chaotic Introduction
Build your first Python Data pipeline using the News API, and make some sense of the...
Apache Iceberg in Production: Compaction, Catalogs, and the Pitfalls Nobody Warns You About
Dev.to · Gabriel Henrique 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Apache Iceberg in Production: Compaction, Catalogs, and the Pitfalls Nobody Warns You About
Apache Iceberg looked like the answer to everything when we first adopted it. Open format, ACID...
Your Data Engineering Take-Home Is Now 20 Hours of Free Work
Dev.to · DataDriven 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Your Data Engineering Take-Home Is Now 20 Hours of Free Work
Take-homes grew from 4 hours to 20. No pay, no feedback, AI banned with no rubric updates. The DE interview is now just unpaid consulting.
Day 32: Configuring ClickHouse® Clusters with ClickHouse Keeper
Dev.to · Kanishga Subramani 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Day 32: Configuring ClickHouse® Clusters with ClickHouse Keeper
As ClickHouse® deployments grow beyond a single server, ensuring high availability, scalability, and...
Understanding Apache Airflow DAGs: Structure, Communication, and Deployment
Dev.to · Wangila russell 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Understanding Apache Airflow DAGs: Structure, Communication, and Deployment
Apache Airflow has become one of the most widely used workflow orchestration platforms for building,...
Stop waking up at 3 AM: Why your data pipelines must be idempotent
Dev.to · Aniket Abhishek Soni 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Stop waking up at 3 AM: Why your data pipelines must be idempotent
Why I chose this topic: In my first year as a junior engineer, I pushed a non-idempotent job that...
I'm opening ContractForge — define data ingestion intent once, run it natively anywhere
Dev.to · Marco Marques 🔄 Data Engineering ⚡ AI Lesson 1mo ago
I'm opening ContractForge — define data ingestion intent once, run it natively anywhere
Every data engineer who works across platforms knows this pain: You build a clean ingestion layer...
GBase 8a Table Design and Modeling: Choosing Data Types, Partitions, Distribution Keys, and Replicated Tables
Dev.to · Michael 🔄 Data Engineering ⚡ AI Lesson 1mo ago
GBase 8a Table Design and Modeling: Choosing Data Types, Partitions, Distribution Keys, and Replicated Tables
In a distributed analytical gbase database, many performance issues are baked in at the table design...
From DataStage and Informatica to Databricks Medallion Architecture: Why Migration Is More Than Code Conversion
Dev.to · Amit Kumar Singh 🔄 Data Engineering ⚡ AI Lesson 1mo ago
From DataStage and Informatica to Databricks Medallion Architecture: Why Migration Is More Than Code Conversion
Legacy ETL modernization is often described as a technology migration. Move DataStage jobs to...
Synthetic Data for Data Engineering: How to test a Pipeline before the real data arrives
Dev.to · Muhammed Rasin O M 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Synthetic Data for Data Engineering: How to test a Pipeline before the real data arrives
There is a quiet absurdity at the center of most data work, and once you notice it you cannot stop...
AWS DEA-C01: What Each Domain Actually Tests (Not What the Blueprint Says)
Dev.to · ExamCert.App 🔄 Data Engineering ⚡ AI Lesson 1mo ago
AWS DEA-C01: What Each Domain Actually Tests (Not What the Blueprint Says)
The AWS Data Engineer Associate is one of the newer associate-level certs, and because it's new, the...
Top 12 Pipeline Architecture Interview Questions, With Answers
Dev.to · DataDriven 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Top 12 Pipeline Architecture Interview Questions, With Answers
12 real pipeline architecture interview questions with answers: batch vs streaming, idempotency, backfills, DAGs, schema evolution, and monitoring.
Why ClickHouse Merges and Mutations Are Difficult to Track in Production
Dev.to · Kanishga Subramani 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Why ClickHouse Merges and Mutations Are Difficult to Track in Production
One of the reasons ClickHouse delivers exceptional analytical performance is its ability to optimize...
India Daily Brief — fault tolerance patterns from 60 days of broken RSS feeds
Dev.to · Aman Sachan 🔄 Data Engineering ⚡ AI Lesson 1mo ago
India Daily Brief — fault tolerance patterns from 60 days of broken RSS feeds
How a 17-feed RSS pipeline stays alive when TOI, NDTV, Moneycontrol, The Wire, and Scroll.in all break differently. Fault-tolerance patterns, source quality sco
day 01 of learning data engineering (step1: sql joins and set operators)
Dev.to · nain 🔄 Data Engineering ⚡ AI Lesson 1mo ago
day 01 of learning data engineering (step1: sql joins and set operators)
So, yes. Today's goal is to get the 30hr SQL Bootcamp completed (or at least as much as I can), I am...
Track Apache Iceberg Schema Changes in AWS Glue Data Catalog with aws glue get-table-versions
Dev.to · Aki 🔄 Data Engineering ⚡ AI Lesson 1mo ago
Track Apache Iceberg Schema Changes in AWS Glue Data Catalog with aws glue get-table-versions
Original Japanese article: Iceberg × Glue Data Catalogのスキーマ変更履歴をaws glue get-table-versionsで確認する ...
GBase 8a Backup and Recovery Guide: gcrcman from Basics to Production
Dev.to · Michael 🔄 Data Engineering ⚡ AI Lesson 1mo ago
GBase 8a Backup and Recovery Guide: gcrcman from Basics to Production
GBase 8a, as an MPP analytical database, does not use WAL transaction logs. Instead, it relies on the...
From STTM to Snowflake SQL: Building a Metadata-Driven Data Engineering Copilot
Dev.to · Amit Kumar Singh 🔄 Data Engineering ⚡ AI Lesson 1mo ago
From STTM to Snowflake SQL: Building a Metadata-Driven Data Engineering Copilot
A practical build-in-public note on automating repetitive data engineering artifacts from source-to-target mapping metadata.
this is scary (day 0 of learning data engineering)
Dev.to · nain 🔄 Data Engineering ⚡ AI Lesson 1mo ago
this is scary (day 0 of learning data engineering)
apparently i need to build in public and create a personal brand to get a job, which tbh is a big big...