Back to jobs
Toptal

Senior Data Engineer — Azure, Databricks & ML Pipelines | Remote

Anywhere in the World, 🇺🇸 United States of America · US Remote
Sign in to apply
Free account · 30 seconds

About the role

Headquarters: Remote URL: https://www.toptal.com/

About the Role

We're looking for a Senior Data Engineer to design, build, and maintain scalable data pipelines and ML-ready infrastructure on Azure and Databricks. This is a hands-on engineering role: you'll own the full data pipeline lifecycle — ingestion, transformation, orchestration, and deployment — while supporting machine learning workflows with clean, reliable data. If you're comfortable owning infrastructure decisions and writing production-quality Python at scale, this role is built for that.

What You'll Do

Design, build, and maintain data pipelines using Databricks and Azure-native data services Develop and optimize ETL/ELT processes to support analytics and machine learning workloads Build and maintain CI/CD pipelines for data engineering and ML deployment workflows Write clean, efficient, production-quality Python for data processing and pipeline automation Support machine learning teams with well-structured, high-quality datasets and feature pipelines Design and manage data architecture across Azure services (e.g., Azure Data Factory, Azure Data Lake, Azure Synapse) Monitor pipeline performance, troubleshoot data quality issues, and implement reliability improvements Implement data governance, security, and access control best practices Collaborate with data scientists, analysts, and software engineers to align data infrastructure with business needs Participate in code reviews, architecture discussions, and technical planning What You Bring Strong hands-on

experience

with Azure cloud data services Proven

experience

building and maintaining pipelines on Databricks Solid

experience

designing and managing CI/CD pipelines for data or ML workflows Strong Python skills for data engineering and pipeline development Working knowledge of machine learning workflows and how data engineering supports them

Experience

with SQL and relational/distributed data systems Understanding of data pipeline orchestration, monitoring, and reliability practices Strong problem-solving skills and ability to work independently on complex data infrastructure challenges Solid communication skills for collaborating with data science and engineering teams Nice to Have

Experience

with MLOps practices and tools (MLflow, Azure ML) Familiarity with Spark internals and performance tuning within Databricks

Experience

with infrastructure-as-code (Terraform, Bicep, ARM templates) Exposure to real-time/streaming data pipelines (Kafka, Event Hubs, Structured Streaming) Relevant Azure or Databricks certifications Why This Role Full pipeline ownership: Own data infrastructure end to end, from ingestion through ML-ready delivery Modern data stack: Work with Azure and Databricks, leading platforms in enterprise data engineering Cross-functional impact: Directly enable machine learning and analytics outcomes, not just move data Flexibility: Remote-friendly engagement structure How to Apply Ready to bring your data engineering expertise to Azure and Databricks-powered ML infrastructure? Apply through Toptal here: https://www.toptal.com/talent/apply To apply: https://weworkremotely.com/remote-jobs/toptal-senior-data-engineer-azure-databricks-ml-pipelines-remote

Skills this role needs

Not yet in DeepCamp: Azure Data Factory Azure Data Lake Azure ML Azure Synapse MLflow

Unlock your 14-step prep roadmap

DeepCamp maps each role to the exact skills that move you from "interested" to "interview-ready". Free account, 30 seconds.