Databricks Data Engineer
Engineering
About the Role
We are seeking an experienced Databricks Data Engineer to design and develop scalable data pipelines and cloud-based data solutions using Databricks, Spark, PySpark, and SQL. The role will focus on building high-performance data platforms, optimizing data workflows, and ensuring data quality and reliability.
The ideal candidate will collaborate with data architects, analysts, and business stakeholders to deliver scalable, efficient, and business-driven data solutions.
Responsibilities
Design and develop scalable data pipelines and ETL/ELT workflows using Databricks.
Build data processing solutions with Apache Spark, PySpark, and SQL.
Develop and optimize batch and near-real-time data pipelines.
Implement solutions using Delta Lake, Delta Live Tables (DLT), and Databricks Workflows.
Develop robust data ingestion, transformation, cleansing, and validation processes.
Integrate data from databases, APIs, files, and cloud storage platforms.
Optimize pipelines and Spark workloads for performance, scalability, and cost efficiency.
Implement data quality, governance, security, monitoring, and validation practices.
Leverage Azure, AWS, or GCP data services to build cloud-native solutions.
Collaborate with architects, analysts, and engineering teams to deliver business-focused data solutions.
Troubleshoot pipeline, data quality, and performance issues.
Follow best practices for Git, CI/CD, testing, and deployment.