The document provides a detailed overview of the role, responsibilities, skills, qualifications, and desirable attributes for a Databricks Engineer position.
Role Overview
The position seeks a highly skilled Databricks Engineer responsible for designing, developing, and optimizing data pipelines and analytics solutions. The ideal candidate should possess strong expertise in Databricks, PySpark, and SQL, with experience in building scalable data solutions across cloud platforms .
Key Responsibilities
- Design and implement data pipelines and ETL processes using Databricks and PySpark.
- Develop and optimize SQL queries for data transformation and analytics.
- Collaborate with data architects, analysts, and business stakeholders to deliver high-quality data solutions.
- Ensure data quality, governance, and compliance across all processes.
- Integrate data from multiple sources and efficiently manage large-scale datasets.
- Troubleshoot and optimize performance for data workflows and queries .
Primary Skills (Must-Have)
- Proficiency in Databricks, including Delta Lake and ML capabilities.
- Expertise in PySpark for distributed data processing.
- Strong SQL skills for data manipulation and analytics .
Good-to-Have Skills
- Cloud services experience with Azure (ADF, Synapse, Purview, Fabric).
- AWS services knowledge including Glue, Lambda, Step Functions.
- Workflow orchestration with Airflow.
- Data transformation and integration tools like DBT, Fivetran, Informatica.
- Streaming and messaging platforms such as Kafka.
- BI and visualization skills with Power BI.
- Familiarity with data governance tools like Collibra and Alation.
- Experience with Google Cloud Platform's BigQuery .
Qualifications
- Bachelor’s or Master’s degree in Computer Science, Data Engineering, or related field.
- Proven experience in building and managing data pipelines in cloud environments.
- Strong problem-solving and analytical skills.
- Excellent communication and collaboration abilities .
Nice to Have
- Experience working in multi-cloud environments such as Azure, AWS, and GCP.
- Familiarity with modern data stack and data governance frameworks