Adastra
Website:
adastracorp.com
Job details:
Role Summary
We are seeking a Senior Databricks Data Engineer to design, build, and operate scalable data platforms on Azure for our manufacturing business. The ideal candidate will lead end-to-end data solutions—data mesh, data lake, and data warehouse architectures—optimize Databricks workloads for performance and cost, enforce governance and security, and mentor engineering teams to deliver reliable, production-grade data pipelines.
Responsibilities
- Design and implement end-to-end data solutions on Azure Databricks, focusing on scalability, reliability, security, and cost optimization.
- Lead the design and rollout of data mesh, data lake, and data warehouse architectures tailored to manufacturing use cases and data domains.
- Translate business and analytics requirements into detailed technical designs, data models, and implementation plans.
- Develop, optimize, and maintain data ingestion and transformation pipelines using Databricks, Apache Spark, PySpark, and Azure Data Factory.
- Tune Databricks clusters, jobs, and Spark workloads for performance and cost efficiency; enforce best practices for partitioning, caching, and resource allocation.
- Implement and enforce data governance, access controls, encryption, and disaster recovery processes across the Databricks environment.
- Produce and maintain technical design documents, runbooks, and operational runbooks that align with business objectives.
- Establish and promote CI/CD practices for data pipelines, notebooks, and infrastructure-as-code (IaC) deployments.
- Collaborate closely with data scientists, analytics, BI, operations, and domain teams to ensure data quality, lineage, and usability.
- Mentor and guide development teams on data engineering standards, Spark/PySpark patterns, testing, and observability.
Requirements
- 6+ years of professional experience as a Data Engineer or similar role working on cloud-based data platforms.
- 4+ years hands-on experience with Azure services including Azure Databricks, Azure Data Factory, and Azure Data Lake Storage.
- Solid experience addressing data modelling requirements and best practices within Azure data platforms.
- Proven experience designing and operating Databricks environments, with strong knowledge of control plane and compute plane concepts.
- Expertise in Apache Spark and PySpark for large-scale data processing, including performance tuning and troubleshooting.
- Solid SQL skills and experience designing data models for analytics and reporting; familiarity with Kimball, Inmon, and Data Vault approaches.
- Experience implementing secure networking and governance strategies for cloud data platforms (VNet, private link, workspace policy controls).
- Practical experience building CI/CD pipelines for data solutions (e.g., using Azure DevOps, GitHub Actions, Terraform, or similar).
- Experience in the manufacturing sector or working with manufacturing data, OT/IoT, and related domain concepts.
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent experience).
- Strong analytical, problem-solving, and communication skills; able to liaise effectively with technical and non-technical stakeholders.
Nice to Have
- Hands-on experience with SAP ERP data extraction, integration patterns, or knowledge of SAP data structures.
- Familiarity with data lineage, metadata management, and tools like Purview, Collibra, or Amundsen.
- Experience with streaming data platforms (Kafka, Event Hubs) and real-time processing on Databricks Structured Streaming.
- Knowledge of containerization and orchestration (Docker, Kubernetes) and infrastructure-as-code (Terraform).
- Advanced degree in a quantitative field or relevant professional certifications (Databricks Certified, Azure certifications).
Click on Apply to know more.