Lead Data Engineer
Wishtree Technologies
- Location
- Pune Division, Maharashtra, India
- Job type
- Full-time
Required skills
- Power BI
- AWS
- automated testing
- Azure
- cloud data services
- compliance
- data modeling
- data warehouse
- Databricks
- DevOps
- end-to-end
- ETL
- GCP
- GitHub
- Oracle
- PostgreSQL
- SAP
- Snowflake
- SQL
- Terraform
- BI tools
About the role
Wishtree Technologies
Website:
wishtreetech.com
Job details:
Key Responsibilities
- Build & Scale: Develop and maintain end-to-end data pipelines spanning source systems, ingestion, transformation, storage, and serving layers for enterprise-scale analytics and reporting.
- Lakehouse Implementation: Implement and optimize modern data platforms using medallion architecture (Bronze/Silver/Gold) on Databricks, Microsoft Fabric, or Snowflake, leveraging Delta Lake / ACID-compliant storage.
- Pipeline Engineering: Build and manage robust ETL/ELT pipelines using Azure Data Factory, Fabric Data Factory, or equivalent orchestration tools, including data extraction from SAP, SAP BusinessObjects, and other enterprise source systems.
- Data Modeling & Quality: Apply advanced data modeling standards (dimensional, normalized, and lakehouse models) and implement automated frameworks for data quality, lineage, and observability.
- BI & Serving Integration: Build and optimize semantic and reporting layers for BI consumption, including Power BI Direct Lake, SQL Analytics Endpoints, and data warehouse / lakehouse serving layers.
- Security & Governance: Implement data governance, security, access controls, and compliance frameworks (such as GDPR/DPDP) within the data pipelines and storage layers.
- Team Leadership & Mentorship: Lead and mentor junior/mid-level data engineers, championing best practices in PySpark/SQL coding standards, performance tuning, and cloud cost optimization.
- DevOps & CI/CD: Drive DevOps best practices for data (DataOps), including automated testing, CI/CD pipeline development, and Infrastructure as Code (IaC).
- PoC to Production: Lead proof-of-concept (PoC) initiatives and rapidly transition them into reliable, fault-tolerant, and high-performance production deployments.
Required Qualifications
- 6+ years of hands-on experience in data engineering, pipeline development, and data platform delivery in production environments.
- Strong expertise implementing and optimizing medallion/lakehouse platforms on Databricks, Microsoft Fabric, or Snowflake.
- Deep technical skills in cloud data services, particularly across Azure (Data Factory, ADLS Gen2, Synapse, Fabric) or equivalent tools on AWS/GCP.
- Deep technical skills in cloud data services, particularly across Azure (Data Factory, ADLS Gen2, Synapse, Fabric) or equivalent tools on AWS/GCP.
- Proven experience integrating enterprise source systems such as SAP, SAP BusinessObjects, and relational databases (SQL Server, Oracle, PostgreSQL).
- Hands-on experience with data modeling techniques, including dimensional modeling, star/snowflake schemas, and lakehouse table design.
- Familiarity with BI tools and semantic layers, specifically Power BI (Direct Lake, Import, DirectQuery).
- Experience with DevOps for data, including CI/CD (Azure DevOps/GitHub Actions), Git, and Infrastructure as Code (Terraform).
- Strong communication and collaboration skills to work effectively with architects, business stakeholders, and delivery teams.
- Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, or a related field (or equivalent practical experience).
Click on Apply to know more.
This page is fully interactive when JavaScript is enabled. Please enable JavaScript to apply or browse related roles.