ORMAE
Website:
ormae.com
Job details:
Senior Data Engineer
Location: Pune
Experience: 4+ Years
Role Overview
We are looking for a seasoned Senior Data Engineer with 4+ years of core data engineering experience using Databricks to design and build scalable, real-time and batch data platforms.
The role requires strong expertise in data engineering, data services, Python, advanced SQL, complex data transformations, system integrations, and Medallion Architecture.
Key Responsibilities
- Design and build real-time, near-real-time, and batch data ingestion pipelines.
- Develop Bronze, Silver, and Gold data layers using Medallion Architecture.
- Build complex data transformation, cleansing, enrichment, and reconciliation workflows.
- Integrate data from REST APIs, SFTP, databases, files, enterprise systems, and streaming sources.
- Develop event-driven pipelines using AWS Lambda and DAGs.
- Build scalable data processing solutions using Azure Databricks, Apache Spark.
- Develop reusable Python and SQL components for ingestion and transformation.
- Implement schema evolution, incremental loads, retries, error handling, and data recovery.
- Optimize pipeline performance, data storage, partitioning, and query execution.
- Implement data quality checks, monitoring, logging, alerting, and audit controls.
- Design data models, schemas, and integration mappings across source and target systems.
- Review technical designs and guide data engineers on implementation standards.
Required Skills
- 4+ years of hands-on data engineering PySpark experience.
- Strong Python and PySpark development skills.
- Strong hands-on experience with: Apache Spark, Kubernetes and AKS, Docker, Databricks.
- Advanced SQL skills, including complex joins, window functions, CTEs, and query optimization.
- Strong experience with real-time and batch ingestion patterns.
- Experience implementing Medallion Architecture.
- Strong understanding of Databricks, Databricks Cluster management, ETL and ELT design patterns.
- Experience with REST APIs, SFTP, JSON, CSV, Parquet, and relational databases.
- Experience with Change Data Capture, incremental processing, and event-driven architectures.
- Strong understanding of data modelling, schema design, partitioning, and schema evolution.
- Experience implementing pipeline observability, data validation, monitoring, and alerting.
- Experience with Git, CI/CD, Docker, and Kubernetes is preferred.
Click on Apply to know more.