Website:
ottomate.global
Job details:
Senior AWS Data Engineer – Databricks/Snowflake & Streaming
Location: India - Remote
The Opportunity
We are seeking a highly skilled, India-based Senior Data Engineer to join our growing data team.
While our architects define the technical blueprint, you will be the lead craftsman responsible for building, optimizing, and maintaining the robust data pipelines that power our real-time analytics, AI/ML initiatives, and enterprise reporting.
You will be a hands-on expert in AWS and modern cloud data platforms, specifically Snowflake or Databricks, with the engineering rigor required to build scalable, reliable, and production-grade data ecosystems.
Work Hours and Collaboration Expectations:
This is a remote position based in India, with required availability during the following core collaboration hours with U.S.-based teams:
- 9:00 AM–2:00 PM EST
- 8:00 AM–1:00 PM CST
- 7:00 AM–12:00 PM MST
- 6:00 AM–11:00 AM PST
The remaining three hours of the workday may be completed before or after the core collaboration window at the employee’s discretion.
What You’ll Do:
Pipeline Development and AWS Data Lake Engineering
- Build and maintain complex data pipelines using AWS Glue, AWS Step Functions, or Databricks Workflows.
- Implement modular data structures using advanced modeling techniques, including Medallion Architecture and Dimensional Modeling.
- Manage scalable data storage solutions using AWS S3 as the primary landing zone and data lake foundation.
- Optimize storage formats, including Delta, Iceberg, and Parquet, to improve processing performance, throughput, and cost efficiency.
- Build decoupled, event-driven architectures using AWS SNS and SQS to support high-throughput messaging between data services.
Real-Time Data Streaming and Ingestion
- Develop and deploy real-time data ingestion pipelines using AWS Kinesis or Kafka.
- Implement Change Data Capture using tools such as Debezium or Fivetran to support low-latency operational analytics.
Data Quality and Automated Validation
- Own end-to-end data validation and quality assurance by building automated data quality checks directly into ETL and ELT pipelines.
- Enforce data contracts and schema-evolution standards to maintain data quality and integrity across domains.
- Implement proactive alerting and observability to identify data drift, pipeline anomalies, and quality degradation before they affect downstream users.
Engineering for ML and AI
- Engineer ML-ready datasets and manage feature stores to support Data Science teams.
- Operationalize ML workflows by integrating with services such as Snowflake Cortex, Databricks AI, or AWS Bedrock.
Technical Leadership and Collaboration
- Mentor junior engineers in coding best practices, SQL optimization, and Python development.
- Collaborate closely with Product, Engineering, Analytics, and ML teams to translate architectural designs into functional, production-ready code.
- Create and maintain technical documentation, including playbooks, technical specifications, operational procedures, and troubleshooting guides.
Required Qualifications:
- 6–8+ years of data engineering experience, with a focus on large-scale distributed systems.
- Expert-level Python and PySpark skills.
- Strong SQL development and optimization experience.
- Deep hands-on experience with either Snowflake or Databricks within an AWS-based ecosystem.
- Proven experience building real-time streaming applications using AWS Kinesis or Kafka.
- Experience implementing automated testing frameworks, data profiling, pipeline validation, and data quality controls.
- Experience working with AWS data services, including S3, Glue, Step Functions, SNS, and SQS.
- Strong understanding of modern data modeling and data lake architecture.
- Strong documentation habits and an ownership-driven engineering mindset.
Collaboration and Ownership:
- Strong communication skills, with the ability to explain technical concepts to both technical and non-technical stakeholders.
- Ability to collaborate effectively across Product, Engineering, Analytics, ML, and leadership teams.
- High standards for quality, maintainability, performance, and operational discipline.
- Strong ownership mindset with the ability to move quickly and solve problems thoughtfully.
Preferred Certifications:
Relevant professional certifications are considered an advantage, including:
- SnowPro Core Certification
- Databricks Certified Data Engineer Professional
- AWS Certified Data Engineer
Click on Apply to know more.