HARP
Website:
harp-india.com
Job details:
Job Title: GCP Data Engineer
Exp-6 to 10 yrs
Location Gurugram and Bangalore
Job Summary
We are looking for an experienced GCP Data Engineer with strong expertise in building scalable data pipelines and modern data solutions on Google Cloud Platform (GCP). The ideal candidate should have hands-on experience in Python, PySpark, and GCP data services, with the ability to design, develop, optimize, and maintain enterprise-grade data engineering solutions.
Roles & Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and PySpark.
- Build and manage data processing workflows on Google Cloud Platform (GCP).
- Develop batch and real-time data ingestion pipelines from multiple structured and unstructured data sources.
- Work with large-scale datasets and optimize data processing for performance, scalability, and reliability.
- Utilize GCP services such as BigQuery, Cloud Storage, Dataflow, Dataproc, Pub/Sub, Cloud Composer, and Cloud Functions.
- Perform data transformation, cleansing, validation, and integration to ensure high-quality data delivery.
- Optimize SQL queries, PySpark jobs, and data models for improved performance.
- Collaborate with business stakeholders, data analysts, architects, and application teams to understand data requirements and deliver robust solutions.
- Implement data governance, security, and best practices across data platforms.
- Troubleshoot production issues and provide timely resolutions.
- Participate in code reviews, testing, deployment, and documentation activities.
- Mentor junior engineers and contribute to improving engineering standards and development practices.
- Stay updated with the latest GCP technologies and recommend improvements to existing data architectures.
Mandatory Skills
- Strong experience in Google Cloud Platform (GCP) Data Engineering.
- Hands-on expertise in Python programming.
- Strong development experience in PySpark.
- Experience with BigQuery.
- Knowledge of Dataflow and/or Dataproc.
- Experience in building ETL/ELT pipelines.
- Strong SQL and database optimization skills.
- Experience working with Cloud Storage (GCS).
- Good understanding of data warehousing concepts and dimensional modeling.
- Experience with workflow orchestration tools such as Cloud Composer (Apache Airflow).
- Strong debugging and performance tuning skills.
Preferred Skills
- Experience with streaming technologies such as Pub/Sub or Kafka.
- Exposure to CI/CD pipelines and DevOps practices.
- Knowledge of Git and version control systems.
- Familiarity with Docker and Kubernetes.
- Experience with Terraform or Infrastructure as Code (IaC).
- Understanding of data quality, metadata management, and data governance.
- Experience working in Agile/Scrum environments.
Key Skills
- Google Cloud Platform (GCP)
- Python
- PySpark
- BigQuery
- Dataflow
- Dataproc
- Cloud Storage (GCS)
- Cloud Composer (Airflow)
- SQL
- ETL/ELT Development
- Data Warehousing
- Data Modeling
- Pub/Sub
- Git
- CI/CD
- Apache Spark
- Performance Optimization
- Data Integration
- Agile Methodology
Qualification
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
Preferred Candidate Profile
- Total experience: 6–10 years
- Relevant experience in GCP Data Engineering: Minimum 5 years
- Strong analytical and problem-solving skills.
- Excellent communication and stakeholder management skills.
- Ability to work independently as well as collaboratively in a fast-paced environment.
Click on Apply to know more.