Jash Data Sciences
Website:
jashds.com
Job details:
Jash Data Sciences: Letting Data Speak, AI Act!
You'll join a data platform team working on the AI-first infrastructure-as-code, enablement tooling, and platform administration that ensure teams across the company can use our Databricks platform effectively. Instead of building product pipelines, you'll contribute to the platform layer that everyone else builds upon, helping keep it reliable, secure, and easy to use. You'll also help maintain existing MLOps tooling and legacy ETL infrastructure. If you like working at the overlap of cloud infrastructure and data platforms, and you want your work to support clean energy, this is a great fit.
We are a cutting-edge Data Sciences and Data Engineering startup based in Pune, India. We believe in continuous learning and evolving together. And we let the data speak!
What will you be doing?
- Build and maintain multi-cloud Databricks infrastructure using Terraform, primarily across AWS and GCP.
- Contribute to Unity Catalog administration, including best practices, metastores, workspaces, catalogs, schema infrastructure, and Databricks Asset Bundles (DABs).
- Help build deployment accelerators, templates, and self-service tooling so other teams can onboard and ship independently on the Databricks platform.
- Act as a platform engineer, building self-service tools and 'golden paths' for other teams, and contribute to community enablement through documentation, training, and cross-team knowledge sharing.
- Maintain existing MLOps tooling and legacy ETL infrastructure.
- Write clean, efficient Python code to orchestrate and support data and platform pipelines
What do we need from you?
- 3 to 8 years of software/data engineering experience, ideally in infrastructure or data platform roles.
- Strong proficiency in Python/PySpark and strong knowledge of SQL, including windowing
functions, subqueries, and various types of joins.
- Experience with Terraform or similar Infrastructure-as-Code (IaC) tools.
- Strong knowledge of at least one cloud platform (AWS/GCP).
- Solid CI/CD experience and familiarity with Github workflows, including container versioning.
- Strong background in Databricks, specifically Unity Catalog, medallion architecture, and Databricks Asset Bundles (DABs).
- Proficiency in Python and API frameworks such as FastAPI or Flask, with strong experience designing, developing, and maintaining RESTful APIs.
- Good understanding of data engineering concepts such as data modeling and workflow orchestration.
- Experience with MLOps and Data Engineering frameworks such as Spark, Delta Lake, MLflow, and ETL tools (e.g., Fivetran, Lakeflow Connect).
- Proficiency with incident management and observability tools (e.g., PagerDuty, system monitoring).
- A collaborative, 'AI as a partner' mindset - comfortable asking questions, navigating unclear requirements, and contributing to knowledge sharing.
- A good team player with the ability to communicate with clarity. Show us your git repo/blog!
Qualification:
- 3 to 8 years of hands-on experience in software/data engineering roles.
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
- Courses or certifications in Data Engineering or Databricks will be given higher preference.
Click on Apply to know more.