Website:
clair-x.ai
Job details:
About the Mission
At ClairX, we are building the future of Enterprise AI with Data-First Foundation. As a founding Data Engineer in our India team, you are not just executing tasks; you are an architect of our global data foundation. We are looking for an inquisitive, high-initiative engineer who loves to experiment, thrives in ambiguity, and views technical challenges as opportunities to build something world-class.
What You Will Do
- Solving problems
- Learning and growing with the team
Why Join Us? (The "Founding Mindset")
- Ownership: You will work directly with our global leadership team to shape our technical roadmap.
- Impact: Your work will directly influence our ability to scale and deliver high-quality datasets to our global stakeholders.
- Learning: We encourage "failing fast" and experimentation. If you want to master the cutting edge of GenAI and Lakehouse architecture, this is the place to be.
Technical Skills
This candidate needs a robust, hands-on engineering toolkit, heavily anchored in the Databricks and big data ecosystem:
- Data Processing & Big Data: Apache Spark (PySpark/Scala), Apache Kafka, and Structured Streaming.
- Databricks Ecosystem: Databricks Auto Loader, Delta Live Tables (DLT), Delta Lake (partitioning, Z-Ordering, OPTIMIZE, VACUUM, data compaction), Unity Catalog, Databricks Workflows, MLflow, and Databricks AI.
- Data Architecture & Modeling: Medallion Architecture (Bronze, Silver, Gold layers), ETL/ELT pipelines, dimensional data models (Star/Snowflake), and Lakehouse architecture.
- Orchestration & Cloud: Azure Data Factory (ADF), Apache Airflow, and general cloud storage services.
- Programming Languages: Python, PySpark, and SQL.
- DevOps & Infrastructure: CI/CD pipelines, Git-based version control, automated deployments, and Infrastructure-as-Code (IaC).
- Governance & Quality: Role-Based Access Control (RBAC), data lineage, data quality checks, validation frameworks, metadata management, and logging.
- Modern Data Stack Integration: dbt (data build tool) and Generative AI integrations.
Foundational Skills
Beyond the technical syntax, this role requires the "startup" mindset - someone who takes ownership of systems, costs, and cross-functional relationships:
- System Design & Architecture: The ability to design, develop, and maintain scalable pipelines and robust data ingestion frameworks from scratch.
- Performance & Cost Optimization: Writing highly efficient code focused on scalability and maintainability. Crucially, they must be able to optimize cloud compute costs, Spark cluster configurations, and resource allocation while meeting business SLAs.
- Troubleshooting & Reliability: A proactive approach to monitoring pipeline health, investigating production issues, and ensuring high availability of the data platform.
- Cross-Functional Collaboration: The ability to communicate and partner effectively with Data Scientists, BI Developers, Data Analysts, and business stakeholders to translate technical data into trusted business assets.
- Continuous Learning & Initiative: A proven willingness to stay at the cutting edge of modern data practices, experiment with GenAI capabilities, and proactively implement better solutions rather than just maintaining the status quo.
- Soft Skills: Ability to communicate complex technical concepts, high self-motivation, and the ability to operate independently in a distributed team environment.
Click on Apply to know more.