Ekfrazo Technologies Private Limited
Website:
ekfrazo.com
Job details:
Role: Data Architect
Exp - 11 - 15 yrs
Mode of work - Hybrid ( 2 days in a week )
Location: Bangalore
Budget: 36 - 42 LPA
Must-haves: Data Architecture, Data Modeling, Data solutioning on Azure, AWS and/or GCP, Data governance, Python, SQL.
Secondary skills: AI, Migration and/or Modernization of data platforms, enterprise architecture, Databricks, Snowflake.
Architect - We are seeking a Data Architect with strong experience in data platform migration and modernization, specifically transitioning workloads from DataIKU to Azure Databricks. The role involves designing, building, and optimizing scalable data pipelines using Databricks, Azure Data Factory (ADF), and PySpark.
Job Description:
Core Skills
Migration & Modernization
• Lead and execute migration of data workflows from DataIKU to Databricks
• Analyze existing DataIKU pipelines, datasets, and workflows
• Re-engineer and optimize pipelines using Databricks (PySpark-based processing)
• Ensure functional parity and improved performance post-migration
________________________________________
Data Engineering & Pipeline Development
• Design and develop scalable ETL/ELT pipelines using:
o Azure Data Factory (ADF)
o Databricks (PySpark, Spark SQL)
• Handle batch and near real-time data processing
• Implement data transformations, validations, and error handling
________________________________________
Data Modeling & Storage
• Design and implement data models in Delta Lake / Data Lake (ADLS)
• Optimize data storage using:
o Partitioning
o File formats (Parquet, Delta)
• Ensure efficient data access and performance
________________________________________
Performance Optimization
• Optimize Spark jobs for performance and cost efficiency
• Tune queries and pipelines (caching, partitioning, joins)
• Monitor and troubleshoot pipeline failures and bottlenecks
________________________________________
Integration & Automation
• Integrate data pipelines with upstream/downstream systems
• Automate workflows using ADF pipelines and triggers
• Implement CI/CD pipelines (Azure DevOps preferred)
________________________________________
Validation & Testing
• Ensure data consistency between DataIKU and Databricks outputs
• Develop and execute reconciliation and validation scripts
• Collaborate with QA teams for end-to-end validation
________________________________________
Documentation & Knowledge Transfer
• Document architecture, pipelines, and workflows
• Support knowledge transition to internal teams
• Create runbooks and operational guidelines
________________________________________
Required Skills & Experience
Core Technical Skills
• Databricks (mandatory)
• PySpark / Spark SQL (strong hands-on experience)
• Azure Data Factory (ADF)
• Experience with DataIKU (DSS) or similar legacy data tools
________________________________________
Data Engineering & Cloud
• Azure Data Lake Storage (ADLS Gen2)
• Delta Lake architecture
• Data modeling (dimensional, normalized)
• ETL/ELT pipeline design
________________________________________
Programming & Scripting
• Python (advanced)
• SQL (advanced)
• Shell scripting (good to have)
________________________________________
DevOps & Automation (Preferred)
• CI/CD pipelines (Azure DevOps / GitHub Actions)
• Version control (Git)
• Infrastructure as Code (Terraform – good to have)
________________________________________
Good-to-Have Skills
• Experience in migration projects (lift-and-shift or re-platforming)
• Knowledge of stream processing (Kafka, Structured Streaming)
• Exposure to data governance and quality frameworks
• Experience with monitoring tools (e.g., Datadog, Azure Monitor)
________________________________________
Soft Skills
• Strong analytical and problem-solving ability
• Ability to work in fast-paced, production environments
• Effective communication with cross-functional teams
• Ownership and accountability mindset
________________________________________
Experience
• 5–10 years in Data Engineering
• At least 2+ years in Databricks / Spark ecosystem
• Proven experience in data platform migration projects
________________________________________
Typical Deliverables
• Migrated pipelines from DataIKU ? Databricks
• ADF orchestrations and workflows
• Optimized PySpark jobs
• Data validation and reconciliation reports
• Documentation and KT artifacts
Click on Apply to know more.