TRUGlobal
Website:
truglobal.com
Job details:
Data Engineer / Analyst
Job Description
Experience Level: 5-7 Years
About the Role
We are looking for a versatile Data Engineer & Analyst who can bridge the gap between robust data infrastructure and actionable business insights. The ideal candidate will design and maintain scalable data pipelines while also delivering meaningful analytics and dashboards to drive decision-making. This role requires strong hands-on expertise in modern data engineering tools combined with a working knowledge of AI/ML concepts and their practical application in data workflows.
Key Responsibilities
- Design, build, and maintain scalable ETL/ELT pipelines using Apache Airflow for orchestration and workflow automation
- Develop and optimize data models and warehousing solutions in Snowflake
- Write efficient, production-grade SQL queries for data transformation, validation, and reporting
- Build data processing and automation scripts using Python
- Leverage Databricks for large-scale data processing, transformation, and advanced analytics workloads
- Design and develop interactive dashboards and reports using Power BI to support business stakeholders
- Collaborate with cross-functional teams (business, product, and engineering) to gather requirements and translate them into data solutions
- Ensure data quality, integrity, and governance across pipelines and reporting layers
- Identify opportunities to apply AI/ML techniques (e.g., predictive modeling, anomaly detection, NLP-based automation) to improve data processes or generate business insights
- Implement basic AI-driven solutions or proof-of-concepts (e.g., using LLMs, forecasting models, or classification models) and integrate them into existing data workflows
- Monitor pipeline performance, troubleshoot issues, and optimize for cost and efficiency
- Document technical processes, data flows, and architecture for team knowledge sharing
Required Skills & Experience
- 5-7 years of experience in data engineering, data analytics, or a related field
- Strong hands-on experience with Apache Airflow for pipeline orchestration
- Proficiency in Snowflake (data modeling, performance tuning, warehouse management)
- Advanced SQL skills (complex queries, optimization, stored procedures)
- Strong programming skills in Python (data manipulation, automation, libraries like Pandas/PySpark)
- Experience with Databricks for big data processing and analytics
- Proficiency in Power BI (DAX, data modeling, dashboard design, report publishing)
- Practical exposure to AI/ML concepts with at least one proven implementation (e.g., a deployed model, automated AI-driven workflow, or use of LLMs/APIs in a business context)
- Hands-on experience with dimensional data modeling (star schema, snowflake schema, fact/dimension tables)
- Solid understanding of OLTP and OLAP systems and the ability to design solutions appropriate to each
- Solid understanding of data warehousing concepts, ETL/ELT design patterns, and cloud data architecture
- Familiarity with AWS (e.g., S3, Glue, Redshift, Lambda, or similar data-related services)
Preferred Qualifications
- Prior experience in the semiconductor industry, working with domain-specific data such as fab/manufacturing data, yield analysis, wafer testing, supply chain, or product engineering datasets
- Experience with version control (Git) and CI/CD for data pipelines
- Familiarity with data governance, security, and compliance best practices
- Exposure to MLOps or LLM integration frameworks (e.g., LangChain, OpenAI/Anthropic APIs)
- Strong problem-solving skills and ability to work independently in a fast-paced environment
- Excellent communication skills to translate technical concepts for non-technical stakeholders
Education
- Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field (or equivalent practical experience)
Click on Apply to know more.