Tata Consultancy Services
Website:
tcs.com
Job details:
Role- PySpark Developer
Year of Experience- 4 to 15 years
Location -Pune, Hyderabad, Chennai, Bangalore
Technical Skills:
- Pyspark
- Python concepts and Framework
- Spark Architecture
- Big Data
- SQL
Job Description:
Job Requirements*
- Good work experience on Big Data Platforms like Hadoop, PySpark, Scala, Hive, Impala, SQL, Python
- Good Python, Pyspark, Big Data experience
- Spark UI/Optimization/debugging techniques
- Good python scripting skills
- Intermediate SQL exposure – Subquery, Joins, CTE’s
- Database technologies
Key Responsibilities
- *Design scalable PySpark-based test architectures for ETL/data pipelines, including modular frameworks for batch processing
- .Architect end-to-end data validation systems in Hadoop environment for lineage, schema evolutio
- nLead system design for Hadoop/Hive test environments, including YARN resource management, dynamic partitioning
- .Exposure to Zephyr-Jira-ServiceNow integrated test management systems with experience on API-driven automation
- .Design CI/CD test pipelines for PySpark/Hadoop jobs, incorporating artifact management, parallel execution, and blue-green deployments
- .Create data quality system designs using PySpark integrated with Hive metadata services
- .Design testing platforms, test data generator
- sMentor juniors on PySpark testing basics, contribute to testing strategy discussion
- sSpark session configurations for memory and core allocations for both local and cluster manager setting
- sData handling with distributed file systems like HDFS and writing back to hive table
- sImplementation of Partitioning, caching techniques in organizing code for transformation pipeline
- sPerformance tuning implementation like salting, minimizing shufflin
g
Click on Apply to know more.