Talentoj
Website:
talentoj.com
Company:
https://www.linkedin.com/company/talentoj
Industries: Software Development
Job details:
Key Responsibilities
- Build and optimize ML model serving infrastructure with a focus on low-latency, cost-efficient inference.
- Architect scalable inference pipelines balancing latency, throughput, and infrastructure cost.
- Develop monitoring, logging, and observability solutions for production ML systems.
- Collaborate with ML Engineers to establish best practices for model deployment and inference optimization.
- Design and implement enterprise-scale, cost-optimized ML infrastructure.
- Work closely with MLEs, QA Engineers, DevOps Engineers, and cross-functional teams.
- Evaluate, benchmark, and adopt new ML infrastructure technologies and tools.
- Contribute to architecture and design decisions for distributed ML systems.
Mandatory Skills & Experience
- 5+ years of software engineering experience with Python.
- Strong hands-on experience with PyTorch.
- Experience optimizing ML models using AWS Neuron, ONNX, and TensorRT.
- Hands-on experience with AWS SageMaker, Inferentia, and Trainium.
- Experience building and operating AWS serverless architectures.
- Strong understanding of event-driven architectures using SQS, SNS, and serverless caching.
- Experience with Docker and container orchestration.
- Strong knowledge of RESTful API design and development.
- Experience writing secure, high-quality, production-grade code and using static code analysis tools.
- Strong understanding of algorithms, data structures, problem-solving, and complexity analysis.
- Excellent verbal and written communication skills.
Click on Apply to know more.