Aligned Automation
Website:
alignedautomation.com
Job details:
About the Job
About Aligned Automation
At Aligned Automation, we live by our "Better Together" philosophy to build a better world. As a strategic service provider to Fortune 500 companies, we help digitize enterprise operations and drive impactful business strategies. Our purpose goes beyond projects—we strive to deliver meaningful, sustainable change that shapes a more optimistic and equitable future.
Our culture is deeply rooted in our 4Cs—Care, Courage, Curiosity, and Collaboration—ensuring that each employee is empowered to grow, innovate, and thrive in an inclusive workplace.
About the Role
Key Responsibilities
- Design, build, and maintain scalable MLOps platforms for training, deploying, monitoring, and managing machine learning models.
- Develop and optimize CI/CD pipelines for ML applications using GitLab CI, Jenkins etc.
- Containerize applications using Docker and orchestrate workloads using Kubernetes.
- Build and maintain ML pipelines using Kubeflow, MLflow, Airflow, Argo Workflows, or similar orchestration tools.
- Implement model versioning, experiment tracking, model registry, and reproducible ML workflows.
- Deploy ML models as REST/gRPC APIs using FastAPI, Flask, or similar frameworks.
- Configure monitoring, logging, and alerting using Prometheus, Grafana, Dynatrace, or Splunk.
- Optimize model serving performance, scalability, and resource utilization.
- Work closely with Data Scientists, Data Engineers, Platform Engineers, and Software Developers to productionize ML solutions.
- Ensure platform security, governance, and compliance following DevSecOps best practices.
- Troubleshoot production issues and perform root cause analysis for ML workloads.
Required Technical Skills
- Strong programming experience in Python.
- Hands-on experience with Docker, Kubernetes, Helm, and containerized deployments.
- Strong knowledge of Git, branching strategies, merge requests, and CI/CD.
- Experience with MLflow, Kubeflow, Airflow, or similar MLOps tools.
- Experience with Spark/PySpark and distributed data processing.
- Knowledge of REST APIs, FastAPI, Flask, and microservices architecture.
- Experience with Linux, Bash scripting, and automation.
- Familiarity with model monitoring, drift detection, feature stores, and model governance.
- Experience working with SQL and NoSQL databases.
- Strong understanding of software engineering principles, design patterns, testing, and code quality.
Preferred Skills
- Experience with Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), vector databases, and AI orchestration frameworks such as LangChain or lang-fuse.
- Experience with NVIDIA GPU workloads and model optimization.
- Knowledge of Kafka, RabbitMQ, or event-driven architectures.
- Exposure to feature stores such as Feast.
- Experience with OpenShift, Red Hat ecosystem, or enterprise Kubernetes platforms.
- Knowledge of security scanning tools such as Trivy, SonarQube, and vulnerability management.
Resource should be technically strong, should have good communication skills and experience in handling customer communications and leading a team working on multiple projects. He should be able to generate ideas drive, customer communications and manage team to ensure quality delivery and great customer experience.
Click on Apply to know more.