Website:
nasugroup.com
Job details:
Role Purpose
This role is responsible for designing, deploying, and operating scalable AI/ML and Generative AI
platforms on AWS. The engineer will enable robust MLOps, LLMOps, and DevOps practices,
supporting production-grade AI solutions (including LLMs and agent-based systems) across lending,
risk, and customer journeys.
Key Responsibilities
1. AI/ML & GenAI Operations (AI Ops)
Operate and support production ML and GenAI workloads on AWS
Monitor:
o Model performance, drift, and reliability
o LLM outputs (hallucination, latency, response quality)
Ensure high availability of AI services powering lending workflows
2. MLOps & LLMOps (AWS Native)
Build and manage end-to-end ML pipelines using:
o Amazon SageMaker (training, deployment, pipelines)
o SageMaker Model Registry for versioning
Implement:
o CI/CD pipelines for ML and GenAI workloads
o Experiment tracking and reproducibility
Enable LLMOps practices:
o Prompt lifecycle management
o Evaluation frameworks for GenAI outputs
3. DevOps & Cloud Engineering (AWS)
Design and manage infrastructure using:
o Terraform / AWS CloudFormation
Implement CI/CD pipelines using:
o AWS CodePipeline, CodeBuild, GitHub Actions
Orchestrate workloads with:
o EKS (Kubernetes) and Docker
Ensure scalability, resilience, and cost optimisation
4. GenAI & Agentic AI Enablement
Deploy and manage:
o Amazon Bedrock (LLMs, foundation models)
o RAG pipelines using vector databases (OpenSearch, Pinecone, etc.)
Enable runtime support for:
o AI agents and multi-agent workflows
Integrate AI systems with:
o APIs, event-driven services, and enterprise platforms
5. Data & Pipeline Integration
Build pipelines using:
o AWS Glue, Lambda, Step Functions
Manage data storage and access via:
o S3, Redshift, DynamoDB
Enable real-time and batch AI workflows
6. Monitoring, Observability & Reliability
Implement monitoring using:
o CloudWatch, Prometheus, Grafana
Track:
o Model metrics, pipeline performance, system health
Define SLAs/SLOs and manage incident response
7. Security, Risk & Compliance
Ensure secure AI deployments using:
o IAM, KMS, Secrets Manager
Implement data governance and privacy controls
Enforce Responsible AI and model governance standards
8. Collaboration & Enablement
Work with:
o AI Architects, Data Scientists, Platform Engineers
Enable teams with:
o Reusable MLOps templates and frameworks
o Self-service AI deployment capabilities
Key Skills and Experience
AWS AI/ML & Cloud Stack
Strong experience with:
o Amazon SageMaker (end-to-end ML lifecycle)
o Amazon Bedrock (GenAI / LLMs)
Familiarity with:
o OpenSearch, S3, Lambda, API Gateway
MLOps / LLMOps
Experience implementing:
o ML pipelines, model registries, CI/CD for ML
Knowledge of:
o Prompt engineering workflows
o GenAI evaluation techniques
DevOps & Platform Engineering
Hands-on experience with:
o Docker, Kubernetes (EKS)
o Terraform / CloudFormation
CI/CD:
o CodePipeline, Jenkins, GitHub Actions
Programming
Python (primary), Bash scripting
Experience building APIs (FastAPI preferred)
Monitoring & Reliability
Experience with:
o CloudWatch, ELK stack, Prometheus
Understanding of:
o AI system observability and logging
Click on Apply to know more.