Customertimes
Website:
customertimes.com
Job details:
Job Description:
We are seeking an Infrastructure Platform Engineer to manage and optimize the platform infrastructure that supports our applications. The ideal candidate will have hands-on experience with compute infrastructure, orchestration, RAG pipelines, CI/CD, observability, load testing, and security best practices. The role focuses on ensuring scalability, reliability, performance, and security across the platform ecosystem.
Location: Bangalore/Pune
Number of openings: 2
Key Responsibilities
- Design, build, and operate production-grade infrastructure for AI and LLM-based products.
- Manage compute infrastructure, container orchestration, RAG pipelines, CI/CD processes, observability, load testing, and security posture.
- Own and maintain production Kubernetes clusters, workloads, RBAC policies, and deployment pipelines.
- Provision and manage GPU compute environments for AI workloads.
- Design, deploy, and optimize vector database infrastructure and embedding pipelines.
- Build and maintain infrastructure-as-code modules and GitOps workflows.
- Collaborate with multiple product teams and manage competing infrastructure priorities.
- Lead technical discussions during InfoSec and customer security reviews, including evidence preparation and control documentation.
- Implement monitoring, logging, distributed tracing, and incident-response mechanisms.
- Participate in on-call rotations, troubleshoot production issues, and prepare post-mortem reports.
- Ensure platform reliability, scalability, performance, and security.
Required Qualifications
- 5–8 years of experience in Platform Engineering, SRE, or Infrastructure Engineering.
- Strong infrastructure background with experience owning production systems.
- Experience delivering infrastructure for at least one AI or LLM product in production.
- Hands-on experience managing infrastructure across multiple product teams.
- Proven experience leading InfoSec or customer-facing security reviews.
- Experience supporting production services with defined SLAs and participating in on-call rotations.
Technical Skills:
Container Orchestration & Runtime
Must-have:
- Kubernetes / EKS: Production cluster operations, workload management, and RBAC.
- Docker: Image creation, multi-stage builds, and container lifecycle management.
- Helm: Chart authoring, templating, and release management.
Cloud Platforms
Must-have:
- Microsoft Azure.
- Azure Key Vault.
- Azure Monitor and Log Analytics Workspace.
- Azure DevOps Pipelines.
Nice-to-have:
- AWS.
- Azure Kubernetes Service (AKS).
- Azure API Management.
Infrastructure as Code
Must-have:
- Terraform: Modules, state management, and remote backends.
- GitOps tools such as ArgoCD or Flux CD.
Nice-to-have:
- Pulumi (Python or TypeScript).
AI / ML Infrastructure
Must-have:
- Vector databases such as Milvus, Weaviate, pgvector, or FAISS.
- GPU compute provisioning using data center
- RAG pipeline infrastructure, including ingestion, chunking, and embedding pipelines.
- LLM inference infrastructure, including serving, scaling, and latency management.
CI/CD & Deployment
Must-have:
- GitHub Actions.
- Azure DevOps Pipelines.
- Container registry management (ECR, ACR).
Observability & Monitoring
Must-have:
- Datadog.
- Elasticsearch or OpenSearch.
- OpenTelemetry.
- Lightstep or distributed tracing tools.
- Azure Monitor and Log Analytics integration.
- Incident management tools such as PagerDuty.
Performance, Security & FinOps
Must-have:
- Load testing using tools such as k6, Locust, or JMeter.
- Security best practices, including secrets management, network isolation, and RBAC audits.
- InfoSec review readiness, controls documentation, and evidence preparation.
Nice-to-have:
- Python scripting for infrastructure automation and ingestion pipelines.
Relevant candidates can reach out to reshma.ponnamma@customertimes.com
Click on Apply to know more.