Website:
Job details:
About Novatro
We're a growing and focused team that moves fast and builds with intention — passionate about creating a great culture and transforming how industrial operations run. Every person here has real ownership and real impact. Novatro is building an AI-powered Intelligent Manufacturing Platform that is reshaping how medium-to-large manufacturers operate across industries including oil & gas, aerospace, and defence.
The Opportunity
Are you someone who thrives on keeping complex operations running smoothly behind the scenes? We're looking for a ProdOps (Production Operations) specialist to help ensure our platform runs reliably in production and that operational issues are caught and resolved before they impact our customers.
Salary and Benefits
- We offer a competitive salary, based on skills and experience with lots of opportunities for career growth
- This is a hybrid, permanent full time position based in India
Key Responsibilities
- Monitor platform performance and reliability in production, identifying and escalating issues proactively
- Coordinate with engineering to triage, prioritize, and resolve production incidents
- Build and maintain operational dashboards, alerts, and runbooks
- Support release and deployment processes to minimize disruption to customers
- Work closely with Customer Experience to understand how production issues are impacting real customer workflows
- Continuously identify opportunities to improve operational stability and reduce recurring issues
- Watch production health against our SLOs. The starting target is 99.9% monthly availability, with alarms on error rate, web latency, Sidekiq queue latency, database pressure, and Redis. Check the web app, background jobs, the load balancer, PostgreSQL, and Redis before a problem spreads.
- Watch production health for the Node.js API, the React application, background workers, the load balancer, PostgreSQL, and Redis. Alarms cover error rate, p95 and p99 latency, default-queue latency over 60 seconds, heavy-queue latency over 15 minutes, database connections, and Redis memory or evictions.
- Triage incidents from CloudWatch alarms, synthetic checks, security findings, and user reports. Set severity from SEV1 to SEV4, open the incident channel, and coordinate diagnosis across dashboards, logs, and traces.
- During a material incident, own the customer update with Customer Experience. Confirm the service has recovered against the SLO before the incident is closed, then write the post-incident review and assign repeat issues to an owner.
- Support the release bake window. Confirm what changed, watch rollback alarms, and roll back the affected release when a deploy alarm or smoke test fails. Freeze deploys during SEV1.
- Build and keep the dashboards and runbooks: executive health, release, application, infrastructure, security, and database. Wire alarms to on-call.
- Run the failure playbooks: platform outage, one customer unavailable, a capacity spike, a background-job backlog, database pressure, Redis pressure, a failed migration, a failed deploy, a security finding, a backup failure, and a regional outage.
- Take part in on-call for production alarms and release windows.
- Use investigation aids, including AI summaries, as input only. Rollbacks outside the pre-approved bake rule, scaling, database changes, and customer messages stay human-approved.
Qualifications and Skills
- Strong analytical and troubleshooting skills, comfortable diagnosing issues across systems
- Clear communicator, able to translate technical incidents into plain language for non-technical stakeholders
- Comfortable working in a fast-paced environment with shifting priorities
- Level C1 or above English proficiency
- Detail-oriented with a bias toward proactive monitoring over reactive firefighting
Experience and Skills Required
- Experience in a production operations, DevOps-adjacent, or technical support role, ideally within a SaaS or platform environment
- Familiarity with monitoring and alerting tools (e.g. Datadog, Grafana, or similar) is a plus
- Basic understanding of cloud infrastructure (AWS or similar) is beneficial
- Exposure to manufacturing or industrial software environments is a plus, but not required
- Experience in production operations, SRE, or technical support for a SaaS or platform product.
- Hands-on familiarity with Amazon CloudWatch (metrics, alarms, logs, and synthetics) and Grafana. Prometheus and tracing (X-Ray or OpenTelemetry) are a plus.
- Working knowledge of AWS: a load balancer, ECS or EC2, RDS PostgreSQL, ElastiCache Redis, and a deployment that can roll back.
- Comfortable running an incident: severity, an incident channel, a runbook, and a clear update for the customer.
- Exposure to manufacturing, ERP, or shop-floor software is a plus
- Experience in production operations, SRE, or technical support for a SaaS or platform product.
- Hands-on support of a Node.js and React application in production: Node.js API logs and errors, React client failures, and background workers.
Interview Process
Candidates who progress will be invited to a multi-stage interview process, including a scenario-based troubleshooting exercise.
Contact
If you're interested, please send your resume and introduction to careers@novatro.ai. Subject line: "ProdOps - [Your Name]".
We look forward to hearing from you!
Click on Apply to know more.