On Arrival | Plan The Perfect Trip
Website:
onarrival.com
Job details:
About OnArrival
OnArrival is building the infrastructure layer that enables banks, fintech companies, consumer brands and enterprises to offer travel products directly within their platforms.
Our systems handle real-time inventory, pricing, bookings, payments, cancellations and post-booking operations across multiple travel partners. As our transaction volumes and customer base grow, reliability, security, scalability and operational efficiency are critical parts of our product.
About the Role
We are looking for a hands-on DevOps Engineer to build and operate reliable, secure and scalable cloud infrastructure.
You will work closely with engineering teams to improve deployment velocity, observability, infrastructure reliability and incident response.
At OnArrival, AI is an essential part of how we work. You will be expected to use AI tools extensively for writing scripts, creating infrastructure configurations, troubleshooting issues, analysing logs, reviewing changes, documenting systems and automating repetitive operational work.
We are looking for someone who uses AI to increase the quality and speed of their work, while applying strong technical judgement before making production changes.
Responsibilities
* Design, maintain and scale our cloud infrastructure on AWS.
* Manage containerised workloads using Docker, ECS and Kubernetes/EKS.
* Build and improve CI/CD pipelines for automated testing, deployment and rollback.
* Provision and manage infrastructure using Terraform or similar Infrastructure as Code tools.
* Maintain development, staging and production environments.
* Improve monitoring, logging, tracing and alerting across distributed systems.
* Define and track infrastructure SLIs, SLOs and operational health metrics.
* Monitor system availability, application latency, errors, resource utilisation and infrastructure costs.
* Operate and troubleshoot systems involving Kafka, Redis, MongoDB, PostgreSQL, ClickHouse, Elasticsearch and SQS.
* Improve system resilience through autoscaling, redundancy, backups and disaster-recovery processes.
* Strengthen cloud security, IAM policies, secrets management and network controls.
* Support SOC 2, ISO 27001 and other security and compliance initiatives.
* Participate in incident response, root-cause analysis and preventive action planning.
* Improve AWS cost efficiency without compromising performance or reliability.
* Enable engineering teams to deploy services independently through reliable self-service tooling.
Using AI at Work
The successful candidate will be expected to use AI tools extensively in their daily workflow to:
* Write and improve Bash, Python and other automation scripts.
* Generate and review Terraform, Kubernetes, Docker and CI/CD configurations.
* Investigate logs, metrics, traces and production incidents.
* Analyse failed deployments and identify probable causes.
* Create debugging plans and accelerate root-cause analysis.
* Review infrastructure changes for reliability, security and cost risks.
* Generate test cases and validation checklists for infrastructure changes.
* Create and maintain runbooks, architecture documentation and incident reports.
* Summarise incidents and produce clear root-cause analyses.
* Identify repetitive operational tasks that can be automated.
* Learn unfamiliar systems and technologies quickly.
* Compare infrastructure approaches and evaluate technical trade-offs.
* Improve personal productivity and reduce manual operational work.
AI-generated output must always be understood, reviewed, tested and validated before being used in production. Blindly copying AI-generated commands or configurations is not acceptable.
Required Skills
* 3 or more years of experience in DevOps, Site Reliability Engineering or Cloud Infrastructure.
* Strong hands-on experience with AWS services such as ECS, EKS, EC2, S3, CloudFront, IAM, Route 53, SQS and CloudWatch.
* Experience with Docker and container orchestration platforms.
* Practical experience managing Kubernetes workloads in production.
* Experience building CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins or similar tools.
* Experience with Terraform, CloudFormation or another Infrastructure as Code framework.
* Strong Linux, networking and shell-scripting knowledge.
* Experience with monitoring and observability tools such as Grafana, Prometheus, OpenTelemetry, Datadog, SigNoz or similar platforms.
* Understanding of HTTP, DNS, load balancing, CDNs, TLS, VPCs, firewalls and cloud networking.
* Ability to investigate production issues across infrastructure, application and database layers.
* Experience supporting highly available, customer-facing production systems.
* Demonstrated ability to use AI tools effectively in technical work.
* Ability to validate AI-generated output and recognise incorrect or unsafe recommendations.
Good to Have
* Experience with Kotlin, Java or Spring Boot applications.
* Experience operating Kafka or Confluent in production.
* Experience with MongoDB Atlas, Redis, ClickHouse or Elasticsearch.
* Familiarity with WAFs, vulnerability management, secrets rotation and zero-trust security practices.
* Experience working on multi-tenant SaaS, fintech, travel or high-transaction platforms.
* Experience preparing infrastructure for SOC 2 or ISO 27001 audits.
* Familiarity with FinOps and cloud cost optimisation.
* Experience building internal developer platforms or deployment self-service tooling.
* Experience using tools such as GitHub Copilot, Claude Code, Cursor, ChatGPT or similar AI coding and reasoning tools.
What Success Looks Like
Within your first six months, you should be able to:
* Improve deployment frequency while reducing deployment-related failures.
* Reduce mean time to detect and resolve production incidents.
* Establish clear dashboards and alerts for critical customer journeys.
* Improve infrastructure security and access-control practices.
* Reduce repetitive manual work through automation and effective use of AI.
* Improve cloud-cost visibility and eliminate avoidable infrastructure expenses.
* Build reliable backup, recovery and incident-response processes.
* Enable engineers to deploy services confidently without depending on DevOps for routine releases.
* Demonstrate measurable improvements in personal and team productivity through responsible use of AI.
What We Value
* Strong ownership and accountability.
* Extensive and responsible use of AI in everyday work.
* A reliability-first mindset without unnecessary process.
* Clear thinking during production incidents.
* Automation over repetitive manual work.
* Curiosity and the ability to learn quickly.
* Practical problem-solving and attention to detail.
* The ability to balance speed, reliability, security and cost.
* Strong technical judgement rather than blind dependence on tools.
Why Join OnArrival
* Build infrastructure supporting a rapidly growing travel technology platform.
* Work on real-time, high-volume systems used by banks, fintech companies and consumer platforms.
* Work in an AI-first engineering environment.
* Own meaningful infrastructure decisions rather than maintaining a narrow operational function.
* Work directly with engineering and company leadership.
* Help establish the reliability and security foundations of a company entering its next stage of scale.
Click on Apply to know more.