Orion Funded
Website:
orionfunded.com
Job details:
About the Role
We are looking for an experienced DevOps Engineer to design, build, and operate our cloud infrastructure and data streaming platform. You will own infrastructure automation on Google Cloud Platform using Terraform and Ansible, manage Kubernetes workloads at scale with GitOps-driven deployments using ArgoCD and Kustomize, and ensure the reliability of our Confluent Kafka streaming pipelines. Deep expertise in logging, monitoring, and observability is essential. You will be the person the team turns to when something needs to be traced, debugged, or made visible.
Key Responsibilities
- Design, provision, and manage cloud infrastructure on GCP using Terraform, following modular, version-controlled, and reusable Infrastructure as Code patterns.
- Automate configuration management, application deployment, and server provisioning using Ansible playbooks, roles, and inventories.
- Deploy, scale, and maintain containerized applications on Kubernetes, preferably GKE, including autoscaling, ingress, networking, and security policies.
- Manage Kubernetes manifests and environment overlays using Kustomize, and drive GitOps-based continuous delivery with ArgoCD, including application sync, rollout strategies, and multi-environment promotion.
- Administer and operate Confluent Kafka clusters, including topic management, Schema Registry, connectors, consumer group monitoring, performance tuning, and capacity planning. This is a mandatory requirement.
- Build and maintain end-to-end logging and observability solutions, including centralized log aggregation, structured logging standards, log-based alerting, dashboards, and distributed tracing using tools such as Cloud Logging, ELK/EFK, Grafana, Prometheus, Loki, or Datadog.
- Troubleshoot production incidents by analyzing logs, metrics, and traces, and drive root-cause analysis and post-incident reviews.
- Build and improve CI/CD pipelines for automated build, testing, and deployment.
- Implement security best practices, including IAM, secrets management, and network policies.
- Participate in the on-call rotation and continuously improve incident response.
Must-Have Skills
- Terraform: Strong hands-on experience with production-grade Infrastructure as Code, including modules, state management, and workspaces.
- Ansible: Proven experience writing and maintaining playbooks and roles for configuration management and automation at scale.
- GCP: Solid experience with GKE, Compute Engine, Cloud Storage, VPC networking, IAM, Cloud Logging, and Cloud Monitoring.
- Kubernetes: Production experience deploying, scaling, and troubleshooting workloads.
- Kustomize: Hands-on experience managing Kubernetes manifests with bases and overlays across multiple environments. This is a mandatory requirement.
- ArgoCD: Production experience with GitOps-based continuous delivery, including app-of-apps patterns, sync policies, rollbacks, and multi-cluster deployments. This is a mandatory requirement.
- Confluent Kafka: Proven production experience with Confluent Platform or Confluent Cloud, including brokers, Schema Registry, and Kafka Connect. Experience with ksqlDB is a plus. This is a mandatory requirement.
- Logging and Observability: Expert-level experience designing logging pipelines, analyzing high-volume logs, building alerts and dashboards, and debugging distributed systems using logs and metrics.
- Strong Bash scripting skills, along with solid Linux and networking fundamentals.
- Cloud SQL: Hands-on experience provisioning, managing, and tuning Cloud SQL instances using MySQL or PostgreSQL, including backups, high availability, replication, and performance troubleshooting. This is a mandatory requirement.
Nice-to-Have
- Confluent or Kafka certification, such as CCDAK or CCAAK, or GCP certifications.
- Experience with Ansible Tower, AWX, or Red Hat Ansible Automation Platform.
- Experience with Helm charts, service mesh technologies such as Istio, and observability tools such as OpenTelemetry or Jaeger.
- Knowledge of SRE practices, including SLIs, SLOs, and error budgets.
Click on Apply to know more.