Website:
refold.ai
Job details:
About Refold AI
Refold AI is the AI integration platform built for agents, made for modern SaaS companies that need to ship enterprise integrations at scale. We're reimagining integrations with AI agents that work like business consultants and integration engineers, building them 10x faster, in hours instead of days. The payoff: integrations stop being a bottleneck that costs you deals and burns margin, and become a growth engine. You win the enterprise deals that hinge on them, onboard faster, and scale without an army of consultants.
We've grown rapidly, reaching seven-figure ARR and 30+ paying customers, powering the product and professional services teams at modern SaaS companies, from Incorta, Kantata, and Coupa to Highradius, Naehas, and OvationCXM.
Refold AI has raised $6.5M from world-class investors including Eniac Ventures (early investors in Airbnb & Uber), Tidal Ventures, and Better Capital.
What you'll do:
- Own our Helm chart as a product: values design, subchart and global merge semantics, secret handling, ingress and gateway topology, probes, resource envelopes, HPAs, and upgrade paths that survive contact with real customer state.
- Own the Terraform that deploys into customer clouds — VPC, EKS/GKE, managed Postgres and Redis, object storage, IRSA/workload identity, storage classes, autoscaling — written well enough that a customer's own platform team can read, review and run it.
- Run releases: one global platform version, every service image built and published, the release stamped, the digest-pinned bill of materials generated, the release branch reconciled. Make it boring enough that nobody hand-patches a cluster to ship.
- Lead enterprise installs and upgrades with customer platform and security teams — scope their network and compliance constraints, do the deployment, and fix the root cause instead of leaving a workaround behind.
- Debug production across the stack: pod crashloops and OOM kills, failed schedules, DNS and ingress behaviour, CSI and PVC issues, stateful store failover, slow queries, memory leaks.
- Extend the observability stack — OpenTelemetry collectors, Prometheus metrics, alert rules, PagerDuty routing — and make the same signals work inside an air-gapped install where nothing phones home.
- Keep the deployment documentation true: install guides, upgrade guides, backup and restore, troubleshooting. Customer teams read these instead of talking to us, so they ship in the same pull request as the change.
- Maintain developer environments — a per-engineer namespace on a shared dev cluster running the whole platform from production images, plus the local Compose stack — and keep both current with the chart.
- Automate operational toil in Terraform, Bash, Python or Go. If you've done it twice by hand, the third time should be code.
- Partner with security on secrets management, image provenance, CVE triage, patching cadence and audit evidence.
What we're looking for:
- 5+ years in DevOps/platform/SRE, with Kubernetes in production at depth — you've upgraded clusters, debugged a failing node, worked through RBAC and admission, and can explain why a pod wouldn't schedule.
- Helm as an author, not a consumer. You've designed and maintained charts other people install: umbrella charts, subchart values and global merges, template helpers, hooks, and versioned upgrades against live state.
- Terraform in anger.Modules, state and backends, provider quirks, drift, import, targeted operations, and the discipline of a plan someone else reviews. Comfortable with the Kubernetes and Helm providers and clear-eyed about where they hurt.
- One major cloud at production depth(AWS or GCP) — VPCs and subnets, security groups, IAM trust policies, managed databases, object storage. Enough to deploy into someone else's account without root.
- Real CI/CD ownership. You've designed or substantially rebuilt pipelines for multi-service systems — build, publish, promotion, environment gating, and rollback that works under pressure.
- Solid Linux, container, and networking fundamentals—processes and memory, file descriptors, DNS, TLS and certificates, load balancing, proxies, multi-stage Docker builds. You can trace a request end to end.
- Scripting proficiency in Bash plus Python or Go for anything longer than a screen.
- Customer-facing composure. You can run an install call with an enterprise's platform and security team, take their constraints seriously, and write down what you promised.
- Clear written communication—runbooks, install guides, and postmortems are part of the job, not an afterthought.
Nice to have
- Self-hosted or on-premises delivery to regulated enterprises: air-gapped installs, VPN-only networks, private registries, customer-supplied certificates.
- Stateful workloads on Kubernetes — MongoDB, Redis, Postgres — including backup, restore and failover drills.
- Temporal, or another durable-execution or queue-heavy system, in production.
- Observability tooling depth: OpenTelemetry, Prometheus, Grafana or SigNoz, Alertmanager, PagerDuty.
- Service mesh or API gateway experience (APISIX, Envoy, Istio, Kong).
- Secrets and supply chain: External Secrets, Vault, sealed secrets, image signing, SBOMs.
- SOC 2, HIPAA or ISO 27001 environments and the operational discipline they require.
- Cost optimisation on a real cloud bill — spot capacity, autoscaling, scale-to-zero, right-sizing.
- Node.js and Python service operations, gRPC/Protobuf toolchains, or monorepo build tooling.
Click on Apply to know more.