MCO (MyComplianceOffice)
Website:
mycomplianceoffice.com
Job details:
The Infrastructure Operations Engineer is responsible for the administration, reliability, security and continuous improvement of production infrastructure. The role sits within L2 Production Support and works closely with TechOps, Platform Monitoring, DBA and Engineering teams to maintain production stability, strengthen platform security, improve infrastructure automation and deliver repeatable operational services. The engineer will provide time zone-aligned Infrastructure and DevOps coverage as part of the organisation’s 24×5 support model and will be expected to investigate complex issues, deliver controlled infrastructure changes and independently own operational work within their areas of competence.
Job Requirements
- Strong Linux systems administration experience in production environments, including system services, logs, packages, processes, permissions, filesystems, resource analysis and performance troubleshooting.
- Practical experience with Linux security controls, including SELinux, SSH, patching, vulnerability remediation and directory-service integration.
- Hands-on experience in administration of production cloud infrastructure; OCI experience is preferred, with solid experience in AWS or other major cloud platforms considered transferable.
- CVE and security awareness.
- Practical knowledge of cloud compute, networking, IAM, storage, tagging, monitoring, certificates, DNS, and host/capacity considerations.
- Experience with automation and configuration management tools or similar technologies.
- Practical understanding of containers and Kubernetes, including workloads, services, configuration, storage, and cluster troubleshooting.
- Ability to investigate and resolve infrastructure issues involving Linux servers, SFTP, storage, permissions and configuration, using available telemetry, documentation and structured troubleshooting methods.
- Experience managing Linux storage, including partitioning, LVM creation and expansion, filesystem growth, mounts, migration, capacity analysis and recovery from common storage issues.
- Practical networking knowledge covering TCP/IP, subnetting, routing, DNS, NAT, firewall configuration, TLS certificates, SSH and system connectivity troubleshooting.
- Strong collaboration skills with the ability to work across Infrastructure, TechOps, SRE, DBA, and Engineering teams.
- Working knowledge of Git, branching and source-control practices, with experience in executing and maintaining CI/CD pipelines.
- Hands-on experience with infrastructure-as-code and configuration management using Terraform or Ansible or similar tools.
- Ability to read, write, and troubleshoot bash scripts and similar operational automation.
- Experience working with incident, problem, and controlled change-management processes, including ticketing systems, technical handovers and change evidence.
Nice to Have
- Experience with Red Hat Enterprise Linux or Oracle Linux; relevant certification is advantageous.
- OCI, AWS, Azure or Kubernetes certification.
- Experience supporting Java middleware or application-server infrastructure.
- Practical AWS administration experience.
- Experience applying CIS standards or comparable security baselines.
- Exposure to enterprise PAM solutions.
- Experience administering Apache HTTP Server, reverse proxies and TLS configuration.
- MongoDB administration or support experience.
- Experience with Ansible Tower, AWX or comparable orchestration platforms.
- Advanced Kubernetes or cluster-administration experience.
- Experience working in regulated, audited or security-sensitive production environments.
Job Responsibilities
- Administer, maintain and troubleshoot Linux-based production infrastructure, ensuring platform availability, security, performance and operational readiness.
- Administer and maintain OCI infrastructure, including compute, networking, IAM, certificates, DNS and NTP, and contribute to capacity, resilience and cost-management activities.
- Plan and execute operating-system patching, hardening and baseline-security activities, including validation, rollback planning and remediation of identified vulnerabilities.
- Administer privileged access, access controls, and platform security processes in accordance with security, audit, and least-privilege requirements.
- Own the investigation and resolution of infrastructure incidents involving Linux, cloud services, networking, storage, SFTP, identity, permissions and environment configuration; coordinate escalation of complex or high-risk issues where specialist intervention is required.
- Actively contribute to infrastructure incident response, service restoration, technical investigation, root-cause analysis and preventive remediation, working closely with TechOps, Platform Monitoring, DBA and Engineering teams.
- Design, maintain and make controlled updates to infrastructure automation and configuration management tools, with emphasis on repeatability, auditability and configuration-drift reduction.
- Maintain infrastructure observability, including metrics, logs, alerts, dashboards and operational health checks; identify performance, capacity and reliability risks.
- Plan and deliver controlled infrastructure changes through established change-management processes, including impact assessment, implementation planning, validation and rollback preparation.
- Create and maintain operational documentation, infrastructure standards, technical procedures, recovery instructions and support runbooks.
- Contribute to capacity planning, infrastructure optimization, backup and recovery validation, disaster-recovery readiness and platform lifecycle management.
Click on Apply to know more.