VMC Soft Technologies, Inc
Website:
vmcsofttech.com
Job details:
Job Title: CockroachDB Database Administrator / Site Reliability Engineer (SRE)
Location: Delhi NCR, Bangalore, Pune, Hyderabad, Kochi
Experience: 5+ Years
Employment Type: Full-Time
Job Summary
We are seeking an experienced
CockroachDB Administrator / SRE to manage, operate, and optimize mission-critical CockroachDB clusters in distributed production environments. The ideal candidate will have strong expertise in distributed database architecture, cloud infrastructure, automation, performance tuning, and high-availability systems. The role involves ensuring reliability, scalability, security, and operational excellence of globally distributed CockroachDB deployments.
Key Responsibilities Cluster Administration & Operations
- Design, deploy, configure, and maintain multi-region CockroachDB clusters.
- Ensure high availability, fault tolerance, and data consistency across geographically distributed environments.
- Monitor database health, replication status, cluster performance, and resource utilization.
- Perform capacity planning and implement proactive scaling strategies.
Performance Optimization
- Analyze and optimize SQL queries, indexes, and schema designs.
- Identify and resolve performance bottlenecks related to throughput, latency, and resource consumption.
- Monitor and mitigate hotspotting, leaseholder imbalance, and range distribution issues.
- Implement best practices for database tuning and workload optimization.
Incident Management & Troubleshooting
- Diagnose and resolve complex production issues including:
- Node failures
- Network partitions
- Replication lag
- Leaseholder/range imbalances
- Cluster connectivity issues
- High latency and throughput bottlenecks
- Participate in incident response and root cause analysis activities.
- Develop preventive measures to improve cluster resilience.
Disaster Recovery & Business Continuity
- Design and implement backup, restore, and Point-in-Time Recovery (PITR) strategies.
- Establish disaster recovery procedures for multi-region deployments.
- Conduct regular DR drills, failover, and fallback testing.
- Ensure compliance with RPO and RTO objectives.
Automation & Infrastructure Management
- Automate cluster provisioning, upgrades, patching, monitoring, and maintenance activities.
- Develop Infrastructure-as-Code solutions using Terraform, Ansible, or similar tools.
- Perform rolling upgrades with minimal or zero downtime.
- Integrate CockroachDB operations with CI/CD pipelines.
Documentation & Governance
- Create and maintain operational runbooks, SOPs, troubleshooting guides, and on-call playbooks.
- Document architecture, deployment procedures, and operational standards.
- Ensure adherence to security, compliance, and database governance requirements.
Required Skills & Experience Must-Have Skills
- 5+ years of experience in Database Administration, SRE, or Production Support roles.
- Hands-on experience with CockroachDB in production environments.
- Strong understanding of distributed database architectures and consensus protocols.
- Experience managing multi-region and high-availability database deployments.
- Expertise in SQL performance tuning, indexing, and query optimization.
- Experience with backup, recovery, replication, and disaster recovery planning.
- Knowledge of Linux system administration and troubleshooting.
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Proficiency in scripting using Python, Shell, or Go.
- Experience with monitoring tools such as Prometheus, Grafana, Datadog, or similar.
Good-to-Have Skills
- Experience with Kubernetes and containerized database deployments.
- Knowledge of Terraform, Ansible, or Infrastructure-as-Code tools.
- Experience with distributed systems troubleshooting and networking concepts.
- Familiarity with CI/CD pipelines and DevOps practices.
- Understanding of database security and compliance standards.
Preferred Qualifications
- Bachelor's degree in Computer Science, Information Technology, or related field.
- CockroachDB certifications or relevant cloud/database certifications are preferred.
- Experience supporting large-scale, mission-critical production environments.
Key Competencies
- Strong analytical and problem-solving skills.
- Excellent troubleshooting and debugging capabilities.
- Ability to work in 24x7 production support environments.
- Effective communication and stakeholder management skills.
- Strong documentation and process-oriented mindset.
Location: Delhi NCR | Bangalore | Pune | Hyderabad | Kochi
Experience: 5+ Years
Notice Period: Immediate to 30 Days Preferred.
Click on Apply to know more.