Mejenta Systems, Inc.
Website:
mejenta.com
Job details:
Job Title: Hadoop Support Engineer – L1/L2: 3 positions
Working hours: 30 hours per week Rotational shifts during U.S. daytime hours, between 8:00 a.m. and 6:00 p.m. CST (6:30 PM - 4:30 AM IST). On-call support is required outside regular working hours
No. of Positions: 3
Coverage: L2 escalation + platform ownership
Position Summary
We are seeking an experienced Hadoop Support Engineer to join our Big Data Operations team, providing Level 1 (L1) and Level 2 (L2) production support across a large-scale, on-premises Hadoop ecosystem. This role is responsible for the day-to-day health, stability, and performance of a multi-cluster environment spanning approximately 700–800 nodes and ~4 PB of storage, built on Acceldata ODP (Apache Hadoop 3.3.x / 3.2.x). The ideal candidate is a hands-on troubleshooter who can monitor, triage, and resolve incidents across HDFS, YARN, Hive, Spark, Trino, ZooKeeper, Ranger, and Ambari — while collaborating with Unix/infrastructure, security, and application teams to keep production workloads running smoothly.
Key Responsibilities
Production Support & Incident Management (L1/L2)
· Provide first- and second-line production support for a multi-cluster Hadoop environment (700–800 nodes) running Acceldata ODP Hadoop 3.3.x / 3.2.x.
· Monitor cluster health and proactively identify, triage, and resolve incidents affecting NameNodes, ResourceManagers, DataNodes, NodeManagers, ZooKeeper ensembles, Client/Gateway nodes, and Ambari-managed services.
· Perform root cause analysis (RCA) for recurring incidents and service outages, and drive permanent fixes in partnership with L3/engineering teams.
· Manage and resolve an average monthly volume of L2/L3 incidents, service outages, and routine service requests, meeting defined SLAs.
· Own ticket lifecycle management (logging, triage, escalation, resolution, and closure) using the enterprise ITSM/ticketing platform.
· Participate in an on-call rotation and provide off-hours support for critical production incidents and planned maintenance windows.
Cluster Administration & Platform Operations
· Administer and support HDFS and OneFS-based storage deployments (~80% of clusters use OneFS as the primary HDFS filesystem) across on-premises bare-metal infrastructure, with Dev/QA clusters running on VMs.
· Maintain and support core Hadoop components: HDFS, YARN/ResourceManager, ZooKeeper, and Ambari-managed services, across dedicated per-cluster blueprints with isolated resource allocation.
· Support master, worker, and edge node configurations (32–160 vCPUs, 250–2000 GB RAM depending on cluster) and coordinate with infrastructure/Unix teams on hardware and OS-level issues.
· Manage capacity, performance tuning, and resource allocation across NameNodes, ResourceManagers, and NodeManagers to prevent contention and ensure workload SLAs.
· Support cluster upgrades, patching, configuration changes, and decommissioning/commissioning of nodes with minimal production disruption.
Compute & Query Engine Support
· Support and troubleshoot Apache Spark job failures, performance issues, and resource allocation problems.
· Support Apache Hive (including LLAP and Tez execution engines), including query performance issues, metastore health, and job failures.
· Provide support for Trino query engine deployments (currently used in select clusters).
· Work with application and data engineering teams to troubleshoot job submission, scheduling, and resource contention issues across YARN queues.
Security & Governance
· Support Apache Ranger for policy management, RBAC, and access governance across the Hadoop ecosystem.
· Manage user and group access via Active Directory / LDAP synchronization; support onboarding, offboarding, and access remediation requests.
· Support Kerberos authentication where enabled and assist in extending secure authentication coverage across additional clusters as required.
· Enforce data access and governance policies in line with Acceldata ODP Apache governance architecture and enterprise security standards.
· Assist with security audits, access reviews, and compliance reporting related to platform governance.
Monitoring, Alerting & Observability
· Monitor cluster and service health using Acceldata Pulse, SolarWinds, and custom cron-based monitoring scripts.
· Respond to alerts and thresholds breaches, escalating to L3/engineering or infrastructure teams as needed.
· Support and improve log aggregation and troubleshooting practices across clusters, given the currently decentralized, cluster-specific logging implementations.
· Contribute to the standardization of monitoring, alerting, and log aggregation practices across the environment.
· Isilon / OneFS Basics: Health monitoring, storage capacity tracking, basic node/drive health ops — all three L1s cross-trained on Isilon
·
Collaboration & Documentation
· Partner with Unix/Infrastructure, Network, Security, and Application Development teams to resolve cross-functional issues.
· Maintain and update runbooks, SOPs, and knowledge base articles for recurring issues and standard operating procedures.
· Participate in change management processes for production changes, patches, and maintenance activities.
· Provide clear, timely communication and status updates during incidents and outages to stakeholders and leadership.
Required Qualifications
· Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field (or equivalent practical experience).
· 3–6+ years of hands-on experience supporting or administering Hadoop environments in a production capacity.
· Solid working knowledge of Hadoop 3.x architecture, including HDFS, YARN, ResourceManager/NodeManager, and ZooKeeper.
· Experience with Apache Ambari for cluster management, monitoring, and service configuration.
· Hands-on experience supporting Apache Spark and Apache Hive (LLAP/Tez) in production.
· Experience with Apache Ranger for RBAC and policy-based access governance.
· Familiarity with Active Directory / LDAP integration for identity and access management in Hadoop ecosystems.
· Strong Linux/Unix systems fundamentals (networking, storage, performance troubleshooting) in an on-premises, bare-metal environment.
· Experience working within an ITSM/ticketing framework, meeting SLA-driven L1/L2 support commitments.
· Strong analytical and root-cause-analysis skills, with the ability to triage and resolve incidents under time pressure.
· Willingness to work rotational shifts and participate in an on-call rotation, including weekends as needed.
Preferred Qualifications
· Experience with Acceldata ODP or similar enterprise Hadoop distributions.
· Experience with OneFS or other scale-out NAS storage integrated as an HDFS-compatible filesystem.
· Exposure to Trino (or Presto) query engine administration and support.
· Experience with Kerberos authentication setup and troubleshooting in distributed systems.
· Experience with monitoring tools such as Acceldata Pulse, SolarWinds, or equivalent enterprise monitoring platforms.
· Scripting experience (Shell, Python, or similar) for automation of monitoring, health checks, and routine operational tasks.
· Experience supporting large-scale, multi-cluster environments (500+ nodes) with high-availability configurations.
· Relevant certifications (e.g., Cloudera/Hortonworks Hadoop Administrator, Linux, or equivalent) are a plus.
Databricks Competency: Familiarity with Databricks workspace operations — job monitoring, cluster health checks, alerting integrations, and basic triage of Databricks Spark job failures before escalating to the SME
Click on Apply to know more.