Ethara AI
Website:
ethara.com
Company:
https://www.linkedin.com/company/ethara-ai
Industries: Artificial Intelligence
Job details:
Principal Architect (AI Infrastructure, Harness & Environment Design)
Role Overview
As the Principal Architect at Ethara.AI, you will be the ultimate technical authority for our Reinforcement Learning as a Service (RLaaS) platform, evaluation harnesses, and multi-agent simulation ecosystems (such as MILO-Bench). With over a decade and a half of engineering mastery, you will design the macro-architecture that bridges frontier foundation models with high-throughput distributed compute.
You will be directly responsible for the technical blueprint, state-machine logic, and execution layers of our simulation environments and evaluation harnesses, ensuring AI agents can safely, deterministically, and realistically execute long-horizon, multi-step reasoning tasks.
Key Responsibilities
- Harness & Environment Design: Own the architectural design of our simulation environments and evaluation harnesses. Define how agent action-spaces, observation spaces, and environment state transitions are modeled, ensuring absolute determinism, scalability, and alignment with real-world software development ecosystems.
- Enterprise System Blueprinting: Design and govern the foundational, highly distributed execution systems capable of running stateful, long-horizon simulations concurrently for thousands of reinforcement learning agents.
- Algorithmic Infrastructure: Architect high-throughput, low-latency pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), ensuring seamless integration between expert human data loops, automated environment rewards, and model training runs.
- Future-Proofing Compute: Drive long-term strategies for GPU/CPU cluster optimization, multi-node communication topologies, and ultra-low-latency inference engines (vLLM, TensorRT-LLM) at extreme scale.
- Technical Stewardship: Act as the primary technical advisor to the executive team, steering the company’s infrastructure roadmap and mentoring senior engineering staff.
Qualifications
- Experience: Minimum 15+ years of hands-on experience in software engineering, distributed systems architecture, and high-performance computing (HPC).
- Simulation & Harness Mastery: Proven track record of designing custom evaluation harnesses, sandboxed execution testing layers, or complex simulation environments (e.g., game engines, physics simulators, or heavy multi-step testing frameworks).
- System Design Expertise: Deep experience scaling and maintaining mission-critical, massive-scale distributed infrastructure (e.g., large-scale Ray, Slurm, Kubernetes, or Spark clusters).
- AI Architecture Depth: Deep structural understanding of foundation model execution, training loops, optimization techniques, and the infrastructure demands of agentic workflows.
- Technology Expert: Mastery of low-level and systems programming (Rust, Go, or C++) alongside deep proficiency in Python and hybrid-cloud/bare-metal configurations.
Click on Apply to know more.