ET Robotics
Company:
https://www.linkedin.com/company/et-robotics
Industries: Robotics Engineering
Job details:
Team AI / Foundation Models
Location On-site
Type Full-time
Level Open — Mid-level through Principal (scoped to experience)
About ET Robotics
ET Robotics is building general-purpose humanoid robots that can perceive, reason, and act over long horizons in the physical world. We pair frontier AI research with real hardware — our robots learn from large-scale multimodal data and are deployed on physical platforms today. We are assembling a small, high-caliber team where research and deployment happen under one roof, and where the models you train run on real robots within weeks, not years.
The Role
We are looking for an AI / VLM Engineer who deeply understands the architecture of Vision-Language Models (VLMs) and Large Language Models (LLMs), and who has hands-on expertise pre-training and post-training these models. You will design and train the foundation models that give our humanoids long-horizon intelligence — the ability to understand a scene, reason about a multi-step goal, and translate that into grounded action over extended time horizons.
This is a build role for someone who wants to own model architecture and the full training stack, from data to pre-training to alignment and post-training, and see it deployed on physical robots. The level and scope of the role will be calibrated to your experience, from strong hands-on mid-level engineers to principal researchers who set the research agenda.
What You'll Do
• Design and train VLM / LLM architectures that serve as the reasoning and perception backbone for humanoid robots, including vision-language-action (VLA) formulations that connect perception to control.
• Own pre-training of multimodal foundation models — data curation and mixing, tokenization, objective design, scaling, and distributed training on large GPU clusters.
• Lead post-training including supervised fine-tuning, instruction tuning, Retrieval-Augmented Generation, RLHF / RLAIF and preference optimization (e.g., DPO), and reward modeling to align model behavior with real-world task execution.
• Build long-horizon capability — hierarchical planning, memory, chain-of-thought / reasoning over multi-step goals, and mechanisms that keep a robot coherent across extended action sequences.
• Ground models in embodiment by integrating proprioception, action tokens, and closed-loop feedback so language and vision translate into reliable physical behavior.
• Drive evaluation by designing benchmarks and metrics for reasoning, grounding, and long-horizon task success on both offline data and live robots.
• Optimize for deployment — distillation, quantization, and inference efficiency so large models run within the latency and compute budgets of onboard hardware.
• Collaborate across the stack with perception, controls, and hardware teams to close the loop between the model and the robot.
What We're Looking For
• Deep understanding of VLM / LLM architecture — transformers, attention variants, multimodal fusion, tokenizers, positional schemes, and the design trade-offs that matter at scale.
• Hands-on pre-training and post-training experience with VLMs or LLMs — you have trained or substantially contributed to training such models, not only fine-tuned APIs. Experience with Retrieval-Augmented Generation and its current state of the art derivatives.
• Strong applied ML engineering in PyTorch (or JAX), with experience in large-scale distributed training (FSDP / DeepSpeed / Megatron-style parallelism) and mixed-precision training. Experience with Agentic AI tools like Langchain/ Langgraphs.
• Solid ML and math fundamentals — optimization, linear algebra, measure theory probability, and deep learning theory sufficient to reason about training dynamics and debug at scale.
• A track record of shipping — research that turned into working systems, strong open-source contributions, or publications at venues such as NeurIPS, ICML, ICLR, CVPR, CoRL, or RSS.
• Ownership and pragmatism — comfort operating with ambiguity in a fast-moving team and taking a model from idea to running on hardware.
Bonus Points
• Experience with Vision-Language-Action (VLA) models or robot foundation models (e.g., RT-2, OpenVLA, π0, GR00T-style systems).
• Background in robot learning, imitation learning, reinforcement learning, or long-horizon planning.
• Experience with reasoning models, agentic frameworks, world models, or memory architectures for extended-horizon tasks.
• Familiarity with real-robot data pipelines, teleoperation data, or sim-to-real transfer.
• Experience deploying models under tight latency and compute constraints on embedded or edge hardware.
Click on Apply to know more.