zenteiq.ai
Website:
zenteiq.ai
Job details:
About ZenteiQ
ZenteiQ is a deep-tech company born out of IISc Bangalore, building Scientific Intelligence Infrastructure: physics-native AI for engineering, manufacturing, energy, mobility and national systems. Rather than wrapping general-purpose language models, we train foundation models (BrahmAI) from scratch to reason over thermal, electromagnetic, structural and materials domains, and put them to work through industrial platforms (KogneX) and a talent OS for engineers and researchers (AhamX). We're backed by the IndiaAI Mission (MeitY) and work closely with IISc, ARTPARK and a national AI Hub Network — building the sovereign AI infrastructure that India's engineering and industrial systems will run on.
About the Role
We are looking for a Performance Engineer, Inference to understand and improve the systems that serve our foundation models. Inference is a tightly coupled system spanning model execution, serving runtimes, distributed systems, accelerators, scheduling, memory and reliability. You will measure the system end to end, identify the highest-leverage performance gaps, and work across teams to close them while preserving correctness.
What You'll Do
- Run cross-layer performance investigations across throughput, latency, memory efficiency, reliability and cost.
- Build profiling, benchmarking and observability tooling that makes inference performance measurable and explainable.
- Identify bottlenecks across model servers, batching and scheduling, distributed execution, memory systems and accelerators.
- Partner with model, platform and infrastructure teams to prioritise and land high-impact optimisations.
- Validate that performance improvements preserve model quality and numerical correctness.
What We're Looking For
- Hands-on experience profiling and optimising ML systems or other performance-critical production systems.
- Production experience with at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, or an equivalent serving runtime.
- Experience serving or operating large models across accelerator-backed infrastructure, including multi-GPU or multi-accelerator systems.
- Strong Python skills and the ability to read, instrument and modify large production codebases.
- Solid understanding of transformer inference, distributed systems, latency/throughput trade-offs and accelerator performance fundamentals.
Good to Have / Bonus Points
- Experience with large-scale or multi-node inference, including tensor, pipeline, data or expert parallelism.
- Experience with GPUs, TPUs, NPUs or other ML accelerators and their associated profiling tools.
- Experience with quantization, low-precision inference, KV-cache optimisation, speculative decoding or long-context serving.
- Experience contributing to or modifying inference runtimes, kernels, compilers or distributed serving components.
- Experience optimising inference for constrained or on-device environments.
Why ZenteiQ
- Build real physics-native foundation models from scratch, not another LLM wrapper — deep, defensible technical work.
- Be part of a nationally recognised mission: one of 8 startups selected under the government's IndiaAI Mission to build a sovereign foundation model.
- Work alongside IISc-trained scientists and researchers, in a company founded by an IISc professor.
- See your work land in real industry pilots across automotive, mobility, defence and industrial R&D, with measurable impact.
- Join a lean, high-caliber team at an early, high-ownership stage.
Click on Apply to know more.