RapidBrains
Website:
rapidbrains.com
Job details:
Job Title: Senior AI Data & Distillation Engineer
Experience: 3-5 Years
Location: Mumbai
Job Type: Full-time
Notice Period: Immediate Joiners preferred
We are hiring a Senior AI Data & Distillation Engineer to develop advanced AI model training pipelines and optimize knowledge transfer from frontier teacher models to student models. The role involves implementing model distillation techniques, building large-scale synthetic data pipelines, optimizing custom tokenizers for programming languages, and streaming massive datasets for efficient GPU training.
Responsibilities
• Implement multi-stage distillation objective functions using KL Divergence, logit extraction, and temperature scaling.
• Build high-throughput synthetic dataset generation pipelines using frontier LLM APIs.
• Create multi-turn, token-dense structural reasoning datasets for AI model training.
• Audit and customize tokenizer vocabularies to preserve critical programming syntax.
• Optimize tokenization for programming languages and legacy mainframe code structures.
• Use Ray or Apache Arrow to deduplicate, clean, and process large-scale linguistic datasets.
• Develop efficient data streaming pipelines to support continuous GPU training without data bottlenecks.
Skills:
- Strong understanding of ML Mathematics, probability distributions, cross-entropy, and KL Divergence.
- Experience with logit-level model distillation and temperature scaling.
- Hands-on experience with Ray, Apache Arrow, and Hugging Face (Tokenizers & Datasets).
- Expertise in Token Engineering and custom tokenizer optimization.
- Experience building large-scale synthetic dataset generation pipelines using LLMs.
- Knowledge of Java, .NET, Unix, and COBOL codebases is an advantage.
Key Skills: ML Mathematics, Data Scaffolding, Token Engineering, Ray, Apache Arrow, Hugging Face, Tokenizers, Datasets, KL Divergence, Cross-Entropy, Logit Distillation, Synthetic Data Generation, LLMs
Click on Apply to know more.