Oriserve
Website:
oriserve.com
Job details:
Job Profile:-
We are looking for a highly skilled, fast-paced, research-oriented Data Scientist with 2–3 years of experience, who has worked extensively with open-source Large Language Models (LLMs) — not just third-party APIs, but models loaded in-memory — and has solid experience working with GPUs. You will be responsible for building, fine-tuning, and deploying AI models to solve complex NLP challenges. The ideal candidate has hands-on experience with the core fundamentals of Deep Learning, has experimented with model layers, and has experience in optimizing and implementing LLM-based solutions in real-world applications.
Typical work week look like:-
- Deep AI research, fine-tuning, and deploying LLMs and other Transformer architecture-based models for various NLP tasks.
- Design and implement custom training pipelines for domain-specific models.
- Optimize model performance, including CUDA acceleration, quantization, and retrieval-augmented generation (RAG).
- Understand and implement the latest research papers across a wide variety of NLP domains, including speech, image, and video.
- Develop and maintain data pipelines for model training and evaluation.
- Collaborate with engineering teams to integrate AI models into production systems.
- Conduct experiments and analyze results to continuously improve model accuracy and efficiency.
- Stay up to date with cutting-edge advancements in Generative AI, LLMs, and NLP.
Our ideal candidate should have:-
- 2–3 years of hands-on experience in data science, with a focus on NLP and deep learning.
- Solid foundational knowledge of Deep Learning core concepts (gradient descent, learning rate schedulers, auto-regressive decoding, auto-encoders, etc.).
- Strong understanding of Transformer architectures (BERT, GPTs, T5, LLaMA, etc.).
- Experience in fine-tuning and deploying LLMs using PEFT, RLHF, Knowledge Distillation, or similar training frameworks.
- Proficiency with Hugging Face Transformers, vLLM, or similar frameworks.
- Strong programming skills in Python and experience with ML frameworks such as PyTorch or TensorFlow.
- Hands-on experience with writing custom model layers, not just using pre-built module imports.
- Experience with Speech AI models such as ASR (e.g., Whisper), TTS, Voice Cloning, Speech-to-Speech, Audio LLMs, and related technologies.
- Experience with CUDA accelerator engines such as ctranslate2, flash-attention, DeepSpeed, etc.
- Familiarity with vector databases (e.g., Pinecone, ChromaDB) and RAG pipelines.
- Knowledge of cloud platforms (AWS, GCP, Azure) and ML model deployment strategies.
- Understanding of MLOps practices, including model monitoring, request load balancing, WebSocket/API streaming, versioning, and CI/CD for ML workflows.
- Strong problem-solving skills and ability to work in a fast-paced environment.
What you can expect from ORI:-
- Opportunity to work on cutting-edge AI/ML projects.
- A collaborative, innovation-driven work environment.
- Passion & happiness in the workplace with great people & open culture with amazing growth opportunities.
- An ecosystem where leadership is fostered which builds an environment where everyone is free to take necessary actions to learn from real experiences.
- Freedom to pursue your ideas and tinker
If you are passionate about LLMs and pushing the boundaries of AI, we’d love to hear from you!
Click on Apply to know more.