QualityKiosk Technologies Pvt. Ltd.
Website:
qualitykiosk.com
Job details:
HIRING ALERT! AI QA ENGINEER !
Role: AI QA Engineer
Work Location: Navi Mumbai/ Bangalore
Experience: 3-6 Years
Preferred Immediate Joiner.
Applicants please share your updated resumes on tanvi.palwankar@qualitykiosk.com
Role Overview
We are looking for an AI QA Engineer with hands-on experience working on GenAI platforms and workflows, particularly AI observability, evaluation, testing, or safety platforms similar in nature to Arize, Cekura, Fiddler, Giskard, WhyLabs, Phoenix, or equivalent systems.
This role requires direct exposure to GenAI bots, RAG pipelines, and AI workflows in real platforms, not theoretical knowledge or generic QA experience. You will evaluate AI behavior, analyze failures, and contribute to improving the quality, reliability, and safety of GenAI systems.
This is not a traditional QA role and not suitable for candidates who have only tested UI, APIs, or deterministic systems.
Key Responsibilities
- Execute hands-on AI evaluation and QA on:
- GenAI bots (chatbots, copilots, agents)
- RAG-based workflows
- AI evaluation and observability platforms
- Design and execute:
- Prompt-based test scenarios
- Multi-turn conversation tests
- Edge-case and adversarial evaluations
- Evaluate AI behavior for:
- Hallucinations
- Grounding failures
- Inconsistency and drift
- Safety and reliability issues
- Work within existing AI platforms (evaluation, observability, monitoring) to:
- Analyze traces, metrics, and outputs
- Interpret model behavior across runs
- Document findings clearly in:
- Evaluation reports
- Platform dashboards
- Structured test summaries
- Collaborate with:
- AI QA Lead on methodology and review
- ML and Product teams to provide actionable feedback
Must-Have Skills & Experience
- 3–5+ years of experience in AI QA, AI evaluation, or GenAI testing
- Hands-on experience working on GenAI platforms, such as:
- AI evaluation platforms
- AI observability or monitoring tools
- GenAI safety or reliability platforms
- Direct experience testing:
- GenAI chatbots or agents
- RAG pipelines and workflows
- Multi-step AI interactions
- Strong understanding of:
- Prompt-based testing
- Non-deterministic AI behavior
- Quality vs correctness in GenAI systems
- Ability to clearly explain why an AI response is good or bad
Good-to-Have
- Experience with platforms like Arize, Cekura, Fiddler, Giskard, Phoenix, WhyLabs, or similar
- Familiarity with evaluation metrics and scoring approaches
- Exposure to AI safety, adversarial testing, or red-teaming
- Basic scripting or data analysis skills
Click on Apply to know more.