Website:
halcer.com
Job details:
Job Title: Gen AI QA Engineer
Location: Bangalore, India (On-site)
Experience Level: 5–8 Years
Role SummaryWe are seeking an experienced Gen AI QA Engineer to lead the quality assurance lifecycle for next-generation generative AI and agentic AI systems. In this role, you will be responsible for ensuring the reliability, safety, functional accuracy, and performance of large language model (LLM) applications across the entire Software Development Life Cycle (SDLC). You will bridge traditional automation and non-deterministic AI validation, building robust test pipelines that safeguard model outputs against bias, hallucinations, and safety violations.
Key Responsibilities- AI Test Strategy & Execution: Design, implement, and maintain end-to-end test strategies tailored to GenAI, Retrieval-Augmented Generation (RAG) architectures, and agentic workflows.
- LLM & Agentic Output Validation: Perform rigorous evaluation of model responses to detect and mitigate hallucinations, bias, toxicity, alignment issues, and non-deterministic drift.
- Prompt & Model Regression Testing: Systematically evaluate prompt iterations, model fine-tuning impact, and version changes to prevent regression across software releases.
- Automation & CI/CD Integration: Build and maintain scalable test automation suites using modern frameworks (Playwright, Selenium) and integrate automated AI quality gates directly into CI/CD pipelines.
- API & Integration Testing: Validate end-to-end data flows, tool bindings, and REST/gRPC API interfaces connecting front-end applications, orchestrators (e.g., LangChain/LlamaIndex), and backend AI services.
- Dataset Management: Curate, synthetic-generate, and maintain golden evaluation datasets to measure baseline AI precision, recall, and safety metrics across varied user edge cases.
- Cross-Functional Collaboration: Partner closely with AI/ML developers, prompt engineers, and product managers to define clear acceptance criteria, benchmark safety thresholds, and enforce quality gates.
Required Experience & Qualifications- Core QA Expertise: 5 to 8 years of hands-on software testing experience spanning both manual exploratory techniques and automated testing paradigms.
- GenAI / LLM Testing: Proven track record of evaluating LLM-powered applications, prompt behavior, non-deterministic outputs, and responsible AI safety boundaries.
- Automation Frameworks: Proficient in automated UI and workflow testing using Playwright, Selenium, or modern JavaScript/Python-based automation frameworks.
- API Testing: Demonstrated experience testing microservices and APIs using tools like Postman, REST Assured, or PyTest.
- CI/CD & DevOps: Practical experience embedding automated quality gates into CI/CD workflows (Jenkins, GitHub Actions, Azure DevOps, or GitLab CI).
- Scripting Skills: Strong coding ability in Python or JavaScript/TypeScript for writing test automation scripts and data validation routines.
Preferred / Nice-to-Have Skills- Familiarity with specialized LLM evaluation frameworks (e.g., Ragas, DeepEval, TruLens, or Promptfoo).
- Hands-on exposure to RAG pipeline evaluation metrics (Faithfulness, Answer Relevance, Context Recall).
- Basic understanding of security testing for LLMs (e.g., prompt injection prevention, jailbreak testing).
Click on Apply to know more.