Ubique Systems
Website:
ubique-systems.com
Job details:
AI Testing Specialist – LLM Evaluation & Quality
Location: Hyderabad (HYD)
Contract Duration: 45 Days
Role: AI Testing Specialist
Experience: Relevant experience in AI/ML testing, LLM evaluation, test automation, or AI quality engineering
Role Overview
We are looking for an AI Testing Specialist with hands-on experience in LLM evaluation, AI/ML testing, test automation, and agentic AI quality assurance. The ideal candidate will build automated evaluation and regression frameworks for LLM-powered agents, identify adversarial and security-related failure modes, and integrate AI quality gates into CI/CD pipelines.
The role will work closely with AI/ML Engineers and DevSecOps teams in Germany and Hyderabad to establish robust testing standards for production-grade, LangGraph-based AI agents.
Key Responsibilities
- Design and implement LLM evaluation frameworks using RAGAS, custom evaluation metrics/scorers, and automated accuracy benchmarking.
- Validate AI/agent outputs against structured specifications, expected outcomes, and Gherkin/BDD acceptance criteria.
- Develop adversarial and red-team testing suites to identify hallucinations, prompt injection, jailbreaks, unsafe outputs, data leakage, and edge-case failures.
- Test AI agents across Refinement, Decision, and Coding workflows within a multi-agent architecture.
- Build and maintain automated regression test suites to validate agent behaviour across:
- Foundation model/LLM upgrades
- Prompt changes
- Skill-file modifications
- Retrieval changes
- LangGraph workflow changes
- Develop Python-based test automation using pytest and related testing frameworks.
- Integrate AI evaluation and testing into CI/CD pipelines, preferably using GitHub Actions.
- Implement automated quality gates for accuracy, reliability, fairness, security, and regression detection before production deployment.
- Define and monitor quantitative AI quality metrics, including:
- Response accuracy
- Faithfulness
- Context/retrieval relevance
- Retrieval precision/recall
- Hallucination rate
- Confidence scores and thresholds
- Response latency
- Robustness
- Model/prompt drift
- Build or contribute to AI quality dashboards and evaluation reports and communicate actionable findings to engineering teams.
- Test RAG pipelines, vector search/retrieval, embeddings, and agent orchestration.
- Collaborate with AI/ML Engineering, DevSecOps, and software engineering teams to establish AI testing standards and release criteria.
- Contribute to compliance-oriented testing aligned with EU AI Act, DORA, AI governance, traceability, reliability, and risk-management requirements.
- Document test strategies, evaluation methodologies, defects, quality metrics, and release-readiness criteria.
Required Skills & Experience
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
- Hands-on experience testing AI/ML systems, LLM applications, GenAI applications, or AI agents.
- Strong programming skills in Python.
- Strong experience with pytest, test automation, scripting, and automated evaluation pipelines.
- Practical experience with LLM evaluation, including RAGAS or comparable evaluation frameworks.
- Experience creating custom evaluation/scoring functions and benchmarking methodologies.
- Experience with LLM red teaming, adversarial testing, prompt injection, hallucination detection, and edge-case testing.
- Strong understanding of CI/CD and automated quality gates, preferably GitHub Actions.
- Familiarity with LangGraph, LangChain, or similar agentic AI frameworks.
- Understanding of RAG architectures, retrieval evaluation, embeddings, vector databases, and multi-agent orchestration.
- Ability to translate functional requirements and Gherkin/BDD acceptance criteria into automated AI test cases.
- Strong understanding of software testing concepts including functional, regression, integration, system, API, negative, exploratory, and performance testing.
Good to Have
- Experience with EU AI Act / AI governance / AI risk management.
- Knowledge of DORA and technology-risk/compliance testing.
- Experience validating vector databases, embeddings, semantic search, and retrieval pipelines.
- Experience with Docker and Kubernetes.
- Exposure to cloud-based AI/ML platforms.
- Experience with AI observability, model monitoring, evaluation dashboards, or drift detection.
- Experience testing MCP, tool-calling agents, autonomous/agentic workflows, or AI coding agents.
- Familiarity with security testing methodologies such as OWASP LLM Top 10.
Click on Apply to know more.