Ramp Infotech
Website:
ramp.agency
Company:
https://www.linkedin.com/company/ramp-infotech-ltd
Seniority: Entry level
Industries: IT System Operations and Maintenance
Job details:
Job Description
AI Engineer (LLM / RAG / Document Intelligence)
Role Summary
We are seeking an AI Engineer to build the AI core of a new B2B platform being delivered to a client. Client and project details are confidential and will be disclosed at onboarding under a confidentiality agreement.
The systems operate in a domain where an incorrectly extracted value is a serious defect: accuracy, grounding and controllability take priority over raw capability. The role covers production LLM engineering end-to-end: retrieval-augmented generation over a restricted document corpus with strict source boundaries, document and PDF data-extraction pipelines that normalize inconsistent real-world specifications, NLP classification pipelines, semantic search, and conversational intake that converts informal user language into precise structured data. The engineer works within a small senior delivery team - a Solution Architect who owns the technical design, and a Senior Full-Stack Developer who consumes the engineer's APIs - and demonstrates completed work in fortnightly sprint reviews with client stakeholders present.
EolasFlow is an AI-native engineering team: AI-assisted development (Claude Code, Cursor, GitHub Copilot or equivalent) is the standard working method, and candidates are expected to already work this way.
Key Responsibilities
• Design and build a production RAG system: chunking and embedding strategy, vector store, retrieval evaluation, citation-grounded answering, and strict source-boundary enforcement with refusal on out-of-bound queries
• Build document-intelligence pipelines: PDF and table extraction from inconsistent source documents, unit and format normalization, deduplication, and human-audit workflows
• Build NLP pipelines for content classification (signal vs noise), entity extraction and enrichment, and automated draft generation matched to a defined editorial voice
• Build semantic search mapping natural-language intent to structured capability data
• Build LLM-guided conversational intake converting informal language into precise structured specifications
• Establish evaluation discipline: evaluation sets and regression harnesses ahead of tuning, evaluations running in CI, quantified quality reporting
• Monitor and optimize cost, latency and quality across all LLM usage; make provider and model trade-offs explicit
• Expose all capabilities as clean, documented APIs for consumption by the application layer
• Present completed work in fortnightly sprint reviews
• Optimize high-volume pipeline tasks (classification, drafting) by fine-tuning and deploying small open-weight language models where they outperform API models on cost and latency
Required Skills and Experience
• Minimum 5 years building ML/NLP/LLM systems in production; strong Python (FastAPI or similar for serving)
• Production RAG experience: candidates must be able to walk through a shipped system — architecture, evaluation results, failure modes and remediation
• LLM engineering: prompt design, structured output (JSON schema / function calling), multi-provider model selection (OpenAI, Anthropic, open-weight models), cost and latency optimization
• Vector stores (pgvector, Qdrant, Pinecone or Weaviate); retrieval evaluation and hallucination control
• Document intelligence: PDF and table extraction from inconsistent real-world documents (Unstructured, Textract, Docling or custom pipelines)
• Evaluation discipline: builds evaluation sets and regression harnesses as standard practice and can quantify quality improvements
• Classic NLP fundamentals beyond prompting classification, named-entity recognition, entity resolution
• Data pipeline orchestration (Airflow, Prefect or similar); compliant API and web data ingestion (rate limiting, terms-of-service awareness)
• Optimize high-volume pipeline tasks (classification, drafting) by fine-tuning and deploying small open-weight language models where they outperform API models on cost and latency
• Experience building agentic pipelines (tool use, multi-step agents) in production
• Daily, fluent use of AI-assisted development tools (Claude Code, Cursor, GitHub Copilot or equivalent); this will be assessed through a live practical exercise during selection
• Fluent written and spoken English; able to present work to non-technical stakeholders
Desirable
• Small language model (SLM) fine-tuning: LoRA/QLoRA adaptation of open-weight models (Llama, Mistral, Phi or similar class) for classification and style/domain adaptation, including serving and deployment (vLLM, Ollama or similar) valued as a cost- and latency-optimization path for high-volume pipeline tasks
• Hybrid retrieval and re-ranking (BM25 combined with dense retrieval)
• Experience with technical or industrial specification data
• Content personalization or recommender systems
Click on Apply to know more.