Senior/ Principal AI Engineer
AI Infrastructure, MLOps & Backend Engineering
Location: Bengaluru
Department: Engineering
Employment Type: Full-time
About Nasiko
Nasiko is building the infrastructure enterprises need to run AI agents securely, efficiently, and reliably in production. Our platform combines TokenOps to understand and optimize model and agent spend, policy enforcement to govern how agents access models, tools, data, and credentials, and multi-agent execution to build and operate complex agent workflows. These capabilities are powered by a common foundation for agent and artifact management, model routing, AI and MCP gateways, observability, identity, and deployment—giving organizations a unified way to operate agentic systems across teams and environments. Nasiko is a Rust-first company, using Rust for our core platform services and Python for AI development, integrations, evaluation, and experimentation.
The Role
We are looking for a Senior/ Principal AI Engineer to build the backend and AI infrastructure behind Nasiko’s platform.
You will work across AI agents, LLM applications, model routing, tool integration, inference, evaluation, and production operations. You will help take AI systems from early prototypes to reliable, observable, and secure production services.
This is a hands-on engineering role for someone who understands both modern AI systems and strong software engineering. You should be comfortable building AI applications in Python, contributing to Rust-based platform services, and owning the deployment and operation of what you build.
What You’ll Work On
- Token usage, cost attribution, forecasting, and optimization systems
- Runtime policy enforcement for agents, models, tools, and credentials
- Multi-agent workflows, execution engines, retries, recovery, and human approvals
- LLM routing and provider-agnostic model access
- AI and MCP gateways for governed access to tools and enterprise systems
- Agent evaluation, monitoring, tracing, and production feedback loops
- Model and inference service integration
- Agent registries, artifact management, and lifecycle tooling
- Production infrastructure for deploying and operating AI services
What You’ll Own
AI and Agentic Engineering
- Design and build production-grade AI agentic infrastructure.
- Develop agent workflow infrastructure involving tools, memory, retrieval, planning, and multi-agent collaboration.
- Build integrations with hosted models, customer-hosted models, inference endpoints, and enterprise systems.
- Develop MCP server infrastructure, clients, tools, and gateway integrations.
- Design reliable patterns for structured output, tool invocation, retries, fallbacks, and failure handling.
- Improve agent quality through evaluation, tracing, experimentation, and production feedback.
Backend Engineering
- Build APIs, services, workers, and platform capabilities using Rust.
- Design clean abstractions for agents, models, tools, workflows, policies, and execution.
- Build event-driven and asynchronous services for long-running AI workloads.
- Work with PostgreSQL, Redis, queues, object storage, and vector databases.
- Write maintainable, well-tested code with clear interfaces and operational visibility.
- Contribute to architectural decisions across the Nasiko platform.
MLOps and Production AI
- Create repeatable workflows for packaging, deploying, versioning, and rolling back AI services.
- Build real-time and batch inference pipelines.
- Implement model and agent evaluation pipelines.
- Monitor quality, latency, errors, token usage, reliability, and cost.
- Manage model, prompt, configuration, and evaluation versions.
- Move prototypes into secure and observable production environments.
DevOps and Reliability
- Own services through development, deployment, monitoring, and production support.
- Build and improve CI/CD pipelines and release automation.
- Deploy and operate services using Docker and Kubernetes.
- Implement health checks, autoscaling, alerting, graceful failure, and rollback strategies.
- Manage cloud infrastructure and infrastructure-as-code.
- Apply strong practices around secrets, access control, service identity, and tenant isolation.
What We’re Looking For
- 10+ years of professional software, AI, backend, or platform engineering experience.
- Strong proficiency in Python/Rust for production AI and backend systems.
- Experience building and deploying LLM-powered applications or AI agents.
- Experience with tool-calling, RAG, embeddings, vector search, evaluation, or agent workflows.
- Strong backend engineering fundamentals, including APIs, asynchronous processing, databases, and service design.
- Production experience with Docker, Kubernetes, cloud infrastructure, and CI/CD.
- Experience operating AI or backend services using logs, metrics, traces, and alerts.
- Familiarity with inference services, model APIs, or model-serving infrastructure.
- Ability to independently take a capability from design through production.
- Strong technical judgment and clear written and verbal communication.
- Interest in working with Rust and contributing to a Rust-based platform.
Strongly Preferred Domain Experience
We will strongly prioritize candidates who have directly built products or infrastructure in one or more of the following areas:
- LLM routers or model gateways
- AI gateways or MCP gateways
- Agent runtimes and orchestration platforms
- Multi-agent frameworks or workflow engines
- Model-serving or inference platforms
- LLM observability and evaluation platforms
- AI security, governance, or policy-enforcement systems
- AI FinOps, token analytics, or model-cost optimization
- Agent registries, tool registries, or AI developer platforms
Experience working at an AI infrastructure, developer tooling, model platform, or agent platform company will be a significant advantage—not merely a bonus.
Relevant Technologies
You do not need experience with every technology listed below.
Languages: Python, Rust
AI: PyTorch, Hugging Face, OpenAI-compatible APIs, structured generation
Agents: MCP, A2A, LangGraph, AutoGen, Agno, or similar frameworks
Inference: vLLM, Triton, ONNX Runtime, Ray Serve, hosted model providers
Data: PostgreSQL, Redis, Kafka or RabbitMQ, object storage, vector databases
Infrastructure: Docker, Kubernetes, Helm, Terraform or Pulumi
Observability: OpenTelemetry, Prometheus, Grafana, distributed tracing
MLOps: MLflow, Weights & Biases, evaluation pipelines, model and prompt registries
Additional Advantages
- Production experience with Rust.
- Experience operating GPU-backed inference workloads.
- Familiarity with semantic routing, knowledge graphs, or advanced retrieval architectures.
- Experience with multi-tenant SaaS or customer-managed deployments.
- Knowledge of authentication, authorization, policy engines, and secrets management.
- Contributions to open-source AI, Rust, Kubernetes, or developer infrastructure projects.
- Public examples of systems, libraries, research, or technical writing you have produced.
What Success Looks Like
Within your first six months, you will:
- Own and ship a meaningful AI or platform capability.
- Move an AI workflow from prototype to reliable production deployment.
- Improve the quality, observability, or operability of agent execution.
- Contribute to Nasiko’s Rust-based backend platform.
- Establish stronger patterns for AI evaluation, deployment, and production monitoring.
- Become a technical owner across product, AI, and engineering decisions.
Why Nasiko
- Build foundational AI infrastructure: Work on the systems enterprises need to operate agents in production.
- Solve emerging problems: TokenOps, policy enforcement, and multi-agent execution are still being defined as categories.
- Work across AI and systems engineering: Build agents and AI applications while contributing to the platform beneath them.
- Shape a Rust-based company: Help define the architecture and engineering practices of a modern infrastructure platform.
- Own meaningful outcomes: Take capabilities from initial design through production.
- Build in the open: Contribute to open-source projects and the broader AI developer ecosystem.
- Join at an early stage: Have significant influence over the product, technology, and engineering culture.