nexocean
Company:
https://www.linkedin.com/company/nexocean
Industries: Business Consulting and Services
Job details:
Position Overview :
We are looking for a Senior Forward Deployed Engineer (FDE) who is passionate about operationalizing production AI-native and agentic workloads at scale. This is a high-impact role designed to serve as the technical tip of the spear for the company's most strategic AI-native customers and platform initiatives. As an FDE, you will operate at the intersection of Product Engineering, AI Infrastructure, and Customer Implementation. You will partner deeply with strategic AI-native enterprises (ANEs), startups, infrastructure vendors, and internal engineering teams to deploy, optimize, and scale production AI systems on the company's AI-native cloud platform. This role extends beyond traditional GPU infrastructure deployment. You will work across Inference Engines, runtime systems, orchestration frameworks, and AI-native applications to help customers operationalize production AI and agentic systems with a strong focus on scalability, reliability, latency, and workload economics. FDE engineers also act as the "first customer" for new AI-native platform capabilities. You will validate products under real-world workloads, surface operational insights and architectural gaps, and help accelerate product maturity through continuous feedback loops with Product Engineering and Research teams. You will build scalable deployment frameworks, benchmarking systems, automation tooling, and AI starter kits that transform field learning into reusable platform intelligence and repeatable deployment patterns across the company's AI ecosystem. Your mission is to accelerate production adoption of AI-native systems while helping shape the future of the company's AI-native cloud platform for the inference and agentic era.
What You'll Do :
Strategic AI Workload Operationalization Partner with strategic AI-native enterprises and AI startups to architect, deploy, optimize, and scale production AI and agentic systems. Support complex migrations, production-ready Proofs of Concept (PoCs), deployment acceleration, and long-term workload expansion across inference and runtime platforms.
AI Performance & Systems Engineering :
Optimize distributed inference and runtime performance through:
Benchmarking
GPU efficiency tuning
KV-cache optimization
Speculative decoding
Prefill/decode disaggregation
Multi-node deployments
Latency and cost optimization
Platform Validation & Product Acceleration :
Act as the "first customer" for new AI-native platform capabilities, including: Inference Engines Runtime systems Orchestration platforms GPU platforms Deployment workflows Surface operational insights, architectural gaps, and scaling bottlenecks directly to Product Engineering and Research teams.
Platform Intelligence & Automation Build scalable deployment assets, including:
Benchmarking systems
Automation tooling
AI starter kits
Deployment frameworks
Operational playbooks
Fine-tuning workflows
Reference architectures
These assets should improve deployment velocity and platform adoption across customers and internal teams.
Ecosystem & Technical Enablement
>Collaborate with GPU vendors, model providers, infrastructure partners, and ISVs on codevelopment, technical validation, optimization, and launch readiness.
>Enable customer-facing technical teams through validated deployment patterns, benchmarking insights, demos, operational playbooks, technical guidance, and reference architectures
Key Metrics
Customer Adoption & Production Success Measured by:
> Production AI workloads launched
>Reduction in time-to-production
>Pilot-to-production conversion Expansion of AI platform adoption
Platform Intelligence & Product Influence Measured by:
>Product improvements
>Roadmap influence
>Customer validation
>Operational insights from production deployments
Asset & Tooling Delivery
Measured through adoption of:
Deployment frameworks
Automation tooling
Benchmarking systems
Operational playbooks
Reference architectures
Field Enablement & Ecosystem Scale
Measured through:
Technical enablement
Ecosystem collaboration
Adoption of deployment standards across customers and partners
What You'll
Bring AI-Native Systems & Architecture:
Experience designing and operating production AI systems, including:
AI inference workloads
Agentic runtimes
AI orchestration frameworks
AI-native applications
Strong experience with inference and serving frameworks such as:
vLLM
SGLang
Ray Serve
NVIDIA
Dynamo
llm-d or equivalent technologies
Hands-on experience with LLM optimization techniques including:
Continuous batching Quantization KV-cache optimization Speculative decoding
Distributed Systems & Infrastructure
Deep expertise with NVIDIA and AMD GPU platforms and related software ecosystems, including:
CUDA
ROCm
TensorRT
Triton
NCCL
RCCL
NVLink
XGMI
RoCE
Strong proficiency in:
Kubernetes (K8s)
Distributed Systems
Networking Storage Systems
Infrastructure as Code (IaC)
Large-scale AI Infrastructure
Runtime & Orchestration Systems:
Experience with AI orchestration and agent frameworks such as:
LangGraph
CrewAI
MCP ecosystems
LlamaIndex
OpenAI Agents SDK
OR equivalent frameworks
Understanding of:
Workflow orchestration
Deployment systems
Memory patterns
AI-native application architectures
Software Engineering & Automation Strong production coding skills in:
Python /Go
Experience building: Automation tools ,Deployment workflows ,Benchmarking, frameworks Operational platforms
Preferred Qualifications
>4+ years of experience in Forward Deployed Engineering, ML Engineering, Applied AI Engineering, AI Infrastructure, Technical Consulting, or equivalent customer-facing engineering roles supporting production AI systems.
>Experience building deployment standards, technical enablement programs, platform adoption frameworks, or ecosystem integration strategies.
>Active contributor to open-source AI, infrastructure, orchestration, or developer tooling communities.
>Experience collaborating with GPU vendors, infrastructure providers, model vendors, or ecosystem partners on benchmarking, optimization, technical validation, or launch readiness initiatives.
Click on Apply to know more.