Axiom Global Technologies
Website:
axiomglobal.com
Company:
https://www.linkedin.com/company/axiomglobaltechnologies
Industries: IT Services and IT Consulting
Job details:
AI/ML & Knowledge Platform Engineer
Retrieval, knowledge modeling, and grounded generation — Full-time
About the rol
eWe are hiring a founding-level platform engineer to build the trusted data and retrieval plane behind our AI products. You will own how unstructured enterprise content and structured records are ingested, modeled, resolved into a canonical knowledge layer, and served to LLM-driven services with full provenance. The central engineering problem is making generated outputs traceable to the exact evidence that supports them, at production quality
.This is a builder role, not a research-only or dashboard role. You will write production pipelines, design canonical schemas, implement entity resolution, and stand up hybrid retrieval and grounded-generation evaluation that other engineers and AI services depend on
.
What you will work on
-
Ingestion and pipeline engineer
- ingBuild connectors for APIs, databases, event feeds, secure file transfer, customer exports, and document repositori
- es.Ingest and normalize PDF, DOCX, PPTX, XLSX, email, and structured records while preserving source metadata, version, access restrictions, and licensing boundari
- es.Design idempotent, restartable pipelines with schema evolution, backfills, retries, dead-letter handling, reconciliation, and observabili
ty.
Canonical models, ontology, and entity resolu
- tionOwn canonical schemas and data contracts across organization, workforce, requirement, solution, pricing, and evidence doma
- ins.Resolve organizations, people, roles, titles, skills, technologies, and identifiers across inconsistent sources while preserving match confidence and reviewabil
- ity.Model time and change: current versus historical attributes, effective dates, revisions, and superseded records using durable IDs and crosswalk tab
les.
Retrieval, knowledge, and provenance ser
- vicesBuild the document/record processing layer for chunking, metadata enrichment, embeddings, and retri
- eval.Support hybrid retrieval across structured filters, full text, vectors, and relationships while enforcing tenant and document-level permiss
- ions.Expose facts, evidence, confidence, and lineage to downstream AI services without coupling them to raw source sch
emas.
Applied LLM quality and eval
- uationBuild grounded-generation datasets, retrieval evaluation (recall/precision, citation coverage), and hallucination/error ana
- lysis.Stand up replayable evaluation corpora and regression tests so model and pipeline changes are measured, not guess
- ed at.Work with self-hosted and API LLMs behind a routing/proxy layer; reason about cost, latency, and where a smaller local model is suffi
- cient.Expose data quality, freshness, failed ingestion, duplicate/match confidence, and provenance coverage as operational me
trics.
What we are lookin
g for -R
- equiredProduction data/backend/search systems: 3+ years building data, backend, search, or knowledge systems with senior ownership of architecture and operations, not ticket-lev
- el ETL.Strong Python and SQL: expert relational modeling plus practical experience with object storage, queues/events, pipeline orchestration, APIs, and cloud depl
- oyment.Entity resolution and modeling: hands-on deduplication, taxonomy/ontology design, slowly-changing/temporal data, schema evolution, lineage, and source reconcil
- iation.Retrieval systems in production: full text, vectors, metadata filters, ranking, and reranking; comfort with graph-shaped models where the
- y help.Data-quality discipline: ability to instrument pipelines, debug silent corruption, perform safe backfills, and reject a convenient pipeline that creates an untraceable data p
- roduct.Azure, Kubernetes, PostgreSQL/pgvector, OpenSearch/Elasticsearch, a graph database, dbt, and an orchestrator (Dagster/Airflow/Temporal); depth in a coherent subset matters more than b
- readth.Evaluation and observability for RAG or agentic systems: groundedness, retrieval recall/precision, citation coverage, and regression t
- esting.Experience fine-tuning, distilling, or serving open-weight models (vLLM or similar) and routing between local and hosted
models.
P
- referredEarly-stage instinct: ship a correct v1, document why, and improve it without waiting for a platform team to
- appear.Unstructured document processing: extracting and serving information from enterprise documents — metadata, section structure, tables, revisions, and ci
tations.
Not the righ
- t profileA warehouse/BI specialist whose primary output is dashboards and
- reports.A research-only knowledge-graph or NLP profile that has not operated pipelines in pr
- oduction.A prompt engineer who treats ingestion, identity, provenance, and quality as someone else’s
problem.
How w
e evaluateExpect a working session: given a sample document set, a structured feed, and an export with inconsistent organizations and roles, design the canonical model, source lineage, entity resolution, ingestion/retry strategy, permission model, and a service that lets an AI layer generate a recommendation with citations — with explicit quality metrics and a plan for a source that silently changes i
ts schema.
Click on Apply to know more.