VIDA Digital Identity
Website:
vida.id
Company:
https://www.linkedin.com/company/vidadigitalid
Seniority: Mid-Senior level
Industries: Insurance and IT Services and IT Consulting
Job details:
About the Role
We are building the machine learning platform that powers real-time decisioning across our business. The platform's first customer is fraud, but the architecture is general: a real-time feature platform that serves both machine learning models and a rule engine from the same feature layer, with a clear path to powering our identity-verification models across the company.
You will own this end to end: streaming pipelines that compute features on live event data, an online feature store serving them at low latency, synchronization to the data lake with point-in-time accuracy for training, model serving on the decision path, and the rule engine that consumes the same features. This is a backend and ML infrastructure role — your customers are the data scientists, fraud analysts, and product teams who build on top of what you create.
What You Will Do
Feature platform (the foundation)
- Build real-time aggregation pipelines: streaming jobs computing windowed and lifetime aggregates over event streams — handling late and out-of-order data, backfills, and exactly-once state correctly.
- Build the online feature store: low-latency serving of features (velocity counters, device and identity aggregates, behavioral signals) on the synchronous decision path, with p99 latency targets under production traffic.
- Guarantee online/offline consistency: sync real-time features to the data lake with point-in-time accuracy, so training data matches exactly what the online system saw at decision time — no label leakage, no training-serving skew.
Decisioning layer (the consumers)
- Own model serving: deploy and operate ML models on the real-time decision path — inference services, feature-to-model plumbing, model versioning, shadow deployments, and rollback.
- Design and build the rule engine: a safe, expressive way for fraud ops to author, version, shadow-test, and roll out detection rules without engineering deploys — consuming the same feature layer as the models.
- Own reliability: the platform sits on the critical path of every transaction and onboarding decision. You will define SLOs, instrument the system, and design for graceful degradation.
- Partner closely with data scientists, fraud analysts, and ML engineers to shape the platform's APIs and abstractions around how they actually work.
What We Are Looking For
Must-have
- 5+ years of backend, data, or ML infrastructure engineering, with systems you built and operated in production at meaningful scale.
- Production stream-processing experience with Flink, Kafka Streams, Spark Structured Streaming, or equivalent — including the hard parts: watermarks, late/out-of-order events, stateful processing, exactly-once semantics, and backfills.
- Event streaming fluency with Kafka or a comparable log (Pulsar, Kinesis): partitioning, consumer group semantics, schema evolution, and replay.
- Low-latency serving experience: designing read paths on stores like Redis, Aerospike, DynamoDB, or ScyllaDB with tight p99 targets; hot-key handling, caching, and capacity planning.
- Understanding of the ML feature lifecycle: you can explain what point-in-time correctness means, why naive feature joins cause label leakage, and the trade-offs between feature logging and recomputation.
- Strong programming skills in Java, Scala, Kotlin, or Go for the streaming and serving layer; solid Python for the ML-facing surface.
- API and abstraction design taste: the feature store, model serving layer, and rule engine are products with internal users — you care about safe, well-versioned interfaces.
Nice-to-have
- Built or operated a feature store (Feast, Tecton, or an in-house equivalent).
- Model serving in production: inference services, A/B and shadow deployments, model registries (e.g., MLflow, Seldon, KServe, or in-house).
- Modern lakehouse table formats (Apache Iceberg, Delta Lake, or Hudi) and CDC pipelines.
- Fraud, risk, payments, or identity-verification domain experience.
- Experience designing DSLs, expression evaluators, or configuration-driven decision systems.
Why This Role
This is a rare greenfield: you define the ML platform architecture for a company whose core product runs on real-time decisions. Every improvement you ship directly reduces fraud losses and unlocks growth by letting good users through faster — and the platform you build becomes the foundation for machine learning across the business. You will work on a small team with direct access to the fraud, data, and product leaders who consume what you build.
Click on Apply to know more.