Sourcebae
Website:
sourcebae.com
Job details:
**SENIOR MDM DEVELOPER**
Role:** Senior MDM Developer
Domain:** Customer / Party MDM
Experience:** 10+ Years
Location : hybrid - Noida(sector3)/Bangalore(hsr layout)
Technology Stack:** Snowflake, dbt, PostgreSQL, Advanced SQL
Architecture:** Custom Warehouse-Native MDM Hub
Method:** Load → Match → Merge
Budget is Upto : 4 LPM
### Position Summary
We are seeking a hands-on Senior MDM Developer to build and enhance a custom, warehouse-native Customer MDM solution using Snowflake, dbt, SQL, and PostgreSQL.
The role will focus on Customer and Individual/Consumer mastering, including data modeling, source-system analysis, data loading, record matching, scoring, survivorship, reversible merges, golden-record creation, data quality, lineage, and auditability.
This is a custom MDM development role and does not involve configuring commercial MDM platforms.
### Key Responsibilities
**1. MDM Solution & Architecture**
* Review and enhance the MDM architecture, data models, source integrations, loading processes, and processing logic.
* Validate support for incremental loads, data quality, matching, merging, exception handling, lineage, auditability, and scalability.
* Identify implementation gaps, risks, dependencies, and improvement opportunities.
**2. Customer & Consumer Data Modeling**
* Design and maintain models for organizations, customers, individuals, consumers, households, accounts, affiliations, and hierarchies.
* Implement identifiers, relationships, lineage, source provenance, match evidence, survivorship decisions, and golden-record history.
* Develop governed party, role, relationship, and hierarchy models.
**3. Source-System Analysis & Mapping**
* Profile and analyze data from D&B, SAP, MTS, Sideline, DAL, and other source systems.
* Identify duplicate, incomplete, inconsistent, conflicting, stale, and invalid data.
* Develop source-to-MDM mappings, transformations, standardization, validations, reference translations, and rejection rules.
* Define how relevant source information should be incorporated into the MDM process.
**4. Snowflake, dbt & SQL Development**
* Develop and optimize Snowflake SQL, dbt models, scripts, procedures, views, and processing logic.
* Build initial, incremental, restartable, and controlled data-loading processes.
* Implement ingestion, staging, profiling, cleansing, standardization, validation, matching, scoring, merging, survivorship, golden-record creation, and publication workflows.
* Build reconciliation, audit, error-handling, recovery, and monitoring controls.
**5. Matching, Scoring & Entity Resolution**
* Analyze and improve deterministic, probabilistic, fuzzy, and composite matching rules.
* Develop blocking and candidate-generation strategies.
* Tune matching thresholds using labelled data or approved ground truth.
* Measure precision, recall, blocking recall, and match quality.
* Implement reason codes, match evidence, negative evidence, and veto rules.
* Apply privacy-aware matching practices for Individual/Consumer records and minimize over-merging risks.
**6. Survivorship, Merge & Golden Records**
* Implement attribute-level survivorship rules based on source trust, quality, completeness, verification, recency, and business precedence.
* Develop crosswalk-based and reversible merge/unmerge processing.
* Preserve immutable source records and source traceability.
* Maintain merge events, redirects, golden-record history, provenance, and rule versions.
* Ensure consent and communication preferences are governed appropriately.
**7. Iterative MDM Execution**
* Execute the complete lifecycle:
**Load → Standardize → Match → Merge → Review → Refine → Validate**
* Investigate duplicate clusters, incorrect merges, missed matches, conflicting values, and incomplete records.
* Refine mappings, standardization, thresholds, matching, and survivorship rules based on measured results.
* Perform regression testing after rule and model changes.
**8. Documentation & Collaboration**
* Document data models, mappings, data-quality rules, matching logic, thresholds, reason codes, survivorship rules, processing sequences, and technical decisions.
* Maintain reconciliation results, validation evidence, exception logs, and operational runbooks.
* Collaborate with data architects, data stewards, business analysts, governance teams, source-system owners, QA teams, and downstream consumers.
### Required Qualifications
* Bachelor’s degree in Computer Science, Information Systems, Engineering, Data Management, or related field.
* **7–12 years** of experience in data engineering, data integration, or enterprise data management.
* Strong hands-on experience with **Customer MDM or Identity Resolution**.
* Strong party and entity-relationship modeling skills.
* Hands-on experience with probabilistic record linkage, blocking, candidate generation, scoring, thresholds, negative evidence, and veto rules.
* Experience measuring matcher performance against labelled data or approved ground truth.
* Experience with privacy-aware person matching and consumer-scale identity resolution.
* Strong understanding of attribute-level survivorship, source provenance, consent handling, and rule governance.
* Experience with reversible, crosswalk-based merge/unmerge processing.
* Advanced SQL skills including window functions, recursive CTEs, MERGE patterns, and query optimization.
* Hands-on **dbt or equivalent transformation-as-code** experience.
* Strong data profiling, validation, reconciliation, and evidence-based problem-solving skills.
### Preferred Qualifications
* Hands-on experience with **Snowflake and PostgreSQL**.
* Experience with D&B and **D-U-N-S identifiers**.
* SAP customer-master experience.
* Python for data profiling, transformation, and match evaluation.
* Graph or network modeling experience.
* Experience with embeddings or vector similarity for entity resolution.
* Experience with data-stewardship workflows and human-in-the-loop review.
* Experience with diverse operational source systems.
### Expected Deliverables
* Reviewed MDM architecture and documented implementation gaps.
* Validated organization and individual party model.
* Source-data profiling and quality assessment.
* Approved source-to-MDM mappings.
* Production-ready Snowflake/dbt/PostgreSQL loading models.
* Measured and improved matching precision, recall, and blocking recall.
* Versioned survivorship rules with provenance and regression testing.
* Reversible crosswalk-based merge/unmerge process.
* Golden-record pipeline with lineage and merged-ID resolution.
* Golden-record validation and exception analysis.
* Technical documentation and operational runbook.
### Success Measures
* Improved blocking recall, match precision, and match recall.
* Privacy-safe person matching with controls for over-merging.
* Accurate, complete, unique, explainable, and traceable golden records.
* Reversible merges with immutable source records and complete crosswalks.
* Attribute-level provenance and defensible golden-value selection.
* Reconciled counts and validation across source, staging, match, merge, golden, and publish stages.
* Repeatable, restartable, monitored, version-controlled, and auditable MDM processing.
* Business and data-steward approval of matching, survivorship, hierarchy, and golden-record outcomes.
Click on Apply to know more.