SourceFuse
Website:
sourcefuse.com
Job details:
We are seeking a highly skilled Data Engineer to architect with 5-9 Years of experience who can build enterprise-scale data platforms for smart metering/utility systems and multi-tenant SaaS applications. The ideal candidate will have deep expertise in modern data engineering patterns, real-time data ingestion, and experience with Databricks Lakehouse architecture.
Key Responsibilities:
● Data Platform Architecture:
○ Design and implement scalable, multi-tenant data platforms following medallion architecture (Bronze/Silver/Gold layers)
○ Build loosely coupled, API-driven microservices for data ingestion, transformation, and serving
○ Ensure platform agnostic design supporting both on-premises and public cloud deployments (AWS, Azure, GCP)
○ Implement zero-trust security with RBAC, encryption at rest and in transit, and tenant isolation
● Data Ingestion & Integration:
○ Build pluggable connector frameworks supporting REST, SOAP, GraphQL, file-based ingestion, database replication, and event-driven sources
○ Implement near real-time data pipelines for streaming meter data, events, and alarms
○ Handle multiple data sources with different schemas and formats (flat files, JSON, XML, database dumps)
○ Design adapters for multiple Head End Systems (HES) or vendor systems with seamless integration
● Data Processing & Transformation:
○ Develop business rule engines for data validation using historical patterns (statistical analysis: deviation, average,
median, standard deviation)
○ Implement data estimation algorithms for handling missing/incomplete data (5-25% gaps)
○ Build aggregation and virtual metering pipelines with arithmetic operations on interval data
○ Create billing determinant calculations and prepay billing processing systems.
● Real-Time & Batch Processing:
○ Design event-driven architectures with messaging queues (Kafka, RabbitMQ, AWS SQS) for alarm/event handling
○ Implement job scheduling for batch processing with monitoring and alerting
○ Build streaming pipelines for near real-time analytics and dashboard updates
● Data Quality & Observability
○ Implement data validation frameworks with configurable business rules
○ Build monitoring and alerting for data freshness, pipeline health, and SLA compliance
○ Create logging, tracing, and APM integrations for system health insights
○ Design reconciliation processes for data integrity verification
● Data Modeling & Analytics
○ Design canonical entity models normalizing common entities across heterogeneous sources
○ Build semantic layers with partner/tenant-specific views using SQL/dbt-like patterns
○ Create self-serve reporting and dashboard surfaces for business users
Skills & Abilities:
● Core Data Engineering:
○ 5+ years experience in data engineering with large-scale distributed systems
○ Strong proficiency in SQL and database design (PostgreSQL, MySQL, or similar)
○ Experience with ETL/ELT pipelines and data orchestration tools (Airflow, dbt, Prefect, Dagster)
○ Knowledge of data modeling principles (star schema, snowflake, dimensional modeling)
● Databricks & Lakehouse
○ 3+ years hands-on experience with Databricks platform
○ Strong expertise in Spark (PySpark/Scala) for distributed data processing
○ Experience with Delta Lake and ACID transactions on data lakes
○ Knowledge of Unity Catalog for data governance and fine-grained access control
○ Experience with Delta Live Tables (DLT) for pipeline orchestration
○ Proficiency in Databricks SQL for analytics and querying
○ Understanding of Databricks Workflows for job scheduling and orchestration
○ Experience with MLOps on Databricks (MLflow integration for model lifecycle)
● Cloud & Infrastructure
○ Hands-on experience with AWS, Azure, or GCP (preferably multi-cloud)
○ Experience with containerization (Docker, Kubernetes)
○ Knowledge of Infrastructure as Code (Terraform, CloudFormation)
○ Understanding of CI/CD pipelines (GitLab CI, Jenkins, GitHub Actions)
Streaming & Real-Time:
○ Experience with Apache Kafka or similar streaming platforms
○ Knowledge of event-driven architecture patterns
○ Familiarity with CDC (Change Data Capture) tools (Debezium, Airbyte)
● Programming
○ Strong proficiency in Python and/or Scala
○ Experience with REST API design and development
○ Knowledge of GraphQL is a plus
● Data Storage
○ Experience with S3/ADLS/GCS for object storage
○ Knowledge of data formats (Parquet, Avro, ORC, JSON)
○ Understanding of partitioning strategies for optimal query performance
● Security & Compliance
○ Experience implementing RBAC and ABAC for multi-tenant systems
○ Knowledge of encryption standards and secure data handling
○ Understanding of audit logging and compliance requirements
Preferred Qualifications:
● Experience in Utilities/Smart Metering/AMI domain knowledge
● Familiarity with IEC CIM standards for utility data exchange
● Experience with SCADA/GIS/MDMS integration
● Knowledge of Commission engines or direct selling industry (bonus)
● Experience with time-series databases (InfluxDB, TimescaleDB)
● Understanding of graph databases for genealogy/network data
● Experience with SaaS multi-tenant architecture patterns
● Knowledge of API gateway solutions (Kong, AWS API Gateway)
● Familiarity with service mesh patterns
Click on Apply to know more.