Source description
About the role
As a highly skilled Data Engineer with 5-9 years of experience, your role involves architecting enterprise-scale data platforms for smart metering/utility systems and multi-tenant SaaS applications. You should have deep expertise in modern data engineering patterns and real-time data ingestion, along with experience with Databricks Lakehouse architecture. Key Responsibilities: - Data Platform Architecture: - Design and implement scalable, multi-tenant data platforms following medallion architecture (Bronze/Silver/Gold layers) - Build loosely coupled, API-driven microservices for data ingestion, transformation, and serving - Ensure platform agnostic design supporting both on-premises and public cloud deployments (AWS, Azure, GCP) - Implement zero-trust security with RBAC, encryption at rest and in transit, and tenant isolation - Data Ingestion & Integration: - Build pluggable connector frameworks supporting REST, SOAP, GraphQL, file-based ingestion, database replication, and event-driven sources - Implement near real-time data pipelines for streaming meter data, events, and alarms - Handle multiple data sources with different schemas and formats (flat files, JSON, XML, database dumps) - Design adapters for multiple Head End Systems (HES) or vendor systems with seamless integration - Data Processing & Transformation: - Develop business rule engines for data validation using historical patterns (statistical analysis: deviation, average, median, standard deviation) - Implement data estimation algorithms for handling missing/incomplete data (5-25% gaps) - Build aggregation and virtual metering pipelines with arithmetic operations on interval data - Create billing determinant calculations and prepay billing processing systems - Real-Time & Batch Processing: - Design event-driven architectures with messaging queues (Kafka, RabbitMQ, AWS SQS) for alarm/event handling - Implement job scheduling for batch processing with monitoring and alerting - Build streaming pipelines for near real-time analytics and dashboard updates - Data Quality & Observability: - Implement data validation frameworks with configurable business rules - Build monitoring and alerting for data freshness, pipeline health, and SLA compliance - Create logging, tracing, and APM integrations for system health insights - Design reconciliation processes for data integrity verification - Data Modeling & Analytics: - Design canonical entity models normalizing common entities across heterogeneous sources - Build semantic layers with partner/tenant-specific views using SQL/dbt-like patterns - Create self-serve reporting and dashboard surfaces for business users Skills & Abilities: - Core Data Engineering: - 5+ years experience in data engineering with large-scale distributed systems - Strong proficiency in SQL and database design (PostgreSQL, MySQL, or similar) - Experience with ETL/ELT pipelines and data orchestration tools (Airflow, dbt, Prefect, Dagster) - Knowledge of data modeling principles (star schema, snowflake, dimensional modeling) - Databricks & Lakehouse: - 3+ years hands-on experience with Databricks platform - Strong expertise in Spark (PySpark/Scala) for distributed data processing - Experience with Delta Lake and ACID transactions on data lakes - Knowledge of Unity Catalog for data governance and fine-grained access control - Experience with Delta Live Tables (DLT) for pipeline orchestration - Proficiency in Databricks SQL for analytics and querying - Understanding of Databricks Workflows for job scheduling and orchestration - Experience with MLOps on Databricks (MLflow integration for model lifecycle) - Cloud & Infrastructure: - Hands-on experience with AWS, Azure, or GCP (preferably multi-cloud) - Experience with containerization (Docker, Kubernetes) - Knowledge of Infrastructure as Code (Terraform, CloudFormation) - Understanding of CI/CD pipelines (GitLab CI, Jenkins, GitHub Actions) - Streaming & Real-Time: - Experience with Apache Kafka or similar streaming platforms - Knowledge of event-driven architecture patterns - Familiarity with CDC (Change Data Capture) tools (Debezium, Airbyte) - Programming: - Strong proficiency in Python and/or Scala - Experience with REST API design and development - Knowledge of GraphQL is a plus - Data Storage: - Experience with S3/ADLS/GCS for object storage - Knowledge of data formats (Parquet, Avro, ORC, JSON) - Understanding of partitioning strategies for optimal query performance - Security & Compliance: - Experience implementing RBAC and ABAC for multi-tenant systems - Knowledge of encryption standards and secure data handling - Understanding of audit logging and compliance requirements Preferred Qualifications: - Experience in Utilities/Smart Metering/AMI domain knowledge - Familiarity with IEC CIM standards for utility data exchange - Experience with SCADA/GIS/MDMS integration - Kno
More at SourceFuse