Source description
About the role
Job Title: Freelance- Kafka Expert Were hiring a Senior Kafka Data Engineer to join a growing data platform team building production-grade lakehouse solutions. Youll design and deliver scalable, observable, and costefficient data pipelines that power analytics and ML products across the business. This role suits handson engineers who have built streaming and batch lakehouse solutions on Azure Databricks and worked extensively with Kafka and event-driven architectures. What youll do Design, build, and maintain production-grade data pipelines (batch, streaming, CDC, API) on Azure Databricks using Python and Spark/PySpark. Implement and operate Bronze/Silver/Gold lakehouse architectures, including transformations, partitioning, incremental processing, and data lifecycle management. Integrate Kafka and other event-streaming platforms into data ingestion and propagation patterns; design event-driven data flows and schemas. Build governed data products and models for analytics and warehousing; collaborate with data scientists, analysts, and product teams. Ensure data quality, validation, observability, monitoring, and alerting across pipelines. Optimize pipeline performance, scalability, and cost on cloud infrastructure. Implement CI/CD for data engineering workflows (GitHub-based), and deploy infrastructure with Infrastructure-as-Code (Terraform). Drive best practices for production readiness: fault tolerance, replayability, schema evolution, and operational runbooks. Must-have skills & experience 710 years of experience in data engineering, with demonstrable production experience (not just POCs). Strong hands-on experience with Azure Databricks and Spark / PySpark. Deep experience with Kafka (event-driven architecture, producers/consumers, partitions, offsets, schema registries, exactly-once/at-least-once patterns). Proven track record implementing batch, streaming, CDC, and API-based ingestion. Experience designing Bronze/Silver/Gold lakehouse architectures and data transformations. Solid data modeling and warehousing knowledge; experience building governed data products. CI/CD experience with GitHub workflows for data pipelines. Experience with Terraform (or similar IaC) to provision cloud resources. Expertise in data quality, validation frameworks, observability (metrics, logging, tracing), and alerting. Strong focus on performance tuning, scalability, and cost optimization in cloud environments. Excellent communication skills; experience working cross-functionally with engineering, analytics, and product teams. Nice-to-have Familiarity with schema registry tools (e.g., Confluent Schema Registry, Apicurio). Experience with Databricks Delta Lake, Unity Catalog, or similar governance tooling. Knowledge of event sourcing, stream processing frameworks (Kafka Streams, ksqlDB, Flink), or message brokers other than Kafka. Experience working in highly regulated or security-conscious environments. Why join Work on high-impact data infrastructure that supports analytics and ML. Collaborate with experienced engineers and cross-functional product teams. Opportunity to influence architecture and operational standards at scale. Competitive compensation and benefits. Job Title: Freelance- Kafka Expert Were hiring a Senior Kafka Data Engineer to join a growing data platform team building production-grade lakehouse solutions. Youll design and deliver scalable, observable, and costefficient data pipelines that power analytics and ML products across the business. This role suits handson engineers who have built streaming and batch lakehouse solutions on Azure Databricks and worked extensively with Kafka and event-driven architectures. What youll do Design, build, and maintain production-grade data pipelines (batch, streaming, CDC, API) on Azure Databricks using Python and Spark/PySpark. Implement and operate Bronze/Silver/Gold lakehouse architectures, including transformations, partitioning, incremental processing, and data lifecycle management. Integrate Kafka and other event-streaming platforms into data ingestion and propagation patterns; design event-driven data flows and schemas. Build governed data products and models for analytics and warehousing; collaborate with data scientists, analysts, and product teams. Ensure data quality, validation, observability, monitoring, and alerting across pipelines. Optimize pipeline performance, scalability, and cost on cloud infrastructure. Implement CI/CD for data engineering workflows (GitHub-based), and deploy infrastructure with Infrastructure-as-Code (Terraform). Drive best practices for production readiness: fault tolerance, replayability, schema evolution, and operational runbooks. Must-have skills & experience 710 years of experience in data engineering, with demonstrable production experience (not just POCs). Strong hands-on experience with Azure Databricks and Spark / PySpark. Deep experience with Kafka (event-driven architecture, producers/consu