Padmi

Data Engineer-Databricks, Python) (Chennai)

ChennaiPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at CG VAK Software & Exports

Opens the source posting on shine.com

Source description

About the role

View original

Role & Responsibilities Key Responsibilities Architect and implement enterprise-grade Lakehouse solutions using Databricks Design and deliver scalable batch and real-time data pipelines using Apache Spark (PySpark/SQL) Build ETL/ELT pipelines, incremental data loads, and metadata-driven ingestion frameworks Implement and optimize Databricks components: Delta Lake, Delta Live Tables, Autoloader, Structured Streaming, and Workflows Design large-scale data warehousing solutions with 3NF and dimensional modeling Establish data governance, security, and data quality frameworks, including Unity Catalog Lead ML lifecycle management using MLflow and drive AI use cases (RAG, AI/BI) Manage cloud-native deployments on Microsoft Azure and integrate with enterprise systems (e.g., ServiceNow) Drive CI/CD, DevOps practices, and performance optimization of Spark workloads Provide technical leadership, mentor teams, and ensure successful delivery Collaborate with stakeholders to translate business requirements into scalable solutions Ideal Candidate Solid Databricks Architect Profile with end-to-end Lakehouse ownership Mandatory (Experience 1) Must have 10+ years of software engineering experience with atleast 5+ years in Data Engineering with hands on exposure to Databricks and strong ownership of end-to-end data pipeline development. Mandatory (Experience 2) Must have atleast 5+ years of expertise across the Databricks ecosystem Delta Lake, Delta Live Tables, Autoloader, Structured Streaming, Workflows, Unity Catalog Mandatory (Tech skill 1) Must have worked at architecture level, owning end-to-end design through deployment Mandatory (Tech skill 2) Must have strong experience with Python and SQL for data processing and Apache Spark for performance tuning & scalability Mandatory (Tech skill 3) Must have experience in large-scale data warehousing & advanced data modeling (3NF and dimensional) across batch and real-time systems Mandatory (AI Exposure) Must have at Role & Responsibilities Key Responsibilities Architect and implement enterprise-grade Lakehouse solutions using Databricks Design and deliver scalable batch and real-time data pipelines using Apache Spark (PySpark/SQL) Build ETL/ELT pipelines, incremental data loads, and metadata-driven ingestion frameworks Implement and optimize Databricks components: Delta Lake, Delta Live Tables, Autoloader, Structured Streaming, and Workflows Design large-scale data warehousing solutions with 3NF and dimensional modeling Establish data governance, security, and data quality frameworks, including Unity Catalog Lead ML lifecycle management using MLflow and drive AI use cases (RAG, AI/BI) Manage cloud-native deployments on Microsoft Azure and integrate with enterprise systems (e.g., ServiceNow) Drive CI/CD, DevOps practices, and performance optimization of Spark workloads Provide technical leadership, mentor teams, and ensure successful delivery Collaborate with stakeholders to translate business requirements into scalable solutions Ideal Candidate Solid Databricks Architect Profile with end-to-end Lakehouse ownership Mandatory (Experience 1) Must have 10+ years of software engineering experience with atleast 5+ years in Data Engineering with hands on exposure to Databricks and strong ownership of end-to-end data pipeline development. Mandatory (Experience 2) Must have atleast 5+ years of expertise across the Databricks ecosystem Delta Lake, Delta Live Tables, Autoloader, Structured Streaming, Workflows, Unity Catalog Mandatory (Tech skill 1) Must have worked at architecture level, owning end-to-end design through deployment Mandatory (Tech skill 2) Must have strong experience with Python and SQL for data processing and Apache Spark for performance tuning & scalability Mandatory (Tech skill 3) Must have experience in large-scale data warehousing & advanced data modeling (3NF and dimensional) across batch and real-time systems Mandatory (AI Exposure) Must have at

One address, no account. We’ll tell you when matching roles go live.