Padmi
Persistent Systems logo
Persistent Systems

Wave Relay MANET · mobile ad-hoc networking

Databricks Data Architect

MumbaiPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Persistent Systems

Opens the source posting on shine.com

Source description

About the role

View original

As a Databricks Data Architect at our company, you will be responsible for designing and implementing enterprise Databricks Lakehouse architecture using Delta Lake, Unity Catalog, and Medallion patterns. You will define end-to-end data architecture standards for batch and streaming data workloads, build and optimize ETL/ELT pipelines, and develop scalable data transformations using PySpark, Spark SQL, and SQL. Additionally, you will implement and manage Databricks Jobs, Workflows, and Auto Loader for ingestion pipelines, design and support real-time and event-driven streaming solutions, and apply best practices for performance tuning, cluster optimization, and cost management. Your role will also involve defining and implementing data modeling strategies, enforcing data governance, security, and access control, integrating CI/CD pipelines, collaborating with cross-functional teams, and supporting technical discussions and reviews. Key Responsibilities: - Design and implement enterprise Databricks Lakehouse architecture - Define data architecture standards for batch and streaming workloads - Build and optimize ETL/ELT pipelines using Databricks and AWS Glue - Develop scalable data transformations using PySpark and Spark SQL - Implement and manage Databricks Jobs, Workflows, and Auto Loader - Design real-time and event-driven streaming solutions - Apply best practices for performance tuning and cost management - Define and implement data modeling strategies - Enforce data governance and security using Unity Catalog - Integrate CI/CD pipelines using Jenkins and Azure/AWS DevOps tools - Collaborate with platform, DevOps, and analytics teams - Support technical discussions, reviews, and interviews Qualifications Required: - 12+ years of experience in Data Engineering and/or Data Architecture - Strong hands-on expertise with Databricks and Apache Spark - Deep understanding of Databricks Lakehouse concepts and platform features - Proficiency in Python (PySpark), SQL, and Spark SQL - Experience with Delta Lake, Unity Catalog, Jobs, Workflows, and Auto Loader - Knowledge of batch and streaming architectures - Hands-on experience with AWS services like Glue, S3, and compute services - Experience in CI/CD pipelines using Jenkins - Strong knowledge of data performance tuning and scalability - Experience in large enterprise or complex data environments - Problem-solving, analytical, and stakeholder communication skills In addition to the above, at our company, you can expect a competitive salary and benefits package, a culture focused on talent development with growth opportunities, company-sponsored education and certifications, engagement initiatives, annual health check-ups, and insurance coverage. We are committed to fostering diversity and inclusion, supporting employees with disabilities, offering hybrid work options, and providing an inclusive work environment. "Persistent is an Equal Opportunity Employer and prohibits discrimination and harassment of any kind." As a Databricks Data Architect at our company, you will be responsible for designing and implementing enterprise Databricks Lakehouse architecture using Delta Lake, Unity Catalog, and Medallion patterns. You will define end-to-end data architecture standards for batch and streaming data workloads, build and optimize ETL/ELT pipelines, and develop scalable data transformations using PySpark, Spark SQL, and SQL. Additionally, you will implement and manage Databricks Jobs, Workflows, and Auto Loader for ingestion pipelines, design and support real-time and event-driven streaming solutions, and apply best practices for performance tuning, cluster optimization, and cost management. Your role will also involve defining and implementing data modeling strategies, enforcing data governance, security, and access control, integrating CI/CD pipelines, collaborating with cross-functional teams, and supporting technical discussions and reviews. Key Responsibilities: - Design and implement enterprise Databricks Lakehouse architecture - Define data architecture standards for batch and streaming workloads - Build and optimize ETL/ELT pipelines using Databricks and AWS Glue - Develop scalable data transformations using PySpark and Spark SQL - Implement and manage Databricks Jobs, Workflows, and Auto Loader - Design real-time and event-driven streaming solutions - Apply best practices for performance tuning and cost management - Define and implement data modeling strategies - Enforce data governance and security using Unity Catalog - Integrate CI/CD pipelines using Jenkins and Azure/AWS DevOps tools - Collaborate with platform, DevOps, and analytics teams - Support technical discussions, reviews, and interviews Qualifications Required: - 12+ years of experience in Data Engineering and/or Data Architecture - Strong hands-on expertise with Databricks and Apache Spark - Deep understanding of Databricks Lakehouse concepts and platform features - Proficiency in Python (PyS

One address, no account. We’ll tell you when matching roles go live.

More at Persistent Systems

Related open roles

View all roles