Padmi

Data Architect - Lead Systems Engineer

Delhi NCRPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at AutoZone

Opens the source posting on shine.com

Source description

About the role

View original

You will be responsible for building and optimizing a robust data platform in the automotive industry as the Senior Data Architect at AutoZone. Your key responsibilities will include: - Leading, mentoring, and managing a team of data engineers to ensure high performance. - Owning the GCP Data Lake architecture and implementation for secure and scalable data processing. - Designing and overseeing the Lakehouse architecture leveraging Delta Lake and Apache Spark. - Implementing and managing GCP Data Lake Unity Catalog for unified data governance. - Collaborating with analytics teams to develop and optimize GCP Data Lake SQL queries and dashboards. - Leading performance tuning initiatives and implementing best practices for incremental data processing. - Working closely with domain analysts, data scientists, and product owners to translate requirements into robust data pipelines. - Integrating GCP Data Lake workflows into the CI/CD pipeline using DevOps principles and Git. - Collaborating with security and compliance teams to uphold data governance standards. - Staying updated with the latest GCP Data Lake features and industry best practices. Qualifications required for this role include: - 10+ years of experience in data engineering, data architecture, or related roles. - Significant hands-on experience with GCP Data Lake and the Apache Spark ecosystem. - Proficiency in building data pipelines using PySpark/Scala and managing data in Delta Lake format. - Strong skills in vector databases and embedding models. - Advanced SQL skills and solid understanding of data warehousing concepts. - Experience optimizing ETL jobs for performance and cost efficiency. - Demonstrated experience implementing data security and governance measures. - Experience leading and mentoring engineering teams. - Strong communication skills and experience working in an Agile environment. Tools & Technologies: - GCP Data Lake Lakehouse Platform: GCP Data Lake Workspace, Apache Spark, Delta Lake, GCP Data Lake SQL, MLflow. - Data Governance: GCP Data Lake Unity Catalog. - Programming & Data Processing: PySpark, Python, SQL. - Cloud Services: GCP Cloud Storage, GCP Pub/Sub technologies, Vector Databases. - DevOps & CI/CD: Git, Jenkins, Terraform. - Other Tools: JIRA, Confluence, Looker Studio, PowerBI. Preferred Certificates, Licenses, and Registrations: - GCP Data Lake Certified Data Engineer Professional or GCP Data Lake Certified Data Engineer Associate. - Exposure to related big data and streaming tools such as Apache Kafka, GCP Pub/Sub services, Apache Airflow, BI/analytics tools is advantageous. You will be responsible for building and optimizing a robust data platform in the automotive industry as the Senior Data Architect at AutoZone. Your key responsibilities will include: - Leading, mentoring, and managing a team of data engineers to ensure high performance. - Owning the GCP Data Lake architecture and implementation for secure and scalable data processing. - Designing and overseeing the Lakehouse architecture leveraging Delta Lake and Apache Spark. - Implementing and managing GCP Data Lake Unity Catalog for unified data governance. - Collaborating with analytics teams to develop and optimize GCP Data Lake SQL queries and dashboards. - Leading performance tuning initiatives and implementing best practices for incremental data processing. - Working closely with domain analysts, data scientists, and product owners to translate requirements into robust data pipelines. - Integrating GCP Data Lake workflows into the CI/CD pipeline using DevOps principles and Git. - Collaborating with security and compliance teams to uphold data governance standards. - Staying updated with the latest GCP Data Lake features and industry best practices. Qualifications required for this role include: - 10+ years of experience in data engineering, data architecture, or related roles. - Significant hands-on experience with GCP Data Lake and the Apache Spark ecosystem. - Proficiency in building data pipelines using PySpark/Scala and managing data in Delta Lake format. - Strong skills in vector databases and embedding models. - Advanced SQL skills and solid understanding of data warehousing concepts. - Experience optimizing ETL jobs for performance and cost efficiency. - Demonstrated experience implementing data security and governance measures. - Experience leading and mentoring engineering teams. - Strong communication skills and experience working in an Agile environment. Tools & Technologies: - GCP Data Lake Lakehouse Platform: GCP Data Lake Workspace, Apache Spark, Delta Lake, GCP Data Lake SQL, MLflow. - Data Governance: GCP Data Lake Unity Catalog. - Programming & Data Processing: PySpark, Python, SQL. - Cloud Services: GCP Cloud Storage, GCP Pub/Sub technologies, Vector Databases. - DevOps & CI/CD: Git, Jenkins, Terraform. - Other Tools: JIRA, Confluence, Looker Studio, PowerBI. Preferred Certificates, Licenses, and Registrations: - GCP Data Lak

One address, no account. We’ll tell you when matching roles go live.

More at AutoZone

Related open roles

View all roles