Padmi

Principal Data Infrastructure Engineer

IndiaPosted 3 months ago
Software engineeringStaff+Full Time; Regular
Apply at Publicis Production

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Data Engineer at our company, you will be responsible for designing and building scalable cloud-based data solutions across multiple platforms, with a primary focus on GCP while also utilizing Snowflake and Databricks. You will play a crucial role in creating scalable data pipelines, optimizing data workflows, and ensuring data quality and availability for production technology. Your responsibilities will include: - Architecting and maintaining robust data pipelines that integrate internal and external data sources through batch and streaming processes (such as APIs, structured streaming, and message queues). - Collaborating with data analysts, scientists, and software engineers to understand data requirements and develop appropriate solutions. - Implementing data quality checks, data governance practices, and monitoring systems to ensure the reliability and trustworthiness of data. - Optimizing the performance of ETL/ELT workflows and enhancing infrastructure scalability. To excel in this role, you should have: - 7+ years of experience in data engineering and solution delivery, demonstrating technical leadership. - Proficiency in Python (including PySpark), SQL, and cloud-based data engineering tools. - Expertise in AWS, GCP, and/or Databricks, along with managing cloud-based data infrastructure. - Strong background in database technologies like SQL Server, Redshift, PostgreSQL, and Oracle. - Familiarity with machine learning pipelines, MLOps practices, and tools like Git, CI/CD pipelines, and DevOps tools. - Hands-on experience with web scraping, REST API integrations, and streaming data pipelines. This full-time role may require limited travel based on team needs, and you should be prepared to operate in a global organization. As a Senior Data Engineer at our company, you will be responsible for designing and building scalable cloud-based data solutions across multiple platforms, with a primary focus on GCP while also utilizing Snowflake and Databricks. You will play a crucial role in creating scalable data pipelines, optimizing data workflows, and ensuring data quality and availability for production technology. Your responsibilities will include: - Architecting and maintaining robust data pipelines that integrate internal and external data sources through batch and streaming processes (such as APIs, structured streaming, and message queues). - Collaborating with data analysts, scientists, and software engineers to understand data requirements and develop appropriate solutions. - Implementing data quality checks, data governance practices, and monitoring systems to ensure the reliability and trustworthiness of data. - Optimizing the performance of ETL/ELT workflows and enhancing infrastructure scalability. To excel in this role, you should have: - 7+ years of experience in data engineering and solution delivery, demonstrating technical leadership. - Proficiency in Python (including PySpark), SQL, and cloud-based data engineering tools. - Expertise in AWS, GCP, and/or Databricks, along with managing cloud-based data infrastructure. - Strong background in database technologies like SQL Server, Redshift, PostgreSQL, and Oracle. - Familiarity with machine learning pipelines, MLOps practices, and tools like Git, CI/CD pipelines, and DevOps tools. - Hands-on experience with web scraping, REST API integrations, and streaming data pipelines. This full-time role may require limited travel based on team needs, and you should be prepared to operate in a global organization.

One address, no account. We’ll tell you when matching roles go live.

More at Publicis Production

Related open roles

View all roles