Padmi
Marsh logo
Marsh

risk management consulting · insurance broking

Senior Principal Data Engineer

Delhi NCRPosted 3 months ago
Software engineeringStaff+Full Time; Regular
Apply at Marsh

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Principal Engineer in Data Engineering at Marsh Risk, your role will involve designing and implementing scalable data pipelines and AI-based solutions using Databricks. You will be responsible for end-to-end ETL/ELT processes, managing large datasets, and utilizing tools like Python, PySpark, and AWS S3 to ensure data transformation and optimization for analytical purposes. Your work will focus on cutting-edge cloud and hybrid data projects, turning raw data into meaningful insights and AI analytics. You will collaborate closely with architects and business stakeholders from day one. Key Responsibilities: - Develop and maintain data pipelines using Databricks and the Medallion Architecture (Bronze, Silver, Gold layers). - Design AI-based solutions using Databricks Genie and ensure end-to-end integration. - Utilize cloud-native tools or other applications to expose/consume Databricks features via API. - Write data transformation scripts using Python and PySpark. - Store and manage real-time data in AWS S3 and integrate with other cloud-based services. - Utilize SQL to query, clean, and manipulate large datasets. - Collaborate with cross-functional teams to ensure data accessibility for business intelligence and analytics. - Monitor and troubleshoot data pipelines for performance and reliability. - Document data processes and adhere to best practices for scalability and maintainability. - Ingest and process structured and unstructured data across batch and streaming sources. Qualifications Required: - Experience with Databricks components like Pipeline, scheduled/event-based jobs, Genie, Unity Catalog, and Datawarehouse. - Proficiency in Python, PySpark, and SQL for data processing and transformation using AWS S3 data. - Knowledge of Data Governance, data access security, and configuring Job compute for different Jobs in Databricks. - Familiarity with version control using Git. - Understanding of Databricks API and its integration with different tools and applications. - Understanding of bulk data and real-time data streaming. - Experience with Delta Lake and other Databricks technologies. - Knowledge of additional AWS services (e.g., Athena, Glue, Lambda, S3, DMS). In addition to the technical aspects of the role, Marsh offers professional development opportunities, an inclusive culture, and a collaborative work environment where you can create new solutions and have an impact on colleagues, clients, and communities. With a global presence and a commitment to diversity and flexibility in the workplace, Marsh provides a range of career opportunities, benefits, and rewards to enhance your well-being. Join a team that helps you thrive through the power of perspective. Please note that Marsh encourages a diverse, inclusive, and flexible work environment, promoting diversity across various aspects. All Marsh colleagues are expected to work at least three days per week in their local office or onsite with clients, with teams identifying an "anchor day" for in-person collaboration. As a Senior Principal Engineer in Data Engineering at Marsh Risk, your role will involve designing and implementing scalable data pipelines and AI-based solutions using Databricks. You will be responsible for end-to-end ETL/ELT processes, managing large datasets, and utilizing tools like Python, PySpark, and AWS S3 to ensure data transformation and optimization for analytical purposes. Your work will focus on cutting-edge cloud and hybrid data projects, turning raw data into meaningful insights and AI analytics. You will collaborate closely with architects and business stakeholders from day one. Key Responsibilities: - Develop and maintain data pipelines using Databricks and the Medallion Architecture (Bronze, Silver, Gold layers). - Design AI-based solutions using Databricks Genie and ensure end-to-end integration. - Utilize cloud-native tools or other applications to expose/consume Databricks features via API. - Write data transformation scripts using Python and PySpark. - Store and manage real-time data in AWS S3 and integrate with other cloud-based services. - Utilize SQL to query, clean, and manipulate large datasets. - Collaborate with cross-functional teams to ensure data accessibility for business intelligence and analytics. - Monitor and troubleshoot data pipelines for performance and reliability. - Document data processes and adhere to best practices for scalability and maintainability. - Ingest and process structured and unstructured data across batch and streaming sources. Qualifications Required: - Experience with Databricks components like Pipeline, scheduled/event-based jobs, Genie, Unity Catalog, and Datawarehouse. - Proficiency in Python, PySpark, and SQL for data processing and transformation using AWS S3 data. - Knowledge of Data Governance, data access security, and configuring Job compute for different Jobs in Databricks. - Familiarity with version control using Git. - Understanding of Databricks AP

One address, no account. We’ll tell you when matching roles go live.

More at Marsh

Related open roles

View all roles