Padmi
Goldman Sachs logo
Goldman Sachs

investment banking · asset management

AI Data Platform Engineer Vice President

IndiaPosted 3 months ago
Software engineeringStaff+Full Time; Regular
Apply at Goldman Sachs

Opens the source posting on shine.com

Source description

About the role

View original

As a Data Engineer in the Lakehouse and AI Data Platform team, your role involves designing, building, testing, and supporting data pipelines and curated datasets on the firms modern data platform. You will work on ingestion, transformation, modelling, optimisation, and data quality to deliver reliable, scalable, and fit-for-purpose data products. Additionally, you may contribute to shared tooling or framework components to enhance the platform's functionality and operations. Key Responsibilities: - Pipeline Engineering - Build, enhance, and support batch and streaming data pipelines on the Lakehouse and AI data platform. - Refactor or modernize existing data flows to improve reliability, performance, and maintainability. - Develop reusable tooling for better delivery, consistency, and operational support. - Ensure data pipelines are production-ready, well tested, and operationally supportable. - Data Modelling and Curation - Develop raw, refined, and curated datasets supporting analytics, reporting, and AI use cases. - Apply sound data modeling principles to accurately represent business entities, relationships, and historical changes. - Collaborate with consumers to shape usable, well-documented data products aligned with business needs. - Data Quality and Reconciliation - Implement controls to validate data completeness, accuracy, and consistency across pipelines and datasets. - Use reconciliation approaches to ensure confidence in production outputs and investigate breaks. - Contribute to testing, monitoring, and issue resolution standards. - Improve testing, monitoring, or reconciliation tooling for platform reliability and day-to-day delivery. - Delivery and Partnership - Work closely with engineers, platform teams, and data consumers to deliver outcomes within set timelines and quality expectations. - Communicate progress, risks, dependencies, and design choices clearly. - For senior candidates, lead technical tasks, breakdowns, and support junior engineers. Required Skills and Experience: - 7-12+ years of experience - Bachelors or masters degree in a relevant discipline or equivalent practical experience - Strong hands-on programming in Python or Java - Good SQL knowledge for troubleshooting, optimization, and data analysis - Ability to quickly learn new tools, platforms, and workflows - Familiarity with software engineering fundamentals, version control, testing, and CI/CD practices - Understanding of temporal data modeling, schema design, and data compatibility considerations - Experience with Apache Spark, distributed data processing, and common data formats - Strong ownership of technical design, code quality, and engineering practices - Ability to lead delivery, manage dependencies, and support less experienced engineers The role involves working with a modern data stack and technologies such as ANSI SQL, Apache Spark, Kafka, JSON, Avro, Parquet, Snowflake, and more. You are not expected to have expertise in every tool immediately but should have the ability to work across similar technologies. In summary, we are looking for engineers who can deliver reliable solutions, take ownership of their work's quality, demonstrate sound judgment in technical decisions, and have a structured approach to problem-solving. Your ability to work in a fast-paced environment and collaborate with stakeholders will be crucial for success in this role. As a Data Engineer in the Lakehouse and AI Data Platform team, your role involves designing, building, testing, and supporting data pipelines and curated datasets on the firms modern data platform. You will work on ingestion, transformation, modelling, optimisation, and data quality to deliver reliable, scalable, and fit-for-purpose data products. Additionally, you may contribute to shared tooling or framework components to enhance the platform's functionality and operations. Key Responsibilities: - Pipeline Engineering - Build, enhance, and support batch and streaming data pipelines on the Lakehouse and AI data platform. - Refactor or modernize existing data flows to improve reliability, performance, and maintainability. - Develop reusable tooling for better delivery, consistency, and operational support. - Ensure data pipelines are production-ready, well tested, and operationally supportable. - Data Modelling and Curation - Develop raw, refined, and curated datasets supporting analytics, reporting, and AI use cases. - Apply sound data modeling principles to accurately represent business entities, relationships, and historical changes. - Collaborate with consumers to shape usable, well-documented data products aligned with business needs. - Data Quality and Reconciliation - Implement controls to validate data completeness, accuracy, and consistency across pipelines and datasets. - Use reconciliation approaches to ensure confidence in production outputs and investigate

One address, no account. We’ll tell you when matching roles go live.

More at Goldman Sachs

Related open roles

View all roles