Padmi
Merck Sharp & Dohme logo
Merck Sharp & Dohme

oncology medicines · vaccines

Senior Specialist, GSF DnA Data Engineer

HyderabadPosted 1 month ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Merck Sharp & Dohme

Opens the source posting on shine.com

Source description

About the role

View original

Job Description The Opportunity Based in Hyderabad, join a global healthcare biopharma company and be part of a 130- year legacy of success backed by ethical integrity, forward momentum, and an inspiring mission to achieve new milestones in global healthcare. Be part of an organisation driven by digital technology and data-backed approaches that support a diversified portfolio of prescription medicines, vaccines, and animal health products. Drive innovation and execution excellence. Be a part of a team with passion for using data, analytics, and insights to drive decision-making, and which creates custom software, allowing us to tackle some of the world's greatest health threats. Our Technology Centers focus on creating a space where teams can come together to deliver business solutions that save and improve lives. An integral part of our companys IT operating model, Tech Centers are globally distributed locations where each IT division has employees to enable our digital transformation journey and drive business outcomes. These locations, in addition to the other sites, are essential to supporting our business and strategy. A focused group of leaders in each Tech Center helps to ensure we can manage and improve each location, from investing in growth, success, and well-being of our people, to making sure colleagues from each IT division feel a sense of belonging to managing critical emergencies. And together, we must leverage the strength of our team to collaborate globally to optimize connections and share best practices across the Tech Centers. Role Overview: We are looking for a Senior Data Engineer to design and build trusted, scalable, and costefficient data platforms on AWS and Databricks. You will lead handson development of batch and streaming data pipelines (ETL/ELT), implement robust data quality and observability practices, and deliver dimensional models that power analytics and reporting. In this role you will also act as a technical mentorsetting engineering standards, coaching other data engineers through design reviews and pair programming, and partnering with stakeholders to translate business needs into well-governed datasets. We are forward-looking and continue evolving our ecosystem with modern patterns such as lakehouse, data mesh, and data fabric. What will you do in this role: Design, develop, and operate end-to-end data pipelines (ETL/ELT) to ingest data from diverse sources into an AWS-based data lakehouse and data warehouse (batch and/or streaming).Build curated datasets using strong data modeling practicesincluding dimensional modeling (star/snowflake), SCD patterns, and conformed dimensionsto support BI and self-service analytics.Partner with product managers, analysts, and data scientists to understand requirements, define source-to-target mappings, and deliver datasets that are accurate, discoverable, and reusable.Define and implement data quality controls (validation rules, reconciliations, anomaly checks), data contracts, and SLAs; partner with governance to maintain catalog, lineage, and business/technical metadata.Implement orchestration, logging, monitoring, and alerting to ensure reliable operations (data observability, pipeline health, backfills, and incident triage).Apply engineering best practices: automated unit/integration tests, code reviews, and CI/CD to promote changes safely across environments.Develop transformations on Databricks using Python/PySpark and Spark SQL; optimize jobs for performance and cost (partitioning, file sizing, caching, tuning).Write and optimize complex SQL for analysis, transformations, and warehouse/lakehouse consumption patterns.Build on AWS using services such as S3, IAM, Glue, Lambda, Step Functions, EMR/ECS/Fargate, and CloudWatch; implement secure access patterns and least-privilege principles.Use Infrastructure as Code (Terraform) to provision and manage cloud resources; promote reusable modules and automated deployments.Package and run workloads using Docker where appropriate, ensuring repeatable environments for development and deployment.Use GitHub for version control; follow branching strategies (e.g., trunk-based or GitFlow) and maintain high-quality pull requests.Process large datasets using PySpark and lakehouse formats (e.g., Delta/Parquet), applying best practices for reliability and scalability.Create reproducible prototypes and analyses using notebooks when appropriate, then productionize solutions with proper packaging, testing, and documentation.Work in an Agile environment (Scrum/Kanban), contributing to sprint planning, estimation, demos, and continuous improvement.Create and maintain technical documentation (data flows, runbooks, data dictionaries) and contribute to team standards and playbooks.Coach and mentor other data engineers through onboarding, pairing, code/design reviews, and knowledge-sharing sessions; raise the teams engineering bar.Provide technical leadership by proposing Job

One address, no account. We’ll tell you when matching roles go live.

More at Merck Sharp & Dohme

Related open roles

View all roles