Padmi

Senior Engineer - Cloud Technologies & Infrastructure

Delhi NCRPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at DUNNHUMBY IT SERVICES INDIA

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Engineer at dunnhumby, you play a crucial role in ensuring the reliability and uptime of our cloud-hosted services. Your responsibilities include: - Maintaining and supporting infrastructure services in development, integration, and production environments. - Designing, implementing, and managing robust, scalable, and high-performance systems and infrastructure. - Ensuring the reliability, availability, and performance of critical services through proactive monitoring, incident response, and root cause analysis. - Driving the adoption of automation, CI/CD practices, and infrastructure as code (IaC) to streamline operations and improve operational efficiency. - Defining and enforcing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) to monitor and improve service health. - Leading incident management, troubleshooting, and postmortems to identify and address operational challenges. - Managing scaling strategies and disaster recovery for cloud-based environments (GCP, Azure). - Driving improvements in operational tooling, monitoring, alerting, and reporting. - Acting as a subject matter expert in Observability, reliability engineering best practices, and promoting these practices across the organization. - Focusing on automation to improve scale and reliability. Qualifications required for this role include: - 8+ years of experience in an engineering role with hands-on experience in the public cloud. - Expertise in cloud technologies (GCP, Azure) and infrastructure automation tools (Terraform, Ansible, etc.). - Proficiency in containerization technologies such as Docker, Kubernetes, and Helm. - Experience with monitoring and observability tools like Prometheus, Grafana, NewRelic, or similar. - Strong knowledge of CI/CD pipelines and related automation tools. - Proficient in scripting languages like Python, Bash, Go. - Strong troubleshooting and problem-solving skills. - Familiarity with incident management processes and tools (e.g., ServiceNow, XMatters). - Ability to learn and adapt in a fast-paced environment while producing quality code. - Ability to work collaboratively on a cross-functional team with a wide range of experience levels. At dunnhumby, you can expect a comprehensive rewards package along with personal flexibility, thoughtful perks like flexible working hours and your birthday off, an investment in cutting-edge technology, and a nimble, small-business feel that gives you the freedom to play, experiment, and learn. The company also emphasizes diversity and inclusion through various networks such as dh Women's Network, dh Proud, dh Parents & Carers, dh One, and dh Thrive. dunnhumby values work/life balance and offers flexible working options to help you achieve a successful career while maintaining your commitments and interests outside of work. If flexibility is important to you, please discuss agile working opportunities with your recruiter during the hiring process. As a Senior Engineer at dunnhumby, you play a crucial role in ensuring the reliability and uptime of our cloud-hosted services. Your responsibilities include: - Maintaining and supporting infrastructure services in development, integration, and production environments. - Designing, implementing, and managing robust, scalable, and high-performance systems and infrastructure. - Ensuring the reliability, availability, and performance of critical services through proactive monitoring, incident response, and root cause analysis. - Driving the adoption of automation, CI/CD practices, and infrastructure as code (IaC) to streamline operations and improve operational efficiency. - Defining and enforcing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) to monitor and improve service health. - Leading incident management, troubleshooting, and postmortems to identify and address operational challenges. - Managing scaling strategies and disaster recovery for cloud-based environments (GCP, Azure). - Driving improvements in operational tooling, monitoring, alerting, and reporting. - Acting as a subject matter expert in Observability, reliability engineering best practices, and promoting these practices across the organization. - Focusing on automation to improve scale and reliability. Qualifications required for this role include: - 8+ years of experience in an engineering role with hands-on experience in the public cloud. - Expertise in cloud technologies (GCP, Azure) and infrastructure automation tools (Terraform, Ansible, etc.). - Proficiency in containerization technologies such as Docker, Kubernetes, and Helm. - Experience with monitoring and observability tools like Prometheus, Grafana, NewRelic, or similar. - Strong knowledge of CI/CD pipelines and related automation tools. - Proficient in scripting languages like Python, Bash, Go. - Strong troubleshooting and problem-solving skills. - Familiarity with incident mana

One address, no account. We’ll tell you when matching roles go live.

More at DUNNHUMBY IT SERVICES INDIA

Related open roles

View all roles