Padmi

Lead Devops Engineer

IndiaPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Société Générale

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: You will be joining the SG Cloud Platform Engineering organization as a Platform Engineer, where your main responsibility will be to ensure the stability, performance, and automation of the cloud platform's core services. This includes overseeing API automation layers, observability components, CI/CD workflows, IaC toolchains, and QA/Documentation systems. Your role will involve operating production services at scale, reducing operational toil through automation, and continuously enhancing the platform's delivery and observability capabilities. Key Responsibilities: - Operate and maintain key platform services such as the Terraform Registry, Tracing infrastructure, SGCP Quality & Observability resources, and documentation & chat-support systems. - Ensure availability, performance, resilience, and secure lifecycle management for all production components. - Perform patching, upgrades, and vulnerability remediation with minimal human intervention on production systems. - Lead incident response, conduct deep root-cause analysis, and implement long-term corrective actions. - Reduce operational toil through automation, workflow industrialization, and proactive reliability engineering. - Operate and evolve the cloud platform's CI/CD pipelines and reusable workflows used by approximately 300 developers. - Manage the lifecycle of base Docker images, including security hardening, automated build pipelines, versioning, and distribution. - Maintain and extend the platform's IaC toolchain, encompassing Terraform workflows, deployment pipelines, and registry management. - Continuously enhance delivery performance, deployment reliability, and overall developer experience. - Maintain and enhance the cloud platform's observability stack across traces and dashboards. - Ensure full visibility into system behavior, performance drifts, errors, and capacity indicators. - Build automation for alerting, anomaly detection, and platform health insights to improve signal quality and reduce noise. - Participate in system demos, validation sessions, and operational readiness reviews. - Act as a partner for SG Cloud engineering teams in troubleshooting and platform enablement. Qualifications Required: - 7+ years of experience in software development and architecture. - Strong experience in Python, Golang object-oriented programming, and microservices architecture. - Solid understanding of cloud infrastructure, operations, and resource definitions with familiarity in AWS and Azure. - Proficiency in PostgreSQL and writing robust SQL queries for database management. - Experience with monitoring and observability at scale using metrics, logs, and traces. - Hands-on experience with Docker and Kubernetes for containerization and deployment. - Proficient with GIT, Jenkins, Sonar, and automation pipelines for DevOps & CI/CD. - Strong knowledge of Linux OS and bash scripting for system skills. - Experience working in Agile/SAFe environments for Agile Practices. - Excellent communication skills for cross-functional collaboration and transversal responsibilities. Role Overview: You will be joining the SG Cloud Platform Engineering organization as a Platform Engineer, where your main responsibility will be to ensure the stability, performance, and automation of the cloud platform's core services. This includes overseeing API automation layers, observability components, CI/CD workflows, IaC toolchains, and QA/Documentation systems. Your role will involve operating production services at scale, reducing operational toil through automation, and continuously enhancing the platform's delivery and observability capabilities. Key Responsibilities: - Operate and maintain key platform services such as the Terraform Registry, Tracing infrastructure, SGCP Quality & Observability resources, and documentation & chat-support systems. - Ensure availability, performance, resilience, and secure lifecycle management for all production components. - Perform patching, upgrades, and vulnerability remediation with minimal human intervention on production systems. - Lead incident response, conduct deep root-cause analysis, and implement long-term corrective actions. - Reduce operational toil through automation, workflow industrialization, and proactive reliability engineering. - Operate and evolve the cloud platform's CI/CD pipelines and reusable workflows used by approximately 300 developers. - Manage the lifecycle of base Docker images, including security hardening, automated build pipelines, versioning, and distribution. - Maintain and extend the platform's IaC toolchain, encompassing Terraform workflows, deployment pipelines, and registry management. - Continuously enhance delivery performance, deployment reliability, and overall developer experience. - Maintain and enhance the cloud platform's observability stack across traces and dashboards. - Ensure full visibility into system behavior, performance drifts, errors, and capacity indicators. - Buil

One address, no account. We’ll tell you when matching roles go live.

More at Société Générale

Related open roles

View all roles