Source description
About the role
Role Overview: As a Senior reliability engineer at Mizuho, you will be responsible for Infrastructure Operations, leveraging observability data, automation, and modern SRE practices. Your role will involve working with platforms such as OpsRamp, InfoSight, VMware Aria Operations, Workspace ONE, NetApp, and ServiceNow to interpret signals, performance, and resiliency data. A strong understanding of AI with automation and toil reduction mindset, along with MELT knowledge, will be critical for this role. Key Responsibilities: - Apply SRE principles (SLIs/SLOs, alert quality, incident reduction) across Infrastructure Operations. - Diagnose cross-platform incidents and interpret health and performance signals from VMware Aria Operations for VDI and VSI environments. - Translate infrastructure signals into actionable metrics, alerts, and observability pipelines for DataMart. - Develop AI-based resolutions and automated approaches for incident resolution. - Collaborate with SRE/Operations teams to eliminate false positives, improve dashboard accuracy, and drive proactive detection. - Support OpenShift/Kubernetes platforms for Day-2 operations and develop automation solutions to improve reliability and response times. Qualification Required: - Bachelor of Engineering or a Degree in Computer Science is required for this role. Organization Overview: Mizuho Global Services (MGS), Pune, is an integral part of Mizuho Financial Group, supporting international businesses with high-quality services. MGS Pune plays a critical role in driving operational excellence, standardization, and innovation for Mizuho Americas, offering competitive compensation and benefits aligned with industry standards. MGS Pune is committed to fostering an inclusive and diverse workplace, subject to background verification checks as per Indian laws and company policies. Role Overview: As a Senior reliability engineer at Mizuho, you will be responsible for Infrastructure Operations, leveraging observability data, automation, and modern SRE practices. Your role will involve working with platforms such as OpsRamp, InfoSight, VMware Aria Operations, Workspace ONE, NetApp, and ServiceNow to interpret signals, performance, and resiliency data. A strong understanding of AI with automation and toil reduction mindset, along with MELT knowledge, will be critical for this role. Key Responsibilities: - Apply SRE principles (SLIs/SLOs, alert quality, incident reduction) across Infrastructure Operations. - Diagnose cross-platform incidents and interpret health and performance signals from VMware Aria Operations for VDI and VSI environments. - Translate infrastructure signals into actionable metrics, alerts, and observability pipelines for DataMart. - Develop AI-based resolutions and automated approaches for incident resolution. - Collaborate with SRE/Operations teams to eliminate false positives, improve dashboard accuracy, and drive proactive detection. - Support OpenShift/Kubernetes platforms for Day-2 operations and develop automation solutions to improve reliability and response times. Qualification Required: - Bachelor of Engineering or a Degree in Computer Science is required for this role. Organization Overview: Mizuho Global Services (MGS), Pune, is an integral part of Mizuho Financial Group, supporting international businesses with high-quality services. MGS Pune plays a critical role in driving operational excellence, standardization, and innovation for Mizuho Americas, offering competitive compensation and benefits aligned with industry standards. MGS Pune is committed to fostering an inclusive and diverse workplace, subject to background verification checks as per Indian laws and company policies.
More at Mizuho