Source description
About the role
Key Responsibilities: Monitor application and infrastructure health using Datadog and other observability tools. • Identify, triage, troubleshoot, and resolve production incidents. • Perform root cause analysis (RCA) and coordinate escalations as required. • Monitor and troubleshoot Kubernetes workloads, pods, deployments, and containers. • Collaborate with Development, DevOps, and Engineering teams to ensure high system availability. • Develop and maintain automation scripts using Python, Java, C#, PowerShell, or Bash. • Support CI/CD pipelines and deployment reliability initiatives. • Create and maintain incident documentation, runbooks, and post-mortem reports. • Ensure compliance with security and regulatory requirements. Required Skills: 34 years of experience in Site Reliability Engineering, DevOps, Production Support, or related roles. • Hands-on experience with Kubernetes and Docker. • Experience with Datadog or similar monitoring tools (Prometheus, Grafana, Splunk, Dynatrace, New Relic). • Strong troubleshooting and incident management skills. • Knowledge of scripting/programming languages such as Python, Java, C#, Bash, or PowerShell. • Experience working with AWS, Azure, or GCP cloud platforms. • Familiarity with SQL, MySQL, or NoSQL databases. • Excellent communication and stakeholder management skills. Preferred Skills: CI/CD tools such as Jenkins, GitLab CI/CD, GitHub Actions, or Azure DevOps. • Helm Charts and deployment automation. • ITIL processes and Agile methodologies. • Fintech, Banking, or Payments domain experience. • Knowledge of PCI DSS, ISO 27001, or related compliance standards. Additional Information: Must be willing to work in a 24/7 rotational support environment. • Strong analytical, problem-solving, and collaboration skills are essential.
More at 3shool Technology Consultants