Source description
About the role
Key Responsibilities Ensure high availability and maximum uptime in SaaS environments by proactively monitoring and managing infrastructure, applications, and network systems using tools like Zabbix, Grafana, Prometheus, Sumo Logic, and Amazon CloudWatch. Proactively identify anomalies, threshold breaches, and performance bottlenecks to prevent potential incidents. Respond to alerts and incidents in real time, perform initial triage and troubleshooting across Windows/Linux servers, applications, and network components, validate service health (APIs/endpoints), and escalate complex issues with detailed analysis. Validate application availability through endpoint checks, API monitoring, and service health verification. Execute routine operations including system health checks, patching, maintenance, backups, and disaster recovery using cloud-native tools (e.g., AWS Backup, S3, RDS), while also monitoring resource utilization and supporting cost optimization initiatives Contribute to the creation, review, and continuous improvement of SOPs, runbooks, and knowledge base articles. Participate in Change Management and Change Control processes by implementing approved changes, validating deployments, and minimizing risk to production systems. Maintain accurate records in ticketing systems, ensuring proper documentation aligned with ITIL processes and audit requirements Participate in shift handovers, ensuring clear communication of ongoing incidents, system status, and pending actions Ensure strict adherence to SLAs, operational processes, security guidelines, and compliance standards Required Qualifications Associate s or Bachelor s degree in Information Technology, Networking, or a related field (or equivalent practical experience) 2 5 years of experience in a NOC, Cloud Operations, or Network Support environment Solid understanding of networking fundamentals, including TCP/IP, DNS, DHCP, and VPN concepts Strong knowledge of operating systems, particularly Linux and Windows environments Hands-on experience or familiarity with monitoring and observability tools such as Zabbix, Grafana, Prometheus, Sumo Logic, Datadog, ELK stack etc. Good understanding of ITIL practices (Incident, Change, and Problem Management) in an operations environment Experience troubleshooting SaaS application performance, system reliability, and cloud-based service disruptions. Willingness to work in a 24/7 shift-based support model Effective communication and documentation skills for incident reporting, escalations, and knowledge sharing Working knowledge of AWS Preferred Qualifications: AWS certifications (e.g., AWS Certified Solutions Architect, AWS Certified DevOps Engineer). Experience with hybrid cloud environments and on-premises-to-cloud migrations. Familiarity with other cloud platforms like Azure or GCP. Knowledge of database management (e.g., RDS, DynamoDB) and caching solutions (e.g., Redis, ElastiCache). Basic scripting knowledge (Shell or Python) is an advantage for automation and operational efficiency Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
More at Five9