Source description
About the role
8+ years' experience in software development or performance Testing and Engineering with systems analysis, including using agile practices• Engage in and improve the whole lifecycle of services - from inception and design, through deployment, operation and refinement. • Support services before they go live through activities such as system design consulting, • Capacity planning, risk assessments and launch reviews. • Maintain services once they are live by measuring and monitoring availability, latency and overall system health. • Run continuous Improvement activities that improve service availability, recovery, efficiency & performance. • Work with team to implement highly available and scalable architectures for core and third-party components of the Service. • Implement metrics, monitoring, and incident response processes. • Align with/improve, change management and capacity planning processes. • Monitor levels of manual effort and drive improvement through automation and stability • Proactively improve service availability, performance & efficiency via delivery of o Implemented metrics, monitoring, and incident response processes. o Improvements in change management and capacity planning processes o Automated feature/story engineering quality measures • Be aware of production degradation of Service level objectives. • Identify the specific system information needed to facilitate incident response / post-mortems and deliver root cause analysis. • Contribute as part of a Scrum team to maintain a deep understanding of system functionality and architecture, with primary focus to operational aspects of the service (availability, performance, change management, incident response, capacity planning, etc.) • Operational Excellence to facilitate improved NPS with customers (internal & external) and address specific operational needs & concerns.
More at Logic Planet