Source description
About the role
Balance feature development with system reliability by defining and maintaining Service Level Objectives (SLOs). Monitor production environments continuously to ensure system availability, performance, and overall health. Lead incident management processes and support a blameless post-mortem culture for continuous improvement. Collaborate with development teams to improve service quality through testing and controlled release processes. Contribute to system architecture, platform management, and capacity planning initiatives. Build scalable, sustainable, and reliable systems using automation and continuous improvement practices. Promote reliability engineering and resilience best practices across teams and systems. Apply programming or scripting skills in languages such as Go, Python, Java, C/C++, Perl, Ruby, or Shell. Utilize strong understanding of software engineering principles, algorithms, data structures, and product engineering practices. Demonstrate expertise in UNIX internals, networking fundamentals, and systems engineering concepts. Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
More at Knowledgesprint Technologies