Padmi
Oracle logo
Oracle

Cloud Infrastructure (OCI) · AI Database

Principal Core Infrastructure Engineer

United KingdomPosted 2 months ago
Infrastructure And DatabasesUnspecified
Apply at Oracle

Opens the source posting on eeho.fa.us2.oraclecloud.com

Source description

About the role

View original

Core Infrastructure Engineering within Oracle Cloud Infrastructure (OCI) is seeking a motivated Principal Site Reliability Engineer (SRE) who thrives in a fast-paced, rapidly evolving technology environment. The ideal candidate should have experience supporting cloud-scale, highly distributed storage or database services on major cloud platforms such as OCI, AWS, GCP, or Azure. You will be focused on improving service reliability, performance and operability of services used by Oracle OCI Tier-0 services and Oracle customers. You will have your hand on the pulse of the services and will play a key role in responding to live service issues. As a hands-on engineer with strong coding skills. You will debug complex production issues, build automation and monitoring tools. You will have the opportunity to create automation and tooling that will allow us to continuously improve our services. You will own the release certification process by determining whether code is ready for production deployment. In this role, you will be responsible for improving the stability, performance, and reliability of database and storage services which are backbone of OCI. You will collaborate with multiple development teams to identify and resolve cross-functional operational risks by combining engineering expertise, troubleshooting skills, and operational best practices. The role requires a high degree of independence, excellent communication and organizational skills, and a strong commitment to improving customer experience by enhancing service reliability, reducing support tickets, and delivering scalable operational solutions. You will also define and deliver mission-critical services with a strong focus on security, resiliency, scalability, capacity planning, performance management, deployment, and release engineering. Qualifications 8+ years’ experience in Site Reliability Engineering and in storage, networking and database troubleshooting for improving application reliability, scalability, availability Excellent troubleshooting skills for resolving critical production issues in cloud services Expertise in developing scripts (linux scripting), utilities and tools to automate routine or manual intensive tasks Solid experience with CI/CD pipelines, Jenkins and Version Control tools (GitHub, Bit Bucket, GIT) Container administration and development experience utilizing Kubernetes, Docker or similar Solid experience with Configuration Management tools Programming languages development experience using Python, Golang, Terraform Experience with monitoring tools such as Grafana Experience in managing 24×7 high-availability production applications Excellent organizational, verbal, and written communication skills Good understanding of Agile software development principles including using common tools such as JIRA

One address, no account. We’ll tell you when matching roles go live.

More at Oracle

Related open roles

View all roles