Source description
About the role
Site reliability engineer Oakland, CA 24 months Required skills: Splunk, SQL, Linux, Shell Scripting, Docker, Kubernetes Basic Qualifications Additional Skills Job Description Site Reliability Engineer Responsibilities Working closely with counterparts in the Infrastructure and Application teams, help build a more sustainable platform through the development of systems that analyze our environments, predict future problems, and actively support the production environment Experience working with automating system administration tasks using scripting tools such as Python or shell (preferred). Experience with monitoring and automation tools further help analyze real time issues. Monitoring and Metrics in Prometheus, Grafana and integrations with ServiceNow Work with other teams to make sure that the infrastructure and applications that depend on it work together seamlessly Support other team’s infrastructure needs on an as-needed basis Use and develop tools for systems continuous delivery automation System Administration on Linux (CentOS, etc..) and Windows Server Proficient with DevOps tools and environments like TeamCity/Jenkins, Git. Experience with monitoring implementations and administration Being available to discuss and resolve technical issues and escalations with other technical staff as the need arises Work to automate detection and resolution of recurring issues in the production environment Experience 5-7 years production support Working experience in using Tomcat, Git, Splunk, Jenkins & TeamCity. Knowledge of Relational Databases & SQL. Experience in Shell Scripting Experience in maintaining a monitoring and alert systems Experience troubleshooting relational databases and distributed platforms Experience in maintaining Java applications Experience in Docker orchestration and management. Experience with Kubernetes
More at Thoughtwave Software and Solutions