Source description
About the role
Role: Site Reliability Engineer (SRE) for Message Queue Location: Temporarily Remote; Preferred San Francisco /LA / Seattle, WA others outside the area must be willing to relocate after 4 months – MUST Duration: 6 + Months Interview: 2 to 3 interviews (Video) Note: 1. This is a support role and a rotational shift (24/7) 2. He / She has to pick the laptop in SFO /LA / Seattle/ NYC once he gets confirmed, Client will take care flight expenses 3. He / She has to relocate to SFO /LA / Seattle after 3 to 5 months. HGS Digital has an opportunity for a Site Reliability Engineer (SRE) for Infrastructure with a major social media platform. Qualified Candidates must be self-motivated and must have learning attitude. Recent and extensive knowledge of Linux basic file systems, memory management, process management, and basic networking along with Linux troubleshooting experience. Must be proactive problem solvers interested in root cause analysis and troubleshooting system issues like performance, debugging process and log analysis. Candidates must have Python programming and shell scripting. Key Responsibilities: • Cluster operations and maintenance, driving on call issues to resolution, troubleshooting, escalation, and documentation • Ensure the reliability, availability, and performance of services through stability and automation product development, disaster recovery plan, emergency response and chaos engineering and system resilience improvements • Manage services, responsible for operational support, 24X7 troubleshooting, automation design and development including deployment • Troubleshoot and diagnose issues, propose, and implement solutions to reduce frequency of occurrence • Meet service-level-agreements (SLAs) or service-level-objective (SLOs) by measuring and monitoring service availability, performance, and overall system health. • Perform various SRE operation including scale up/down, build and maintain clusters • Automate various services and workflows • Available for on-call rotation for production impacting incidents or key customer events Core Experience: 5+ years of experience in the following areas: • Linux Systems Knowledge. e.g. file-systems, memory management, process management, basic networking skills. • Linux Troubleshooting. Debug Linux systems. e.g. file-system level, systems performance issues troubleshooting etc. • Experience in Python programming and Shell scripting. Should be able to code simple programs comfortably. • Strong technical operations, devops and infrastructure support with excellent Linux troubleshooting skills to resolve application issue. • Good knowledge of Kafka. How it works and experience supporting the environment/application. Minimum qualifications: • Bachelor's degree or above, majoring in Computer Science or related fields • Must be responsible, interpersonal self-starters, comfortable with ambiguity, excellent communicators, and problem solvers • Must have the ability to work in a fast paced environment without constant supervision • Motivated learner without requiring constant supervision. • Must have good troubleshooting skills Thanks & Regards, Pavan Kumar| Jr.US IT Recruiter Thought wave Software and Solutions Desk: 6303810025 Email:pavan@thoughtwavesoft.com Website:www.thoughtwavesoft.com https://www.linkedin.com/in/pavan-kumar-madivada-416274170/
More at Thoughtwave Software and Solutions