Source description
About the role
Role- Site Reliability Engineer Duration- 8 Month C2H (can be converted full-time after 3 month mark) Location- Sunnyvale, CA/ Bentonville, AR/ Dallas, TX (expected to be onsite 1-2 days per week) Description: Site Reliability Engineers are hybrid systems and software engineers who are responsible and take ownership for reliability, scalability, automation, and other issues related to uptime and availability of Sam's e-commerce/Retail and Enterprise platform. Our goal is to build, scale, and guard the systems that delight the customers. To do so, you will need to have strong skills in the following areas: Required: Must have at minimum 6 years of US experience (related to the below responsibilities) Responsibilities: · Design, write, and build tools to improve the reliability, latency, availability, and scalability of Sam's e-commerce/Retail and Enterprise products. · Engender reliability and availability starting with metrics and measurements · Build tools/automate to prevent the re-occurrence of problems to mission-critical products/services. · Augment existing instrumentation to build a cohesive picture of the characteristics of our systems with special attention to points of failure. · Develop a deep understanding of the various services and applications that come together to deliver Sam's e-commerce/Retail and Enterprise products. · Design new tools to monitor and smart alerts that help discover failures/issues in a timely fashion and work with engineers to identify root causes and fix issues · Root-cause analysis of complex problems involving multiple parties, networks, hardware, and software that relate to scaling and performance · Secure the system from issues, be they real, perceived, or notional · High focus on collecting and inferring metric documentation to be used by others to build and maintain systems. · Scripting and Development responsibilities · Build and drive the automation systems that maintain system health Duties and Tech Skills: -Working with backend teams -Analyzing anomaly issues on API’s that connect to multiple clients -Needs to troubleshoot through logs and figuring out what the root cause is and working with teams to get issues resolved. -Find an error, look at multiple sources, find out root cause. -Will work with other leads to put in any process improvement plans in to place. Candidate Profile: -Kubernetes -Azure: OnPrem Azure -Pod failures -Bad partitions with Azure Setup -Building out scripts, monitoring alerts, modifying dashboards -Node, NodeJS, some Java --Splunk, promethus and Azure products -Grafana- building dashboards (Plus) -Experience with Scripting/Code -JAVA, REACT, SQL -Hackerrank- will do a scripting challenge during interview with 3 team members -- ---- Thanks & Regards, Tejash - Technical Recruiter Thought wave Software and Solutions 314 N. Lake St, Suite 6, Aurora IL 60506 Desk: 6304918494 EXTN: 157 Email :tejash@thoughtwavesoft.com Website:www.thoughtwavesoft.com https://www.linkedin.com/in/tejash-chettipalli-342aa2246 A Certified Minority Business Enterprise, Disadvantaged Business Enterprise, SAM.gov, SOC2 & ISO2005, Vendors for: · STATES: IL, PA, TN, AK, OR, CT, GA, VA, ID, IA, UT, FL, MN & CO · NATIONAL LABS: ARGONNE & FERMI · COUNTIES: HENNEPIN, MN, FULTON & GA · PUBLIC SCHOOLS: ATLANTA. SECURITY NOTICE: This communication, including any accompanying document(s), is for the sole use of the intended recipient and may contain confidential information. Unauthorized use, distribution, disclosure or any action taken or omitted to be taken in reliance on this communication is prohibited, and may be unlawful. If you are not the intended recipient, please notify the sender (admin@thoughtwavesoft.com) by return e-mail or telephone (630-448-6681) and permanently delete or destroy all electronic and hard copies of this e-mail.
More at Thoughtwave Software and Solutions