Padmi

Site Reliability Engineer

San Francisco Bay AreaPosted 1 month ago
Software engineeringUnspecified
Apply at Thoughtwave Software and Solutions

Opens the source posting on atsapp.swarmhr.com

Source description

About the role

View original

Job Title: Site Reliability Engineer Job Location: Will be remote till covid clears (Sunnyvale, CA) Job Duration: 6 Months Contract (Possible CTH) Rate: $70/hr (H1B/USC/GC/H4/GC-EAD) Description: • As a member of the SRE team, you will work with other DevOps practitioners to produce mission-critical infrastructure, tools, and processes that will ensure highest levels of availability and reliability of all our websites, systems, and services. As a senior member of the team, you will be expected to work with management, peers, and customers to define and implement the technical vision of the team. • You are right for the job if you are comfortable with deep technical Linux, networking topics, and distributed architectures. You will work cross-functionally amongst a variety of teams and be a core contributor in every significant engineering service or solution that we deliver to our stakeholders. You will excel if you have enthusiasm for digging deep, and a flare for sharp technical communication, prioritization, and organization. You will work directly with our Software Engineering teams to build our next generation always up and highly available cloud-based e-commerce/Retail and Enterprise platform. • Site Reliability Engineers are hybrid systems and software engineers who are responsible and take ownership for reliability, scalability, automation, and other issues related to uptime and availability of client’s e-commerce/Retail and Enterprise platform. Our goal is to build, scale and guard the systems that delights the customers. Required skills / Experience: • Design, write and build tools to improve the reliability, latency, availability, and scalability of Walmart e-commerce/Retail and Enterprise products. • Engender reliability and availability starting with metrics and measurements. • Enable scaling by providing tools, developing training and/or augmenting processes. • Build tools/automate to prevent re-occurrence of problem to mission critical products/services. • Augment existing instrumentation to build a cohesive picture of the characteristics of our systems with special attention to points of failure. • Participate in capacity planning, demand forecasting, software performance analysis and system tuning. • Design new tools to monitor and smart alerts that help discover failures/issues in a timely fashion and work with engineers to identify root cause and fix issues. • Influence, design and create new architectures, standards, and methods for large-scale enterprise systems. • Root-cause analysis complex problems involving multiple parties, networks, hardware, and software that relate to scaling and performance. • Participate in on-call rotation. • Secure the system from issues, be they real, perceived, or notional. • High focus on collecting and inferring metric documentation to be used by others to build and maintain systems. • Scripting and Development responsibilities • Experience with configuration management tools such as Ansible, Saltstack, Chef and Puppet • Build and drive the automation systems that maintain system health • Eliminate Single Point of failure and test disaster recovery and HA regularly. Note: •Technical screening and live coding is required for the qualification of the candidates - --- Thanks & Regards, Mohan Sai|Technical Recruiter Thoughtwave Software and Solutions 314 N. Lake St, Suite 6, Aurora IL 60506 Mobile: 8023474210 Desk: 6304804004 EXT:143 Email:Mohan@thoughtwavesoft.com Website:www.thoughtwavesoft.com linkedin: https://www.linkedin.com/in/mohan-badaru-112b581b9/

One address, no account. We’ll tell you when matching roles go live.

More at Thoughtwave Software and Solutions

Related open roles

View all roles