Source description
About the role
Location: Irving, TX
Duration: 6+ Month Contract to hire
Interview: Onsite. (mandatory)
Term: Hybrid (2 days in office)
Responsibilities
-
Lead architecture and development teams to ensure applications are highly available, reliable, and performant at a global scale.
-
Partner with the architecture team to ensure operability, measurability, and manageability are integrated into business features and enablers.
-
Collaborate with product owners and managers to establish service level objectives (SLOs) for applications and define consequences if objectives are not met.
-
Work with development team members to identify monitoring gaps, improve application performance, and assist with troubleshooting issues.
-
Drive Root Cause Analysis (RCA) of production issues and other failures within the product software, pipeline, or other DevOps support processes or technology.
-
Design, build, and advocate for automated solutions to optimize application/service/platform uptime with minimal human intervention.
-
Participate in an on-call rotation to support troubleshooting and communication efforts outside of normal business hours.
-
Create and implement standards and best practices, driving adoption across development teams and external vendors as applicable.
-
Ensure compliance with all company policies and procedures.
Qualifications
-
Bachelor of Computer Science or related Engineering field required.
-
Master's Degree preferred.
-
Required Skills
-
5-7 years of hands-on SRE experience.
-
1-2 years of leading and mentoring others.
-
Hands-on experience supporting Linux production environments, hands-on administration on Spark, and hands-on experience with MS Azure Cloud technologies.
-
3-5 years hands-on experience with scripting with bash, perl, ruby, or python required.
-
3-5 years experience with Docker Datacenter required.
-
2-4 years of hands-on administration experience on Machine learning platforms required.
-
Minimum of 1 year of experience in Mesos, Kubernetes, OpenShift and/or Deis or other such container/platform-as-a-service orchestrator required.
-
Minimum of 1 year of hands-on experience on CICD tools & Technologies required.
-
Minimum of 1 year of lead experience of site reliability engineering team required.
-
Proven leadership skills and the ability to guide and mentor a team.
-
Strong collaboration and communication skills.
-
A proactive approach to problem-solving and continuous improvement.
-
Passion for automation and operational excellence.
-
Deep expertise in cloud technologies and software development, with a strong technical background.
-
Experience with Java
-
Proficiency in SQL and Powershell.
-
Expertise in defining, implementing, and evaluating Service Level Objectives (SLOs) and Service Level Indicators (SLIs), and associated consequences.
-
Strong skills in performing Root Cause Analysis (RCA) and Problem Management.
-
Extensive experience in cloud native applications Azure/AWS (monitoring, networking, containerization, infrastructure).
-
Proficiency in containerization technologies such as Azure Kubernetes Service, Kubernetes (open source), and Docker.
-
Knowledge of metrics and monitoring tools like Azure Application Insights and Azure Monitor.
-
Familiarity with networking technologies relevant to Azure and AWS, including Azure DNS, Virtual Networks, Azure API Manager, Azure Application Gateway, Akamai WAF/CDN, AWS Route 53, AWS VPC, AWS API Gateway, and AWS CloudFront.
-
Strong experience with Terraform for infrastructure as code.
-
Ability to establish and maintain a culture of learning through the development and sharing of skills, knowledge, processes, and tools; combat traditional silos that create "us and them" environments.
More at 3B Staffing