Source description
About the role
Role:Cloud Engineer
Duration:12 months contract Location:100% remote
Day to Day: Work closely with Leidos Engineering and Operations staff as well as the customer’s application owners to solve technical problems at the network, system, and application levels. Lead the team in all areas of telemetry and observability. Help design, test and deploy technical solutions that are innovative and that leverage new both technologies and new methods to shape customer operations. You will have the opportunity to grow in your career as you help us to grow in our value and breadth of services to CMS. Responsible and accountable for managing and following up on incidents, changes, and application release problems through the management channels. Participate in on-call rotation and respond to incident alerts. Building software and systems while managing the platform infrastructure and applications. Creating and maintaining various continuous integration/continuous development pipeline (CI/CD). Focus on proactivity and enablement of self-healing systems. Serve as the expert in creation of KPI’s and alerting thresholds for meaningful metrics relative to the health and performance of the applications the team manages. Ensure availability, reliability, and security and performance of all resources across various applications; and reporting them to owners in a timely manner. Must be a team player, but able to work independently on large, complex projects and assignments in fast paced environment. Provide leadership in problem determination/analysis, isolating system problems utilizing diagnostic and system management tools. Always provide professional and courteous service with excellent verbal and written communications skills. Model inclusive leadership to teammates by building diversity into activities and meetings.
Basic Qualifications: BS degree in in computer science or some equivalent, highly technical discipline. Experience may be substituted in lieu of degree. 5+ years in technical engineering relative to the responsibilities of the Site Reliability Engineer position. Thorough understanding of microservice based architecture. Through understanding of coding best practices, including knowing how to code, typically in a variety of languages, both in a structured and OOP way (e.g., Python, Golang, Ruby, C/C++). Proficient in programming languages for automation (e.g., python) and shell scripting (e.g., bash). Deep knowledge of version control (e.g., Git) and ability to create GitOps practices. Extensive experience with configuring and maintaining monitoring and alerting tools such as Nagios, CloudWatch, Grafana, Prometheus, Splunk ITSI. Proficient in incident management tools (e.g., Splunk On-Call, PagerDuty) Experience with variety of relational and non-relational databases/RDS (e.g., DynamoDB, MongoDB, CosmoDB, PostgreSQL). Strong and relevant experience in cloud technologies, cloud services, IaC, cloud storage, cloud networking and cloud security. Strong knowledge and experience with Cloud IaaS, PaaS, and SaaS offerings. Strong experience with automation and CI/CD tools (e.g., Argo, Jenkins, Travis, Ansible). Knowledge of cloud-based security tools, best practices and policies including demonstrated experience protecting all layers of the application stack. Knowledge of the Software Delivery Life Cycle (SDLC). Excellent writing and verbal communication skills. Ability to manage conflict effectively. Ability to adapt and be productive in a fast-paced dynamic environment. Excellent communication and collaboration skills supporting multiple stakeholders and business operations. Self-starter, self-managed, and a team player.
Preferred Qualifications
Cloud certification (e.g., AWS Solutions Architect Associate, Azure Administrator). Monitoring certification (e.g., Splunk, Prometheus, DataDog) Experience with containerization and orchestration tools (e.g., Kubernetes, Docker). Experience with orchestrating ChatOps. Experience with setting up self-healing components within an application’s infrastructure. Agile-based knowledge and skill, including experience with Scrum Ceremonies and work management tools (e.g., (JIRA, Confluence). Security Skills—Knowledge of information assurance compliance and information security basics within CMS.
--
Thanks & Regards, Pallavi Reddy| Technical Recruiter Thoughtwave Software and Solutions Desk: 6302289745 ,EXTN:167 Email: pallavi@thoughtwavesoft.com Linked in: https://www.linkedin.com/in/pallavi-batchu-2756a8296/
More at Thoughtwave Software and Solutions