Source description
About the role
Strong understanding of critical concepts in site reliability Engineering, DevSecOps, and Agile principles, with a focus on enabling developer productivity and improving operational efficiency. Some tools in production Kubernetes, LGTM, Istio, *Every AWS Service, Advanced Networking, CI/CD, Karpenter. Cloud computing Expertise Advanced working knowledge of cloud platforms such as Amazon Web Services (AWS), Google Cloud Platform (GCP), including proficiency with cloud-native services like serverless computing, EKS, AWS Datastores, Distributed Systems. Continuous Integration and Continuous Delivery (CI/CD) Expertise in designing and implementing CI/CD pipelines, integrating security, testing, and automation into all stages of the software development lifecycle. Tools could include GitHub Actions, , Jenkins, ArgoCD, Flux or equivalent platforms. Infrastructure-as-Code (IaC) Proficiency in using IaC tools such as Terraform, Ansible, CloudFormation, while adhering to everything-as-code principles for infrastructure automation and standardization. Strategic Leadership Ability to shape and drive the strategic direction of your assigned area of expertise via Roadmaps, Presentations, and Documentation. Participate in SCRUM practices, stand-up, sprint planning and close collaboration with Application Dev team sprint and SCRUM ceremonies. Advanced knowledge of containerization technologies such as Docker and Kubernetes. Capable of managing and orchestrating containerized applications at scale. Observability and Monitoring Practical experience with observability tools and practices (e.g., Grafana, Prometheus, ELK Stack, OpenTelemetry) to ensure system health and optimize performance via logging, metrics, and tracing. DevSecOps and Security Strong expertise in DevSecOps practices, including automated security testing, policy enforcement, and implementing secure coding practices. Understanding of cloud-native security tools. Site Reliability Engineering (SRE) Familiarity with SRE principles like Service Level Objectives (SLOs), error budgets, and incident management to maintain system reliability and minimize downtime.
More at Virtusa