Source description
About the role
As a high-calibre DevOps / SRE / Platform Engineer, you will play a crucial role in designing, building, and operating platforms and infrastructure to facilitate quick and safe product shipping at scale. Your responsibilities will include: - Platform & Cloud Engineering: - Designing and operating cloud-native platforms to enhance developer productivity. - Building reusable infrastructure using Infrastructure-as-Code. - Managing Kubernetes clusters, networking, service orchestration, and workload reliability. - Reliability Engineering: - Defining and driving SLIs, SLOs, and error budgets for critical services. - Leading incident response, participating in on-call rotations, and conducting blameless RCAs. - Automating operational tasks to reduce toil. - CI/CD & DevOps: - Developing secure, automated CI/CD pipelines following GitOps principles. - Ensuring safe production deployments with strong rollback capabilities. - Collaborating with application teams to embed reliability and operational excellence in the lifecycle. - Observability & Operations: - Implementing top-notch logging, metrics, tracing, and alerting systems. - Building actionable alerts and enhancing service health visibility. - Creating dashboards, runbooks, and self-healing mechanisms for faster incident resolution. - Architecture & Collaboration: - Working closely with software engineers, architects, and security teams to influence system design. - Reviewing infrastructure and architecture for scalability, resilience, and cost efficiency. - Advocating for DevOps, SRE, and cloud-native best practices within the organization. In addition to the above responsibilities, you are expected to have: - Strong foundations in Linux, networking, and distributed systems. - Proficiency in Python or Go for automation. - Experience with AWS, Kubernetes, Helm, Terraform, and CI/CD tools. - Knowledge of observability tools like Prometheus, Grafana, and Datadog. - Familiarity with datastores, messaging systems, and cloud security practices. The company values your experience in building internal developer platforms, cloud security, operating high-traffic systems, and strong written communication skills. As a high-calibre DevOps / SRE / Platform Engineer, you will play a crucial role in designing, building, and operating platforms and infrastructure to facilitate quick and safe product shipping at scale. Your responsibilities will include: - Platform & Cloud Engineering: - Designing and operating cloud-native platforms to enhance developer productivity. - Building reusable infrastructure using Infrastructure-as-Code. - Managing Kubernetes clusters, networking, service orchestration, and workload reliability. - Reliability Engineering: - Defining and driving SLIs, SLOs, and error budgets for critical services. - Leading incident response, participating in on-call rotations, and conducting blameless RCAs. - Automating operational tasks to reduce toil. - CI/CD & DevOps: - Developing secure, automated CI/CD pipelines following GitOps principles. - Ensuring safe production deployments with strong rollback capabilities. - Collaborating with application teams to embed reliability and operational excellence in the lifecycle. - Observability & Operations: - Implementing top-notch logging, metrics, tracing, and alerting systems. - Building actionable alerts and enhancing service health visibility. - Creating dashboards, runbooks, and self-healing mechanisms for faster incident resolution. - Architecture & Collaboration: - Working closely with software engineers, architects, and security teams to influence system design. - Reviewing infrastructure and architecture for scalability, resilience, and cost efficiency. - Advocating for DevOps, SRE, and cloud-native best practices within the organization. In addition to the above responsibilities, you are expected to have: - Strong foundations in Linux, networking, and distributed systems. - Proficiency in Python or Go for automation. - Experience with AWS, Kubernetes, Helm, Terraform, and CI/CD tools. - Knowledge of observability tools like Prometheus, Grafana, and Datadog. - Familiarity with datastores, messaging systems, and cloud security practices. The company values your experience in building internal developer platforms, cloud security, operating high-traffic systems, and strong written communication skills.
More at Lenskart.com