Source description
About the role
You will be working with Adobe Pass, a leading authentication and authorization platform that enables seamless access to premium TV and video content across devices. Your role will involve defining and driving the long-term reliability and scalability strategy for the platform, aligning with product and business goals. You will be responsible for architecting large-scale, distributed, and multi-region systems designed for resiliency, observability, and self-healing. Anticipating systemic risks and designing proactive mitigation strategies will be crucial to ensure zero single points of failure across critical services. Collaboration with software architecture and infrastructure teams will be necessary to evolve the platform towards greater reliability, efficiency, and cost optimization. Key Responsibilities: - Define and drive the long-term reliability and scalability strategy for the Adobe Pass platform. - Architect large-scale, distributed, and multi-region systems designed for resiliency and observability. - Anticipate systemic risks and design proactive mitigation strategies. - Partner with software architecture and infrastructure teams to evolve the platform towards greater reliability, efficiency, and cost optimization. Qualifications: - Bachelors or Masters degree in Computer Science, Engineering, or a related field. - 12+ years of experience in site reliability, production engineering, or large-scale distributed system operations. - Proficiency in programming/scripting languages such as Python, Go, Java, or Bash for automation and tooling. - Deep understanding of Kubernetes, microservices, and service mesh architectures. - Experience with Infrastructure as Code (Terraform, CloudFormation) and CI/CD automation frameworks. - Mastery in observability and monitoring stacks like Prometheus, Grafana, Datadog, and OpenTelemetry. - Strong expertise in networking, storage, and distributed databases. - Exceptional communication, leadership, and stakeholder management skills. You will have the opportunity to work with a company like Adobe, which empowers everyone to create through creative platforms and tools that unleash creativity, productivity, and personalized customer experiences. Adobe's industry-leading offerings enable people and businesses to turn ideas into impact, powered by AI and driven by human ingenuity. Adobe is proud to be an Equal Employment Opportunity employer, not discriminating based on various factors. If you are eager to innovate with AI, Adobe is looking for candidates like you who are ready to do the same. You will be working with Adobe Pass, a leading authentication and authorization platform that enables seamless access to premium TV and video content across devices. Your role will involve defining and driving the long-term reliability and scalability strategy for the platform, aligning with product and business goals. You will be responsible for architecting large-scale, distributed, and multi-region systems designed for resiliency, observability, and self-healing. Anticipating systemic risks and designing proactive mitigation strategies will be crucial to ensure zero single points of failure across critical services. Collaboration with software architecture and infrastructure teams will be necessary to evolve the platform towards greater reliability, efficiency, and cost optimization. Key Responsibilities: - Define and drive the long-term reliability and scalability strategy for the Adobe Pass platform. - Architect large-scale, distributed, and multi-region systems designed for resiliency and observability. - Anticipate systemic risks and design proactive mitigation strategies. - Partner with software architecture and infrastructure teams to evolve the platform towards greater reliability, efficiency, and cost optimization. Qualifications: - Bachelors or Masters degree in Computer Science, Engineering, or a related field. - 12+ years of experience in site reliability, production engineering, or large-scale distributed system operations. - Proficiency in programming/scripting languages such as Python, Go, Java, or Bash for automation and tooling. - Deep understanding of Kubernetes, microservices, and service mesh architectures. - Experience with Infrastructure as Code (Terraform, CloudFormation) and CI/CD automation frameworks. - Mastery in observability and monitoring stacks like Prometheus, Grafana, Datadog, and OpenTelemetry. - Strong expertise in networking, storage, and distributed databases. - Exceptional communication, leadership, and stakeholder management skills. You will have the opportunity to work with a company like Adobe, which empowers everyone to create through creative platforms and tools that unleash creativity, productivity, and personalized customer experiences. Adobe's industry-leading offerings enable people and businesses to turn ideas into impact, powered by AI and driven by human ingenuity. Adobe is proud to be an Equal Employment Opportunity empl
More at Adobe