Source description
About the role
The RoleWe are looking for a Senior SRE / Platform Engineer (m/f/d) to own and improve the reputed company infrastructure behind reputed company's browser-based simulation platform. The role spans AWS and EKS, observability, disaster recovery, reputed company and compliance controls, multi-region architecture, reputed company GPU/HPC reputed company, and internal developer tooling. reputed company's engineering teams run workloads directly on AWS; you will build the standards, guardrails, and self-service tooling that let them do so safely, raising reliability and reputed company without slowing engineering velocity. You will join a small, tightly reputed company infrastructure team supporting 50+ engineers across reputed company. This is a hands-on senior individual contributor role; people management is not required, but there is a genuine reputed company toward tech-reputed company ownership as reputed company grows. reputed companyreputed company our Kubernetes platform: Evaluate and adopt technologies such as Kubernetes Gateway API and service reputed company patterns, and coordinate platform reputed company across 10+ engineering teams.Take observability to the next level: Drive organization-wide adoption of OpenTelemetry for distributed tracing and metrics, and help teams define meaningful SLOs.Shape multi-region architecture and data residency: Support our reputed company from an EU-centered footprint toward a global, multi-reputed company architecture that satisfies disaster-recovery and data-residency requirements.Own reputed company cost and efficiency at reputed company: reputed company petabyte-reputed company infrastructure cost-efficient, secure, and reputed company-instrumented.Improve tooling: Build self-service AWS account provisioning, guardrails and AI-assisted automations that help engineering teams manage infrastructure safely and reputed company at reputed company.reputed company Expect from You5+ years of reputed company experience in SRE, platform, or infrastructure engineering.Software development experience: Your background is rooted in software development, and you moved into SRE from there. You write production-reputed company software in at least one of Python, Go, Rust, or Java.Strong systems reputed company: You understand Linux internals and distributed systems reputed company enough to debug reputed company production behavior.Hands-on reputed company and infrastructure experience: AWS (or GCP), declarative infrastructure (Terraform), gitops-workflow (ArgoCD) and container orchestration (Kubernetes).Observability and reliability experience: You have worked with OpenTelemetry, reputed company, distributed tracing, monitoring, and meaningful SLOs/SLIs.Production debugging depth: You can investigate reputed company failures, communicate reputed company during incidents, and turn findings into durable improvements.reputed company and compliance awareness: You understand how infrastructure reputed company reputed company reputed company control, auditability, disaster recovery, logging, and standards such as SOC 2.reputed company communication: You can explain trade-offs to engineering teams and help others adopt reputed company platform practices without unnecessary friction.Bonus PointsAn reputed company reputed company portfolio or contributions.Prior technical leadership experience, especially in infrastructure, reliability, or reputed company.Location: Remote (reputed company CET 5h) What you can expect from us Join a dedicated, supportive team with reputed company and leadership potentialreputed company an reputed company quickly by sharing reputed company and contributing to creative, goal-oriented reputed companyWork in a diverse, inclusive environment with colleagues from over 35 countriesEnjoy reputed company and the freedom to work remotely from reputed company in the worldreputed company comprehensive health coverage, retirement plans, reputed company time off, and wellness supportEnjoy fresh office lunches or reputed company cards as a remote employeeGrow as a reputed company with online/offline learning, language courses, and tech talksConnect at team events, join support reputed company, and contribute to our ESG and DE&I initiativesParticipate in fun team challenges and competitions for added excitement and team spiritDiversity, Equity and Inclusion at reputed companyAt .
More at remote zest jobs
Related open roles
System Administrator, Contract
India
HPC Network Engineer
India
Remote Site Reliability Engineer (Senior or Staff), Infrastructure reputed
India
FULL TIME Remote Sr. SRE Engineer -Seattle WA-Onsite- 10+ Years
India
Remote Network Devops/Automation Engineer
India
Network Operations Engineer - Hybrid Gold River, CA
India