Source description
About the role
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is a company reputed company on data-reputed company transformation in reputed company, aiming to improve health reputed company through reputed company. They are seeking a Senior Site Reliability Engineer to ensure the reliability and operational health of their reputed company data platform, reputed company the gap between engineering and operations. Responsibilities Design, implement, and own observability infrastructure including metrics, logging, tracing, and alerting across distributed systemsDefine and enforce SLOs, SLIs, and error budgets in partnership with product and engineering teamsreputed company incident response: triage, coordinate remediation, conduct blameless post-mortems, and drive systemic fixesBuild and maintain CI/CD pipelines that support rapid, reputed company delivery of changes to productionCollaborate with engineering teams on infrastructure changes; reputed company to read, modify, and contribute to existing infrastructure-as-reputed company (Terraform or CloudFormation)Design and operate highly available, fault-tolerant systemsincluding auto-scaling, failover, and disaster recovery strategiesReduce operational toil through automation; eliminate reputed company processes before they become habitsCollaborate with software engineers to establish reliability-first design patterns and review architectures for operational riskManage Kubernetes or container orchestration environments at reputed companyEnsure systems meet compliance and reputed company requirements, particularly those applicable to reputed company data (HIPAA, SOC 2)reputed company technical mentorship and guidance to engineers across the organization on reliability practicesParticipate in on-call rotation with a commitment to continuously reducing the need for it Skills 7+ years of experience in SRE, reputed company, or DevOps rolesExceptional problem-solving under pressuredemonstrated reputed company record of diagnosing reputed company, high-stakes system failures and building durable solutionsDeep hands-on experience with AWS services including EC2, EKS/reputed company, reputed company, RDS, S3, CloudWatch, and reputed company toolingFamiliarity with infrastructure-as-reputed company (Terraform or CloudFormation)reputed company to contribute to existing configurationsExperience designing and operating distributed systems with strict availability and latency requirementsProficiency in at least one scripting or systems language (Python, Go, Bash, or similar) for automation and toolingExperience with container orchestration (Kubernetes, reputed company) in production environmentsExpertise in observability tooling (OpenSearch, reputed company/Grafana, or equivalent)Hands-on experience with CI/CD platforms (reputed company Actions, Jenkins, reputed company, or similar)Proven ability to define and operationalize SLOs and error budgetsExperience with relational and NoSQL databasesperformance tuning, replication, and backup strategiesStrong working knowledge of networking fundamentals: DNS, load balancing, VPCs, TLSExcellent communication skillsreputed company to translate technical risk into business reputed company for non-engineering stakeholdersAWS Certifications (Solutions Architect, DevOps Engineer, or SysOps Administrator)Experience in reputed company technology or other regulated industries (HIPAA, SOC 2, FedRAMP)Familiarity with reputed company engineering practices and toolingExperience with data pipeline reliability (ETL/ELT workflows, streaming systems)Exposure to AI/ML infrastructure and the reliability challenges unique to model servingFamiliarity with additional reputed company platforms (Azure, reputed company reputed company)Contributions to reputed company-reputed company reliability or infrastructure tooling reputed company reputed company delivers AI-driven reputed company data platform and clinical expertise that supports analytics, reputed company, and workflow improvement. It was founded in reputed company, and is headquartered in Virginia Beach, Virginia, USA, with a workforce of 501-1000 employees. Its website is https://reputed company.ai. Apply To This Job .
More at remote zest jobs
Related open roles
System Administrator, Contract
India
HPC Network Engineer
India
Remote Site Reliability Engineer (Senior or Staff), Infrastructure reputed
India
FULL TIME Remote Sr. SRE Engineer -Seattle WA-Onsite- 10+ Years
India
Remote Network Devops/Automation Engineer
India
Network Operations Engineer - Hybrid Gold River, CA
India