Source description
About the role
Role Overview: As a Staff Site Reliability Engineer (SRE) specializing in S3 storage at our Bengaluru location, your main responsibility will be to ensure the system reliability and monitoring of S3 storage clusters. You will play a crucial role in designing and implementing monitoring, alerting, and automation processes to achieve a 99.99%+ uptime. Your expertise in tools like Prometheus, Grafana, or Catchpoint will be essential for tracking performance metrics, capacity utilization, and anomaly detection. Key Responsibilities: - System Reliability and Monitoring: Design and implement monitoring, alerting, and automation for S3 storage clusters to achieve 99.99%+ uptime. Utilize tools like Prometheus, Grafana, or Catchpoint for tracking performance metrics, capacity utilization, and anomaly detection. - Capacity Planning and Scaling: Forecast storage needs based on data growth trends, proactively scale S3 buckets, lifecycle policies, and multi-region replication to support up to 150 PB+ capacities. - Incident Management: Lead on-call rotations, troubleshoot storage-related incidents, and perform root cause analysis using methodologies like blameless post-mortems. - Automation and Infrastructure as Code: Develop and maintain automation scripts for provisioning, configuring, and managing S3 resources, including security policies, encryption, and access controls. - Performance Optimization: Optimize data ingestion, retrieval, and archival processes to handle high-throughput workloads, reducing costs through intelligent tiering and data compression. - Security and Compliance: Ensure storage systems comply with data protection standards, implementing features like bucket policies, versioning, and encryption at rest/transit. - Collaboration and Innovation: Work with data engineering, AI, and energy teams to integrate S3 with other systems. Contribute to open-source tools or internal projects for advanced storage solutions. - Documentation and Knowledge Sharing: Maintain runbooks, contribute to knowledge bases, and mentor junior engineers on best practices for object storage reliability. Qualification Required: - Experience: 5+ years in SRE, Dev Ops, or systems engineering roles, with at least 3 years focused on AWS S3 or similar object storage. Proven track record managing large-scale storage systems. - Technical Skills: Expertise in AWS services and infrastructure tools. Proficiency in scripting/programming for automation and tooling. Strong understanding of distributed systems, networking, and storage concepts. Experience with monitoring and logging tools. - Soft Skills: Excellent problem-solving abilities, strong communication skills, and a collaborative mindset. - Education: Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience). Role Overview: As a Staff Site Reliability Engineer (SRE) specializing in S3 storage at our Bengaluru location, your main responsibility will be to ensure the system reliability and monitoring of S3 storage clusters. You will play a crucial role in designing and implementing monitoring, alerting, and automation processes to achieve a 99.99%+ uptime. Your expertise in tools like Prometheus, Grafana, or Catchpoint will be essential for tracking performance metrics, capacity utilization, and anomaly detection. Key Responsibilities: - System Reliability and Monitoring: Design and implement monitoring, alerting, and automation for S3 storage clusters to achieve 99.99%+ uptime. Utilize tools like Prometheus, Grafana, or Catchpoint for tracking performance metrics, capacity utilization, and anomaly detection. - Capacity Planning and Scaling: Forecast storage needs based on data growth trends, proactively scale S3 buckets, lifecycle policies, and multi-region replication to support up to 150 PB+ capacities. - Incident Management: Lead on-call rotations, troubleshoot storage-related incidents, and perform root cause analysis using methodologies like blameless post-mortems. - Automation and Infrastructure as Code: Develop and maintain automation scripts for provisioning, configuring, and managing S3 resources, including security policies, encryption, and access controls. - Performance Optimization: Optimize data ingestion, retrieval, and archival processes to handle high-throughput workloads, reducing costs through intelligent tiering and data compression. - Security and Compliance: Ensure storage systems comply with data protection standards, implementing features like bucket policies, versioning, and encryption at rest/transit. - Collaboration and Innovation: Work with data engineering, AI, and energy teams to integrate S3 with other systems. Contribute to open-source tools or internal projects for advanced storage solutions. - Documentation and Knowledge Sharing: Maintain runbooks, contribute to knowledge bases, and mentor junior engineers on best practices for object storage reliability. Qualification Required: - Experience: 5+ years in SRE,
More at Tesla
Related open roles
Sr. Software Engineer, Vehicle Pricing, Digital Experience
San Francisco Bay Area · Onsite
Sr. Fullstack Developer (m/w/d)
Prüm, Rhineland-palatinate · Onsite
Software Engineer, C++ Generalist, AI Systems & Infrastructure
San Francisco Bay Area · Onsite
Staff Manufacturing Test Software Development Engineer, Optimus
San Francisco Bay Area · Onsite
Product Analyst, Megapack Sales Tools
San Francisco Bay Area · Onsite
Diagnostics Engineer, Autonomous Driving
San Francisco Bay Area · Onsite