Padmi

Senior Site Reliability Engineer

United States · OnsitePosted 4 months ago
Infrastructure And DatabasesUnspecified
Apply at DPR Solutions

Opens the source posting on dprsolutionsinc.com

Source description

About the role

View original

Senior Site Reliability Engineer – Join Our Team .job-posting-container{font-family:Arial,Helvetica,sans-serif;line-height:1.6;color:#333;margin:0;padding:20px;background-color:#ffffff;box-sizing:border-box;} .job-posting-container *{box-sizing:border-box;} .job-posting-content{background-color:#ffffff;padding:30px;max-width:800px;margin:0 auto;box-shadow:0 0 10px rgba(0,0,0,0.1);border-radius:10px;} .job-posting-container h1,.job-posting-container h2{color:#004080;margin:15px 0;} .job-posting-container h1{font-size:22px;} .job-posting-container h2{font-size:20px;border-bottom:1px solid #ccc;padding-bottom:5px;margin-top:30px;} .job-posting-container ul{margin-left:20px;} .job-posting-container .section{margin-bottom:25px;} .job-posting-container .highlight{background-color:#eef6ff;padding:15px;border-left:4px solid #004080;margin-bottom:15px;border-radius:5px;} .job-posting-container strong{color:#333;} .job-posting-container li{margin-bottom:8px;} .job-meta{font-size:13px;color:#555;margin-top:6px;} .job-badge{display:inline-block;background:#f0f6ff;color:#004080;padding:4px 8px;border-radius:4px;font-size:12px;margin-right:8px;border:1px solid rgba(0,64,128,0.06);} .muted{color:#666;font-size:13px;} a.job-email{color:#004080;text-decoration:none;font-weight:600;} Senior Site Reliability Engineer Location: Stamford, CT Type: CTC, W2 Tax: C2C Ref: JB-OQ3H6GXJ On-site role based in Stamford, Connecticut. Pay rate: approximately $70/hr. Overview We are seeking a Senior Site Reliability Engineer to lead reliability and operational excellence for cloud-native infrastructure. In this role you will partner with engineering, security, and platform teams to design resilient architectures, implement Infrastructure as Code and CI/CD practices, and drive measurable improvements in availability, scalability and recovery readiness. Why Join Us? Work on critical cloud platforms with a focus on reliability, observability, and automation. Opportunities for professional development and industry certifications in cloud and SRE practices. Collaborative engineering culture that values documentation, structured incident learning, and continuous improvement. Meaningful impact on operations, disaster recovery readiness, and business continuity. Key Responsibilities Define and own service-level objectives (SLOs/SLIs), manage error budgets, and drive reliability initiatives. Design, implement and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR and enable proactive detection. Lead incident response improvements including on-call practices, runbooks, post-incident reviews and preventative actions. Build and operate resilient cloud architectures (multi-AZ and where applicable multi-region) with autoscaling, load balancing, backups and replication. Develop and maintain Infrastructure as Code (Terraform, CloudFormation, CDK) and standardize CI/CD pipelines for safe, repeatable deployments. Define and validate RTO/RPO targets, implement BCP/DR plans and coordinate structured DR testing and remediation activities. Qualifications 7+ years of experience in SRE, DevOps, Platform or Systems Engineering supporting production environments. Proven hands-on experience with observability platforms (metrics, logging, tracing) and distributed tracing concepts. Strong AWS experience building and operating production systems; familiarity with multi-AZ patterns and cloud governance. Practical expertise with Infrastructure as Code (Terraform and/or CloudFormation/CDK) and version-controlled infrastructure. Solid CI/CD and automation background, including pipeline design and deployment strategies. Experience designing, testing and operating BCP/DR processes with documented runbooks and recovery checklists. Experience with Kubernetes (EKS, ECS, or Kubernetes on-prem), autoscaling and container platforms. Strong Linux fundamentals, networking (DNS, TCP/IP) and troubleshooting skills; proficiency in one or more scripting languages (Python, Go, Bash, etc.). Clear technical documentation skills and ability to work effectively in a fast-paced, high-intensity environment, including participation in on-call rotations when required. Preferred Qualifications Familiarity with Azure and/or Oracle Cloud (OCI). Experience with Service Mesh, API Gateways and distributed tracing tools (OpenTelemetry experience a plus). Security and compliance experience for cloud environments (IAM, secrets management, audit logging). Experience implementing progressive delivery (blue/green, canary), feature flags and automated rollbacks. Relevant certifications (AWS Solutions Architect/DevOps, CKA/CKAD) and exposure to ArgoCD & Karpenter. Ready to Apply? If you are interested, please email a resume and brief note referencing job code JB-OQ3H6GXJ to the recruiter below. Strong candidates will be contacted for next steps promptly. Contact: kevin@dprsolutionsinc.com Phone: +1 (571) 498-9280 Take the next step in your career!

One address, no account. We’ll tell you when matching roles go live.

More at DPR Solutions

Related open roles

View all roles