Source description
About the role
SRE Engineer 100% Remote Job Summary We are looking for a Site Reliability Engineer (SRE) to ensure the reliability, availability, and performance of our applications and infrastructure. The ideal candidate will work closely with development and operations teams to automate processes, monitor systems, troubleshoot issues, and improve platform stability. Key Responsibilities Monitor application and infrastructure health using monitoring and alerting tools. Troubleshoot production issues and perform root cause analysis. Automate routine operational tasks using scripting and Infrastructure as Code (IaC). Manage cloud infrastructure and deployment pipelines. Improve system reliability, scalability, and performance. Support CI/CD processes and release management. Implement and maintain backup, disaster recovery, and high-availability solutions. Collaborate with development teams to improve application resilience and operational efficiency. Maintain system documentation and operational runbooks. Participate in on-call support and incident response. Required Skills Understanding of Linux/Unix operating systems. Knowledge of cloud platforms such as AWS, Azure, or Google Cloud. Experience with container technologies like Docker and Kubernetes. Familiarity with CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI. Knowledge of Infrastructure as Code tools (Terraform, Ansible, or CloudFormation). Experience with monitoring and logging tools such as Prometheus, Grafana, ELK Stack, Datadog, or Splunk. Basic scripting skills in Bash, Python, or PowerShell. Understanding of networking, DNS, load balancing, and security best practices. Strong problem-solving and troubleshooting skills. .
More at Apetan Consulting