Source description
About the role
OverviewRole: Site Reliability Engineer (SRE) Core IT Infrastructure Work mode: On-site (full Time) Experience: 6+ year's ResponsibilitiesInfrastructure Reliability & OperationsDesign, implement, and maintain highly available and fault-tolerant infrastructureEnsure reliability, performance, scalability, and security of core IT systemsMonitor system health, capacity, and performance using proactive observability practicesLead incident response, root cause analysis (RCA), and post-incident reviewsAutomation & SRE DevelopmentDevelop and maintain automation tools, scripts, and frameworks to reduce manual operationsApply Infrastructure as Code (IaC) principles using tools such as Terraform, Ansible, or CloudFormationBuild self-healing systems and automate repetitive operational tasksImprove deployment pipelines and operational workflows through engineering solutionsCollaborate with DevOps, development, and security teams to support CI/CD pipelinesEnable seamless application deployments with minimal downtimeSupport containerized and orchestration platforms (Docker, Kubernetes, OpenShift)Implement best practices for configuration management and environment consistencyMonitoring, Observability & PerformanceDesign and maintain monitoring, logging, and alerting systemsDefine and track SLIs, SLOs, and SLAsOptimize system performance, capacity planning, and cost efficiencyEnhance observability using tools such as Prometheus, Grafana, ELK, Datadog, or similarSecurity & ComplianceImplement infrastructure security best practicesCollaborate with security teams on vulnerability management and compliance requirementsEnsure secure access, identity management, and audit readinessRequired Skills & QualificationsTechnical Skills Strong experience in Linux/Unix system administrationProficiency in programming/scripting (Python, Go, Bash, Shell, or similar)Experience with cloud platforms (AWS, Azure, or GCP)Hands-on experience with containerization and orchestrationKnowledge of networking concepts (DNS, TCP/IP, load balancing, firewalls)Experience with monitoring, logging, and alerting tools OverviewRole: Site Reliability Engineer (SRE) Core IT Infrastructure Work mode: On-site (full Time) Experience: 6+ year's ResponsibilitiesInfrastructure Reliability & OperationsDesign, implement, and maintain highly available and fault-tolerant infrastructureEnsure reliability, performance, scalability, and security of core IT systemsMonitor system health, capacity, and performance using proactive observability practicesLead incident response, root cause analysis (RCA), and post-incident reviewsAutomation & SRE DevelopmentDevelop and maintain automation tools, scripts, and frameworks to reduce manual operationsApply Infrastructure as Code (IaC) principles using tools such as Terraform, Ansible, or CloudFormationBuild self-healing systems and automate repetitive operational tasksImprove deployment pipelines and operational workflows through engineering solutionsCollaborate with DevOps, development, and security teams to support CI/CD pipelinesEnable seamless application deployments with minimal downtimeSupport containerized and orchestration platforms (Docker, Kubernetes, OpenShift)Implement best practices for configuration management and environment consistencyMonitoring, Observability & PerformanceDesign and maintain monitoring, logging, and alerting systemsDefine and track SLIs, SLOs, and SLAsOptimize system performance, capacity planning, and cost efficiencyEnhance observability using tools such as Prometheus, Grafana, ELK, Datadog, or similarSecurity & ComplianceImplement infrastructure security best practicesCollaborate with security teams on vulnerability management and compliance requirementsEnsure secure access, identity management, and audit readinessRequired Skills & QualificationsTechnical Skills Strong experience in Linux/Unix system administrationProficiency in programming/scripting (Python, Go, Bash, Shell, or similar)Experience with cloud platforms (AWS, Azure, or GCP)Hands-on experience with containerization and orchestrationKnowledge of networking concepts (DNS, TCP/IP, load balancing, firewalls)Experience with monitoring, logging, and alerting tools
More at TECEZE