Padmi
Unlimit logo
Unlimit

payment processing · banking as a service (BaaS)

Site Reliability Engineer (SRE)

Belgrade · OnsitePosted 6 months ago
InfrastructureUnspecifiedFull Time
Apply at Unlimit

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Platform reliability & operations:

Ensure the availability, resilience, and performance of the platform and supporting services.

Own and improve incident management, including troubleshooting, escalation handling, and follow-ups aligned to SLAs.

Participate in an on-call rotation, supporting production systems and driving reliability improvements from real incidents.

Infrastructure engineering (Linux / Cloud / Kubernetes):

Design, deploy, configure, and manage Linux-based system architecture across environments.

Build and support platform implementations using AWS and other cloud technologies (compute-centric services and related infrastructure).

Design and implement large and complex technology projects, from initial design through production rollout and operational handover.

Support and maintain Kubernetes-based workloads and platform components.

Automation & Infrastructure as Code:

Build tooling and solutions to automate recurring operational tasks.

Use Infrastructure as Code (IaC) to standardize and scale: Terraform for provisioning , Ansible for configuration management and automation

Improve reliability by reducing manual steps and enabling repeatable deployments.

CI/CD & developer enablement:

Manage and maintain CI/CD pipelines across 20+ repositories spanning multiple technology stacks.

Partner with Engineering teams to improve build/release consistency, pipeline reliability, and deployment safety.

Observability & operational readiness:

Implement and enhance monitoring, logging, and alerting , using tools such as: Prometheus , Grafana, Zabbix , Splunk, PagerDuty (or equivalent incident alerting/response tooling).

Use metrics and incident learnings to reduce noise, improve signal, and shorten time-to-detect/time-to-recover.

Documentation & standards:

Produce clear, formal documentation including: Configuration standards, Troubleshooting runbooks, Infrastructure and architecture design documentation.

Contribute to internal standards that improve consistency, security, and operational maturity.

More at Unlimit

Related open roles

View all roles