Padmi

SAP DevOps SRE Observability Engineer

Remote · United StatesPosted 1 month ago
Infrastructure And DatabasesUnspecified
Apply at 3B Staffing

Opens the source posting on 3bstaffing.com

Source description

About the role

View original

Job Type: Remote Long Term Contract

Interview: Virtual, 1-2 rounds

Possible assessment required prior to client submission. Make sure that candidates are serious about taking a technical assessment for this role.

JOB DESCRIPTION:

Duties:

Provides counsel and advice to top management on significant Application Development matters, often requiring coordination between organizations.

Acts as the principal designer for complex major systems and their subsystems utilizing a thorough understanding of available technology, tools and existing designs.

Provides comprehensive consultation to business unit and IT management and staff at the highest technical level on all phases of application programming and processes for diverse development platforms, computing environments (e.g., host based, distributed systems, client server, software, hardware, technologies and tools, etc.).

Works closely with client and IT management and staff to identify application development solutions, new or modified programs, reuse of existing code through the use of program development software alternatives, or integration of purchased solutions or a combination of the available alternatives.

Researches and evaluates alternative solutions and recommends the most efficient and cost effective application programming solution. May code new or modified programs, reuse existing code through the use of program development software alternatives and/or integrates purchased solutions.

Documents, tests, implements and provides on-going support for the applications.

Focuses on providing thought leadership and technical expertise across multiple disciplines. Recognized internally as "the go-to person" for the most complex Application Development assignments.

Required Skills:

Application Development

System Design

Application Programming

Diverse Development Platforms

Host-Based Computing

Additional Skills:

Distributed Systems

Solution Integration

Application Testing

Application Implementation

Application Support

Technical Expertise

Thought Leadership

Client-Server Computing

Software Technologies

Hardware Technologies

Program Development Software

Documentation

Top 4 Skills are in bold below:

  1. Hands-On Observability Engineering & Actionable Alerting

This is not a traditional passive infrastructure monitoring role; the candidate must build an active, end-to-end observability ecosystem.

Core Tooling: Deep technical proficiency in Dynatrace (APM and Synthetic Monitoring creation) and Splunk (log ingestion, query optimization, trace/metric correlation).

SAP-Specific Telemetry: Experience leveraging SAP Cloud ALM (CALM) and Tealeaf for Real User Monitoring (RUM).

Key Deliverable: The ability to move from basic monitoring to actionable alerting—reducing alert fatigue by designing custom operational dashboards that correlate application traces, infrastructure logs, and business metrics to track core SLIs and enforce SLOs (availability, response times, job success rates).

  1. Deep Troubleshooting Across Distributed SAP & Supply Chain Landscapes

The engineer must serve as the technical bridge across a complex, multi-tiered enterprise architecture.

Platform Fluency: Working knowledge of SAP S/4HANA , SAP EWM (Extended Warehouse Management), Commerce Cloud , and Fiori .

Integration Tracing: Proven ability to navigate distributed layers to isolate bottlenecks. This includes tracing cross-platform integrations, specifically EDI transactions, IDocs, RFCs , and data flowing through SAP Integration Suite / CPI .

Domain Relevance: Understanding how technical errors impact core supply chain processes like Order-to-Cash (O2C) and Procure-to-Pay (P2P) in a specialty distribution environment.

  1. SRE Incident Management & Root Cause Analysis (RCA) Capabilities

The role requires an engineer who can own production stability during high-stakes outages and drive long-term systemic fixes.

Incident Response: Experience participating in on-call rotations, coordinating major incident response, and rapidly driving down MTTD (Mean Time to Detect) and MTTR (Mean Time to Resolve).

Post-Incident Governance: Strong communication skills to lead blameless postmortems/post-incident reviews, rigorously documenting root causes and translating findings into architectural or operational improvements to extend MTBF (Mean Time Between Failures).

  1. Automation-First Mindset & Proactive Reliability Engineering

A core expectation is shifting the team away from reactive firefighting toward proactive self-healing and automated operations.

Proactive Validation: Writing automated synthetic monitoring scripts and continuous health checks to simulate user journeys and catch degradations before end-users or warehouse operators are impacted.

Toil Reduction: Automating repetitive operational and Level 1/Level 2 support tasks through scripting and modern DevOps practices to drive continuous service improvement across the SAP ecosystem.

One address, no account. We’ll tell you when matching roles go live.

More at 3B Staffing

Related open roles

View all roles