Source description
About the role
What youll do:Support reliability, availability, and performance of enterprise platforms, with initial focus on Supplier Lifecycle Platform (SLP) and associated systemsMonitor system health using key service metrics (latency, error rates, traffic, and saturation) and respond to issues impacting production stabilityExecute incident, problem, and change management processes to restore service and prevent recurrenceDevelop, enhance, and maintain automation, workflows, and monitoring solutions to reduce manual effort and improve system reliabilityParticipate in design, testing, and deployment activities to ensure solutions meet reliability and performance expectations across the lifecycleCollaborate with development, integration, and business teams to support system enhancements, onboarding workflows, and data flows across platformsCreate and maintain system documentation, runbooks, and operational playbooks to support consistent execution and knowledge transferPerform root cause analysis for incidents and implement corrective and preventive actions to improve long-term system stabilityContribute to backlog management for reliability improvements, defect remediation, and continuous system optimizationSupport secure development and operational practices aligned with enterprise standards across the engineering lifecycleSupport supplier onboarding and qualification processes by ensuring stability and effectiveness of onboarding workflows and associated system interactionsParticipate in test strategy execution, validation activities, and release readiness to ensure solutions meet operational and performance expectationsAnalyze system performance and operational data to identify trends, root causes, and opportunities for improvementDevelop automation, monitoring, and reporting solutions to improve system reliability and reduce manual interventionCollaborate with cross-functional teams to understand business processes and ensure solutions align with procurement and supplier management workflowsEnsure solutions and operational practices align with enterprise governance, security, and compliance standardsQualifications:Bachelor's Degree from an accredited institution58 years of experience in IT, software engineering, system support, or related technical rolesSkills:Experience supporting production systems, including incident resolution and continuous improvementExperience working in Agile or Scrum-based delivery environmentsSolid understanding of system reliability concepts including availability, performance, latency, monitoring, and incident managementExperience with system integration concepts, APIs, and data flows across enterprise platformsExperience with automation, monitoring tools, logging, and dashboards for system health and performanceFamiliarity with cloud-based platforms and service models (e.g., SaaS, PaaS, microservices, APIs)Ability to write or support development in one or more programming or scripting languages (e.g., Python, JavaScript, SQL, C#)Understanding of ITSM processes including incident, problem, and change managementComfortable with the English language (written and verbal)Effective communication skills with the ability to explain technical issues clearlyStrong problem-solving and analytical thinkingAbility to manage multiple priorities and respond effectively in high-pressure situationsCollaborative mindset, working across global teams and functional areasContinuous learning mindset and adaptability to changing technologies and business needs What youll do:Support reliability, availability, and performance of enterprise platforms, with initial focus on Supplier Lifecycle Platform (SLP) and associated systemsMonitor system health using key service metrics (latency, error rates, traffic, and saturation) and respond to issues impacting production stabilityExecute incident, problem, and change management processes to restore service and prevent recurrenceDevelop, enhance, and maintain automation, workflows, and monitoring solutions to reduce manual effort and improve system reliabilityParticipate in design, testing, and deployment activities to ensure solutions meet reliability and performance expectations across the lifecycleCollaborate with development, integration, and business teams to support system enhancements, onboarding workflows, and data flows across platformsCreate and maintain system documentation, runbooks, and operational playbooks to support consistent execution and knowledge transferPerform root cause analysis for incidents and implement corrective and preventive actions to improve long-term system stabilityContribute to backlog management for reliability improvements, defect remediation, and continuous system optimizationSupport secure development and operational practices aligned with enterprise standards across the engineering lifecycleSupport supplier onboarding and qualification processes by ensuring stability and effectiveness of onboarding workflows and associated sys
More at Eaton