Padmi

Production and Support Engineer

BangalorePosted 1 month ago
IT supportMid-level
Apply at Devon Software Services

Opens the source posting on naukri.com

Source description

About the role

View original

What are we looking for? An L2 System & Application Support Engineer with strong production operations experience across applications hosted on cloud infrastructure. As a Technical Operations Engineer, you are responsible for the monitoring, incident management, reporting, release management and continuous operational improvement. Youll partner with L3/engineering teams, stakeholders, product teams and clients to ensure highest availability of the services/features. You will be responsible for root cause identification ensuring implementation of processes to prevent the recurrences. You’ll also help us continuously improve observability, runbooks, and operational readiness. The candidate, Must have skills: 3+ years of experience in Level 2 Application support Expertise in application and backend issue triage & analysis. Should have analyzed failures, error logs, done basic analysis, gather inputs and coordinate with Level 3 Own day-to-day product operations for cloud-hosted applications, including proactive monitoring, incident triaging, outage management, and resolution while ensuring adherence to defined SLAs. Demonstrate strong analytical capabilities to diagnose issues, drive incidents to closure, and adapt to evolving tools and techniques during real-time incident response. Having worked upon support layers – L1, L2, L2.5, worked closely with the engineering L3 teams, Product teams and the Business Stakeholders Troubleshooting knowledge in AWS, containerized workloads, and distributed microservices architecture. Understanding the root causes, restoration of the services and final reporting Monitor application and infrastructure using Datadog/Logicmonitor/Newrelic/Grafana or any monitoring tool Usage of multiple tools and techniques to analyse the production issues across Applications, backend, data and infrastructure layers Correlate logs, traces and metrics to isolate the failures across multiple services, understand the impact on business and generate high quality artifacts aiding L3 investigations Support and lead release management, operational readiness, deployment support and postproduction deployment validations Collaborate with the Stakeholders, Project teams and Vendors, handle the incident triages and meetings ensuring SLA adherences, escalation flows and operational excellence Good command in Microsoft Office suite Willing to work in rotational based 24X7 multiple Shifts Willing to learn and implement automation or AIOps Features Tools (Related should be fine too) Monitoring and Observability: Datadog, ElasticSearch ELK / Open Search Dashboards, AWS Cloudwatch Cloud & Infra: AWS, Lambda, API Gateway, Open Search, Athena, Web services, API Management, AWS Cognito, EKS, SQS Incident Management ITSM tools Jira, ServiceNow Databases PostgreSQL, Mongo No SQL, SQL Good to have skills: Familiarity with Gitlab/Github, Jenkins, CI/CD Pipelines or any source code version control. Any Automation Skillset / Willing to learn Automation tools Webservices and API Management Experience in Change Management Process Connected car features understanding What will be the roles and responsibilities? Incident handling, Issue analysis, Investigate & Gather necessary data/insights Incident escalation based on workflows Own L2 incident handling: triage, isolate, restore, and document; run bridges for P1/P2 and steer the right resolver group. Monitor & act on CloudWatch/Datadog/ELK dashboards & autoalerts using runbooks; tune noise vs. signal. Deepdive production issues across Lambda/API Gateway/OpenSearch/Athena; create actionable debug artifacts (timelines, queries, log snippets, correlation IDs). Partner with L3: supply highquality inputs (repro steps, traces, request samples, dashboards), verify fixes, and drive postincident actions. Operational hygiene: maintain runbooks, workflows, SOPs, and knowledge articles; contribute to SLA reporting and weekly ops reviews. Crossgeo coordination: collaborate with business, SupportDesk, QA, SecOps, vendors; close tickets endtoend. Continuous improvement: raise kaizens to reduce toil (alerts, autoremediations, dashboards, cost/perf optimizations in Athena/OpenSearch). In case of priority incidents , bridge the calls with multiple stakeholders and ensure resolution stakeholder is identified and assigned to incident Monitoring Dashboard of cloud applications Monitoring and acting based on the auto-alerts generated from system using runbooks Monitoring and Handling of the slack Datadog alerts Report creation and publishing to required stake holders Responsible for maintaining the procedures/run books, workflows, work instructions used by the ops team Fulfil any ad-hoc data or report request queries from different functional groups Shifts Rotational 247 (multiple shifts), including weekends oncall / holidays support as rostered

One address, no account. We’ll tell you when matching roles go live.