Source description
About the role
Role & responsibilities Provide hands on technical and functional troubleshooting for a suite of applications within TST domain. Build up technical and functional subject matter expertise on the applications/platforms being supported including business flows, the application architecture and the hardware configuration. Work to build awareness of Prod KPIs and help drive initiatives to achieve the same. Identify pro-actively opportunities for automation and toil reduction. Drive to resolution service requests submitted by the application end users to the best of L2 ability and escalate any issues that cannot be resolved to L3. Conduct review on our monitoring platforms for alert rationalization as well as addition of alerts to our real time monitoring tools to ensure application SLAs are achieved and maximum application availability (up time). Ensure all knowledge is documented and that support runbooks and knowledge articles are kept up to date. Approach support with a proactive attitude, working to improve the environment before issues occur. Update the RUN Book and KEDB as & when required. Participate in all BCP (DR, EDR, SSRTO tests) and component failure tests based on the run books. Understand flow of data through the application infrastructure. It is critical to understand the dataflow so as to best provide operational support. Drive and own accountability for managed risk and better control KPIs. Team mentoring and coaching of a global team. SRE mindset – implementation knowledge of SLO and SLI’s Preferred candidate profile 14+ years of experience in IT in large corporate environments, specifically in controlled production environments or in Financial Services Technology in a client-facing function. Experience in securities services\asset management\payments domains will be a plus. Expert level knowledge of GCP and cloud native architecture and applications Intermediate level knowledge of AI, LLM working principles & agentic AI. Working knowledge of Scripting - UNIX shell and PowerShell, PERL, Python Flexibility for incident callouts (after office hours as well as weekends). Flexible weekend cover on a rotation basis. Programming Language - Understanding of Java Operating systems – Understanding of UNIX, LINUX and the underlying infrastructure environments. Understanding of Middleware - (e.g. MQ, Kafka or similar) Understanding of WebLogic, Webserver environment - Apache, Tomcat Understanding of how RDBMS work - Oracle, MS-SQL, Sybase, No SQL Understanding of Batch Monitoring tools - Control-M /Autosys Understanding of Monitoring Tools – GCO, Geneos or App Dynamics or New Relic or Dynatrace or Grafana ITIL Service Management framework such as Incident, Problem, and Change processes
More at Deutsche Bank