Source description
About the role
Title: Production support engineer msrt platform/infrastructure
Description:
We are seeking a highly motivated, self-starting individual eager to learn and grow as a key member of our global MSRT Platform/Infrastructure Application Production Support team within a leading financial org. This role is crucial for ensuring the stability, performance, and availability of critical customer-facing applications. The ideal candidate will possess a strong technical foundation, a passion for automation, and a proactive approach to problem-solving and continuous improvement, operating primarily in a back-end capacity without direct user interaction.
responsibilities
-
Key Responsibilities
-
As a Production Support Engineer on the MSRT Platform/Infrastructure team, your responsibilities will include:
-
Application Operations & Stability:
-
Adhere to and utilize internal application production support standards, tools, and processes for daily operations.
-
Manage and maintain the primary customer-facing application environment, including responsive web applications and platforms with mobility support.
-
Proactively go through tickets, performing incident and problem resolution, and conducting in-depth root cause analysis of outages.
-
Liaise closely with development teams to understand "why there is a problem" and collaboratively drive application stability.
-
Deploy application updates, with a strong focus on using Ansible for deployments.
-
Audit and review infrastructure capacity, planning and implementing necessary enhancements, and developing application-specific instrumentation tools.
-
Individually manage the application lifecycle for web platforms, middleware, and databases.
-
Follow all best practices, policies, and procedures established by the company.
-
Monitoring & Alerting:
-
Monitor applications with Application Performance Monitoring (APM) tools.
-
Lead efforts in the migration of all monitoring to Dynatrace, building new dashboards and alerting systems for applications and infrastructure. This is not just plain application support; it involves strategic migration work.
-
Thoroughly go through logs to identify issues, trends, and performance bottlenecks.
-
Automation & Process Improvement:
-
Lead and contribute significantly to efforts in automation, deployment, and configuration management.
-
Always brainstorm about automations and other projects, driving process improvements by developing and maintaining/enhancing tools to assist with production support.
-
Strategic Planning & Collaboration:
-
Contribute to planning efforts for disaster recovery, capacity expansion, and system upgrading.
-
Research new promising technologies, strategies, and ways to solve complex technical issues.
-
Work across IT functions and collaborate effectively with various teams, including the network team, to deliver complete technical solutions.
-
Establish strong partnerships with application-specific developers, infrastructure teams, and Subject Matter Experts (SMEs).
-
On-Call & Incident Management:
-
Participate in an on-call rotation, driving incident resolution and improving platform resiliency.
-
Technical Environment & Required Skills
-
This role operates within a diverse and complex technical environment. The following lists represent the technologies you will be working with; specific proficiency levels are highlighted where critical:
-
Platform/Operating Systems:
-
Strong proficiency in Unix/Linux and Windows Operating Systems, including advanced troubleshooting skills on these platforms.
-
Networking:
-
Solid understanding of networking concepts including DNS, DMZ, Networking, Load balancing (LTM/GTM/F5), Firewalls/Proxies, SFTP/FTPs, Active Directory, and SharePoint. This includes understanding the functions of LPM and GTMs. Ability to effectively work with the network team is essential.
-
Scripting & Automation:
-
Required: Strong experience with Ansible (5-7 years preferred for deployments and automation), Python, and Shell scripting.
-
Helpful: Perl, and general proficiency in any scripting language for automation.
-
Security:
-
Required: Experience with SSL Certificates.
-
Helpful: Kerberos.
-
Environment: SAML, OAuth2, RBAC, and other security technologies.
-
Databases:
-
Helpful: SQL, particularly for connecting and configuring applications to databases.
-
Environment: Oracle, Sybase.
-
Monitoring Tools:
-
Critical: Experience with Dynatrace, especially given the ongoing migration and focus on building new monitoring solutions.
-
Environment: Geneos ITRS, Splunk, AppDynamics.
-
Middleware:
-
Environment: MQ, WebLogic, Apache, WebSphere, Splunk, Hadoop, Big Data, Oracle Coherence, Vertica.
-
General Programming Knowledge:
-
Environment: C/C++, Java, Python.
-
Scheduler:
-
Environment: Autosys and Unix level scheduling or equivalent Batch scheduling products.
-
Other Technologies:
-
Environment: OpenShift.
-
Desired Qualifications
-
Understanding of Sales, Research, Capital Markets, and Trading systems (financial knowledge is helpful).
-
A deep understanding of client-server and distributed systems architecture (Clustering, High Availability, etc.).
-
A working knowledge and understanding of security technologies to support applications (SSL, Kerberos, SAML, OAuth2, RBAC etc.), Windows Web servers, and Linux.
qualifications
- Support, Networking, Security, Application Support, Automation
skills: Support, Networking, Security, Application Support, Automation
More at 3B Staffing