Source description
About the role
DevOps Platform Engineer with our client in the financial industry located in Charlotte, NC and New York, NY. This is a 12 + month contract position.
Responsibilities
-
Collaborate with team and with partners in QSDG and Platform to define, build, test, and deploy platform meeting requirements
-
Define & enforce standards & best practices related to platform management
-
Evaluate third-party products to meet scalability, resiliency, and performance
-
Build new or leverage existing platforms (Lab, SDLC) by automating setup, installation, verification, monitoring & provisioning processes
-
Maintain a central, version controlled, inventory of all environments, including their current versions and configuration settings
-
Plan & allocate environments to teams depending on their delivery lifecycle
-
Analyze data to identify and proactively address environment-related issues
-
Work with project teams to manage costs & improve efficiency of environments
-
Partner closely with Prod Support and Engineering to deploy & support applications
Requirements
-
7-10 years in similar roles; preferably in the financial industry
-
Higher education in IT field or relevant previous work experience
-
Prior experience designing, implementing, and maintaining end to end environments, from POC to production
-
Deep understanding of hardware, software, network, data & application configuration
-
DevOps processes and CICD tooling (Jira, Git/Bitbucket, Jenkins, Datival, Artifactory, Ansible), orchestration & automation
-
Multi-tier (Python based) web application stack microservices/serverless/loosely coupled architecture
-
Mix of on-premises and cloud based, containerized (Docker/Kubernetes/OpenShift) deployment models
-
Familiarity with no-SQL (MongoDB) and relational (SQL Server/Oracle) databases, and other various forms of Object, Vector, and file stores
-
Unix scripting, SQL, work scheduling tools
-
Setting up infrastructure monitoring & reporting for GPU/CPU & memory consumption, inference latency and model performance
-
Performance profiling & optimization techniques to maximize performance & resource consumption / throughput and minimize latency
-
Load balancing, high availability & backup recovery strategies/techniques
-
Ability to communicate effectively to a wide range of audience (business stakeholders, developer & support teams)
-
Meticulous & highly organized
-
Adaptable to shifting & competing priorities
-
Skilled at delegating, mentoring & setting expectations
-
Critical thinking skills to diagnose & resolve complex issues
-
Desired skills:
-
Familiarity with AI & Deep learning, modeling techniques, Generative AI application stack
-
Proficiency in Python and familiarity with AI frameworks (TensorFlow/PyTorch)
-
GPU cluster management (CUDA/Kubernetes), auto-scaling & scheduling (Triton Inference Server)
More at 3B Staffing