Padmi
BlackRock logo
BlackRock

asset management · Aladdin platform

Vice President, DevOps Engineer, Lead Engineer

IndiaPosted 3 months ago
Infrastructure And DatabasesStaff+Full Time; Regular
Apply at BlackRock

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: As a Data Platform Cloud/DevOps Engineer in the Data Engineering team at BlackRock, you will be responsible for designing, building, and maintaining the cloud-native infrastructure that powers Aladdin's Enterprise Data Platform. Your primary focus will be enabling data engineers, AI engineers, and application developers by providing scalable, reliable, and cost-efficient infrastructure for data processing, AI/ML workloads, and analytics services. Key Responsibilities: - Design, deploy, and manage cloud-native infrastructure across AWS, Azure, and private clouds - Implement Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation for repeatable, auditable deployments - Manage Kubernetes clusters for scalable, reliable, and secure application and data workloads - Deploy and configure service mesh, HashiCorp Vault, cert-manager, and other Kubernetes-native frameworks - Design and implement network architectures including VPCs, load balancers, and ingress/egress controls - Deploy and configure LLM serving platforms like MCP/agent orchestrators, chatbots, vector embedding services and secured API gateways for generative AI applications CI/CD and Automation: - Build and maintain CI/CD pipelines using ArgoCD, Azure DevOps, Jenkins, and GitHub Actions - Implement GitOps workflows for automated, auditable infrastructure and application deployments - Automate repetitive operational tasks using Python and Bash to improve team efficiency and reduce manual errors - Develop self-service infrastructure provisioning capabilities for engineering teams - Maintain version control best practices and collaborative development workflows - Build and maintain MLOps CI/CD pipelines for automated model deployment to production environments Site Reliability Engineering (SRE): - Implement monitoring, logging, and observability solutions using Prometheus, Grafana, ELK Stack, and Datadog - Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for data platform services - Build automated alerting systems to proactively detect infrastructure issues and performance degradation - Perform capacity planning and performance tuning for production infrastructure - Conduct reliability analysis and implement preventive measures to improve system uptime - Collaborate with operational teams on incident escalation and system reliability improvements - Implement chaos engineering practices to test infrastructure resilience and fault tolerance Cloud Cost Optimization and FinOps: - Monitor and optimize cloud infrastructure costs across AWS, Azure, and private cloud environments - Right-size compute, storage, and networking resources based on utilization metrics and cost-performance analysis - Develop cost dashboards and reports to provide visibility into infrastructure spending trends - Collaborate with finance and engineering teams on cloud budget planning and forecasting - Evaluate and recommend cost-effective architectural alternatives (e.g., spot instances, reserved capacity, serverless options) Qualification Required: - Expert-level experience with AWS, Azure, or GCP cloud platforms and services - Proficiency with Infrastructure as Code tools (Terraform, Ansible, CloudFormation) - Templating with Helm, ArgoCD, Ansible, and Terraform - Deep knowledge of Kubernetes (K8s) APIs, controllers, operators, and stateful workloads - Understanding of the K8s Operator Pattern -- comfort and courage to wade into (predominantly golang based) operator implementation Role Overview: As a Data Platform Cloud/DevOps Engineer in the Data Engineering team at BlackRock, you will be responsible for designing, building, and maintaining the cloud-native infrastructure that powers Aladdin's Enterprise Data Platform. Your primary focus will be enabling data engineers, AI engineers, and application developers by providing scalable, reliable, and cost-efficient infrastructure for data processing, AI/ML workloads, and analytics services. Key Responsibilities: - Design, deploy, and manage cloud-native infrastructure across AWS, Azure, and private clouds - Implement Infrastructure as Code (IaC) using Terraform, Ansible, and CloudFormation for repeatable, auditable deployments - Manage Kubernetes clusters for scalable, reliable, and secure application and data workloads - Deploy and configure service mesh, HashiCorp Vault, cert-manager, and other Kubernetes-native frameworks - Design and implement network architectures including VPCs, load balancers, and ingress/egress controls - Deploy and configure LLM serving platforms like MCP/agent orchestrators, chatbots, vector embedding services and secured API gateways for generative AI applications CI/CD and Automation: - Build and maintain CI/CD pipelines using ArgoCD, Azure DevOps, Jenkins, and GitHub Actions - Implement GitOps workflows for automated, auditable infrastructure and application deployments - Automate repetitive operational tasks using P

One address, no account. We’ll tell you when matching roles go live.

More at BlackRock

Related open roles

View all roles