Source description
About the role
Description: We are looking for a skilled and pragmatic DevOps Engineer to own and evolve our infrastructure across the EMEIA region. This is a dual-horizon role: you will keep our existing VM-based systems healthy while leading a greenfield effort to design and build the managed environment that those solutions will migrate onto. A significant proportion of what we build is produced rapidly using AI-assisted, structured development. That means our solutions can move from idea to deployment faster than ever, and our infrastructure needs to keep pace. We need someone who thrives in a fast-moving, ambiguous environment, can absorb change quickly, and treats adaptability as a core part of the job rather than an occasional demand. The new managed environment is most likely to be based on Kube client's internal Kubernetes (EKS) deployment though the final architecture will be a team decision and client specific AWS remains an option for workloads requiring greater control. You will help inform that decision and then own the build-out, regardless of which direction is chosen. You will work closely with data engineers, developers, and analysts, acting as the infrastructure backbone for a team that moves quickly and expects you to move with it. The role also involves working directly with third-party vendors who support some of the tools being deployed, and collaborating with teams outside of EMEIA including WorldWide to align on standards, share solutions, and resolve cross-regional dependencies. KEY RESPONSIBILITIES Platform Migration & Environment Design Lead the design and build-out of a new managed container environment to replace existing VM-based infrastructure the most likely candidate is Kube (client's internal Kubernetes/EKS cluster), but the final decision will be made collaboratively as a team Contribute meaningfully to the environment selection decision: weigh trade-offs between managed solutions (Kube) and more directly controlled alternatives (client specific AWS), considering maintenance overhead, operational control, and team capability Own the migration of existing VM-based workloads onto the new platform, managing sequencing, risk, and continuity of service throughout Establish and maintain the standard workflow for deploying solutions: build locally containerise publish to Kube configure connectivity to client internal system dependencies Client Internal Networking & Connectivity Configure and maintain networking between Kube and client's internal systems, including Shield, Snowflake, Floodgate, and any other platform dependencies the team relies on Own namespace and compute provisioning on the shared Kube cluster, ensuring workloads are appropriately isolated and correctly configured Manage credentials, service accounts, and access controls across the full connectivity chain from container to downstream service Act as the go-to expert on how things connect within client's internal network topology Infrastructure Management Own and manage cloud infrastructure across EMEIA using internal cloud tooling (client cloud and connected systems including Shield) Manage certificates, firewalls, resource pools, networking, and access controls Ensure infrastructure is appropriately sized, resilient, and cost-efficient Maintain accurate documentation of infrastructure topology and configuration VM Provisioning & Automation (Existing Estate) Maintain and operate existing virtual machines, primarily on RHEL, while migration to the new environment is in progress Build and maintain standardised, repeatable provisioning processes (e.g. via Ansible, Terraform, or equivalent IaC tooling) Manage package deployment, software repositories, databases, and web servers Own the patching and update lifecycle for managed systems Monitoring & Reliability Implement and maintain monitoring, alerting, and observability across both the existing VM estate and the new container environment Proactively identify risks, bottlenecks, and failure patterns before they impact users Define and track appropriate SLIs/SLOs for critical services Conduct post-incident reviews and drive lasting improvements Supporting AI-Augmented Development A large proportion of the solutions you will support are built rapidly using structured AI-assisted development you must be comfortable working with codebases and configurations that evolve quickly, may not have deep documentation histories, and may have been substantially generated with AI tooling Provide the infrastructure scaffold that allows AI-assisted solutions to move from local development to production reliably and safely Be a pragmatic partner to developers: unblock deployment quickly .
More at Diverse Lynx India