Padmi
Vonage logo
Vonage

Communications APIs · Cloud Contact Centers

Staff Platform Engineer (Hands-on, IDP, K8)

BangalorePosted 5 months ago
Software engineeringStaff+Full Time, Permanent
Apply at Vonage

Opens the source posting on naukri.com

Source description

About the role

View original

Why this role matters: As a Staff Platform Engineer, you are the "engineers engineer" and a primary architect of Vonage s engineering culture. You operate with complete autonomy, designing the ecosystem that defines how hundreds of developers interact with the cloud. You will lead the long-term technical roadmap for our Cloud Native Kubernetes platform , balancing massive-scale infrastructure management with the evolution of our Internal Developer Portal . Your mission is to eliminate systemic friction through software engineering and AI-driven operations, ensuring our global production APIs are resilient, cost-efficient, and secure by design. Your key responsibilities: Autonomous Technical Leadership: Operate with full independence to scope, size, and execute large-scale initiatives. You are the ultimate technical authority and escalation point, expected to deliver results without requiring assistance from others. Predictive Problem Solving: Use deep systems knowledge to foresee architectural bottlenecks and operational issues before they impact production. You proactively design solutions for future scale, ensuring the platform stays ahead of business demands. Platform Strategy & IDP Ownership: Lead the organizational roadmap for the platform while working closely with the Architecture Team to ensure alignment with global standards. Act as the strategic owner for the Internal Developer Portal (IDP) - managing stakeholders, UX, and onboarding to provide a "single pane of glass" for costs, reporting, and self-service. Mentorship & Multiplier: Act as the primary mentor for junior and senior engineers. You provide the support and guidance needed to raise the engineering bar, fostering a culture of self-sufficiency and technical excellence. GitHub Actions Expertise & CI/CD Modernization: Serve as the subject matter expert for GitHub Actions , designing enterprise-grade reusable workflows and managing high-performance runner infrastructure. Lead the roadmap to migrate away from legacy Jenkins toward modern, automated CI/CD and "no-code" automation. IaC Culture & Cloud-Native Infrastructure: Drive the evolution of our IaC culture using tools like Terraform and Crossplane , moving the organization toward a Kubernetes-native management style. Reliability & Deep Systems Debugging: Respond effectively to service failures by diving into Unix/Linux OS internals (filesystems, system calls, networking). Lead the team in blame-free Root Cause Analysis. What youll bring Experience: 12+ years of progressive experience in software engineering, systems design, and cloud architecture. Self-Direction: A proven track record of delivering complex, multi-region infrastructure projects from concept to completion without oversight or assistance. Linux & Systems Mastery: Expert-level knowledge of Unix/Linux internals and networking. You troubleshoot "black box" failures at the system level and perform deep complexity analysis. Kubernetes Excellence: Expert-level experience managing massive, high-traffic, multi-tenant Kubernetes environments (EKS/GKE). CI/CD Transformation: Expertise in GitHub Actions (custom actions, API integration, and self-hosted runners) coupled with deep experience in Jenkins architecture and migration strategies. Coding & Automation: Professional-grade proficiency in Go (preferred) or Python . You approach infrastructure with a software developer s mindset - focusing on clean code and integrated bot solutions. Artifact & Security Governance: Ownership of Harbor/Artifactory lifecycles, ensuring cost-governance and security compliance. Observability & AI: Experience designing observability infrastructure (VictoriaMetrics, Thanos, Prometheus, Grafana) and applying AIOps to automate incident response, predictive capacity planning , and automated FinOps resource optimization . Whats required for application Systems Thinking: The ability to troubleshoot complex, distributed system failures across networking, storage, and application layers. FinOps & Efficiency: A proven track record of optimizing large-scale cluster topologies to balance top-tier performance with aggressive cost-efficiency goals. Observability Architecture: Experience designing observability stacks using open-source tools like Prometheus, Grafana, VictoriaMetrics, or Thanos . Cloud Native Fluency: Deep familiarity with GCP (Cloud Run Functions, GKE) and AWS (EKS) in a multi-cloud or hybrid-cloud context. How you ll benefit: Attractive Discretionary Time Off Private Medical Insurance with optional dependent coverage Educational Assistance Reimbursement Program Opportunities for reimbursement for conferences, trainings, and other personal development events Maternity and Paternity Leave Ask recruiter for country specific information

One address, no account. We’ll tell you when matching roles go live.

More at Vonage

Related open roles

View all roles