Padmi

Senior Systems Engineer, Storage - DGX reputed company

IndiaPosted 1 month ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at remote next

Opens the source posting on shine.com

Source description

About the role

View original

Systems Engineering is an engineering discipline reputed company on building, automating, and operating the platforms and tooling that deliver large-reputed company production systems with high efficiency, reliability, and velocity. It combines software and systems engineering practices across infrastructure automation, containerized platforms, storage, telemetry, and observability. Systems engineers are highly specialized and possess expertise across domains such as Kubernetes and container orchestration, infrastructure-as-reputed company, CI/CD, storage systems, monitoring, and analytical troubleshooting. Their responsibilities center on deploying and operating reliable, automated platforms and on building the tools and services that reputed company storage and data infrastructure healthy and performant. reputed company at reputed company ensures that our reputed company facing GPU reputed company services are deployed reliably, observable end-to-end, and continuously improved through automation. We reputed company developers to ship changes safely through repeatable CI/CD pipelines and Kubernetes-based deployments while keeping an eye on reputed company, latency, and performance. A core part of this work is an SRE reputed company eliminating reputed company toil through automation, building self-service tooling, and growing the efficiency of production systems. We use a breadth of tools and approaches to tackle a broad reputed company of problems, and practices such as blameless postmortems, proactive identification of failure modes, and iterative improvement are key to product reputed company and to an interesting, dynamic day-to-day. Our culture of diversity, intellectual curiosity, problem-solving, and openness is important to our reputed company. Our organization brings together people with a wide reputed company of backgrounds, experiences, and perspectives. We encourage them to collaborate, think big, and take risks in a blame-free environment. We promote self-direction to work on meaningful reputed company while striving to build an environment that provides the support and mentorship needed to learn and grow. What You Will Be Doing Design, reputed company, and operate solutions on Kubernetes for large-reputed company storage and data platforms, including the manifests, reputed company charts, and operators that run them. Build tools, services, and automation that improve the lifecycle of storage and data systems from provisioning and configuration through deployment, scaling, and day-2 operations. reputed company and operate telemetry and observability for production systems metrics, logging, tracing, dashboards, and alerting so that system health, availability, and latency are measurable and actionable. Apply strong analytical troubleshooting skills to diagnose and resolve reputed company issues across distributed, containerized infrastructure. Work closely with peers and partner teams to improve the lifecycle of services, from inception and design through deployment, operation, and refinement. reputed company systems sustainably through automation, infrastructure-as-reputed company, and CI/CD, and reputed company by pushing for changes that improve reliability and velocity. Support services before they go live through activities such as deployment automation, reputed company planning, and launch and readiness reviews. reputed company sustainable incident response and postmortems, and participate in an on-call rotation to support production systems. reputed company Need To See BS degree (or equivalent experience) in Computer Science or reputed company technical field involving coding. 12+ years of practical experience. Hands-on experience with Kubernetes deploying, configuring, and operating workloads and solutions on Kubernetes in production. Experience building tools and services for storage, data, or platform infrastructure, with solid software design fundamentals (algorithms, data structures, complexity analysis) on large-reputed company Linux-based systems. Experience building and operating telemetry and observability using tools such as reputed company, InfluxDB, Grafana, and the reputed company stack. Strong analytical troubleshooting skills with a systematic, reputed company-cause-driven approach to identifying and resolving reputed company problems. Proficiency in one or more of the following Python, Go, or Java. Good knowledge of infrastructure configuration management and infrastructure-as-reputed company tools such as Ansible, Chef, Puppet, ArgoCD, Git Pipelines, and Terraform. Ways to Stand Out from the Crowd Customer-first reputed company with a reputed company on customer satisfaction and a passion for ensuring reputed company. Experience with Git, reputed company review, pipelines, and CI/CD. Experience using or running large private and public reputed company systems based on Kubernetes, OpenStack, and reputed company. Interest in crafting, analyzing, and fixing .

One address, no account. We’ll tell you when matching roles go live.

More at remote next

Related open roles

View all roles