Padmi

SRE; Blockchain, Web3 & Ai Native

IndiaPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at InfraSingularity

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Site Reliability Engineer (SRE) at Infra Singularity, your role will be critical in taking ownership of the multi-cloud blockchain infrastructure and validator node operations. Your responsibilities will include: - Owning and operating validator nodes across multiple blockchain networks to ensure uptime, security, and cost-efficiency. - Architecting, deploying, and maintaining infrastructure on AWS, GCP, and bare-metal for protocol scalability and performance. - Implementing Kubernetes-native tooling such as Helm, FluxCD, Prometheus, and Thanos to manage deployments and observability. - Collaborating with the Protocol R&D team to onboard new blockchains and participate in testnets, mainnets, and governance. - Ensuring secure infrastructure with best-in-class secrets management using tools like Hashi Corp Vault, KMS, and incident response protocols. - Contributing to a robust monitoring and alerting stack to detect anomalies, performance drops, or protocol-level issues. - Acting as a bridge between software, protocol, and product teams to communicate infra constraints or deployment risks clearly. - Continuously improving deployment pipelines using Terraform, Terragrunt, and Git Ops practices. - Participating in on-call rotations and incident retrospectives, driving post-mortem analysis and long-term fixes. Our tech stack includes: - Cloud & Infra: AWS, GCP, bare-metal - Containerization: Kubernetes, Helm, FluxCD - IaC: Terraform, Terragrunt - Monitoring: Prometheus, Thanos, Grafana, Loki - Secrets & Security: Hashi Corp Vault, AWS KMS - Languages: Go, Bash, Python, Types - Blockchain: Ethereum, Polygon, Cosmos, Solana, Foundry, Open Zeppelin Qualifications required: - 4+ years of experience in SRE/Dev Ops/Infra roles, ideally within Fin Tech, Cloud, or high-reliability environments. - Proven expertise managing Kubernetes in production at scale. - Strong hands-on experience with Terraform, Helm, Git Ops workflows. - Deep understanding of system reliability, incident management, fault tolerance, and monitoring best practices. - Proficiency with Prometheus and PromQL for custom dashboards, metrics, and alerting. - Experience operating secure infrastructure and implementing SOC2/ISO 27001-aligned practices. - Solid scripting skills in Bash, Python, or Go. - Clear and confident communicator capable of interfacing with both technical and non-technical stakeholders. Nice-to-have qualifications: - First-hand experience in Web3/blockchain/crypto environments. - Understanding of staking, validator economics, slashing conditions, or L1/L2 governance mechanisms. - Exposure to smart contract deployments or working with Solidity, Foundry, or similar tool chains. - Experience with compliance-heavy or security-certified environments (SOC2, ISO 27001, HIPAA). Join Infra Singularity to work at the bleeding edge of Web3 infrastructure and validator tech, collaborate with a fast-moving team that values ownership, performance, and reliability, and get exposure to some of the most interesting blockchain ecosystems in the world. As a Senior Site Reliability Engineer (SRE) at Infra Singularity, your role will be critical in taking ownership of the multi-cloud blockchain infrastructure and validator node operations. Your responsibilities will include: - Owning and operating validator nodes across multiple blockchain networks to ensure uptime, security, and cost-efficiency. - Architecting, deploying, and maintaining infrastructure on AWS, GCP, and bare-metal for protocol scalability and performance. - Implementing Kubernetes-native tooling such as Helm, FluxCD, Prometheus, and Thanos to manage deployments and observability. - Collaborating with the Protocol R&D team to onboard new blockchains and participate in testnets, mainnets, and governance. - Ensuring secure infrastructure with best-in-class secrets management using tools like Hashi Corp Vault, KMS, and incident response protocols. - Contributing to a robust monitoring and alerting stack to detect anomalies, performance drops, or protocol-level issues. - Acting as a bridge between software, protocol, and product teams to communicate infra constraints or deployment risks clearly. - Continuously improving deployment pipelines using Terraform, Terragrunt, and Git Ops practices. - Participating in on-call rotations and incident retrospectives, driving post-mortem analysis and long-term fixes. Our tech stack includes: - Cloud & Infra: AWS, GCP, bare-metal - Containerization: Kubernetes, Helm, FluxCD - IaC: Terraform, Terragrunt - Monitoring: Prometheus, Thanos, Grafana, Loki - Secrets & Security: Hashi Corp Vault, AWS KMS - Languages: Go, Bash, Python, Types - Blockchain: Ethereum, Polygon, Cosmos, Solana, Foundry, Open Zeppelin Qualifications required: - 4+ years of experience in SRE/Dev Ops/Infra roles, ideally within Fin Tech, Cloud, or high-reliability environments. - Proven expertise managing Kubernetes in production at scale. - Strong hands-on experience wi

One address, no account. We’ll tell you when matching roles go live.

More at InfraSingularity

Related open roles

View all roles