Padmi

DevOps / Platform & Reliability Engineer

MumbaiPosted 2 months ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at VotersAI

Opens the source posting on shine.com

Source description

About the role

View original

As a DevOps / platform engineer at Voters AI, a silicon valley based start-up, your role will involve the following responsibilities: - Build and operate CI/CD: You will be responsible for building and operating CI/CD pipelines that deploy server-less services to AWS through GitHub Actions and OIDC, ensuring safe one-at-a-time deploys, verification, and fast rollback. - Operate and harden AWS estate: Your duties will include operating and hardening a production AWS estate consisting of various services such as Lambda, API Gateway, SQS with dead-letter queues, DynamoDB, RDS PostgreSQL, VPC networking, IAM, S3, CloudFront, Route 53, Cognito, Secrets Manager, and SES. - Build out observability: You will need to build out observability by setting up CloudWatch metrics, alarms, dashboards, structured logging, and alerting. You will also create an operations and monitoring surface for the team. - Engineer for resilience and scale: Your responsibilities will involve engineering for resilience and scale by conducting load and chaos testing, implementing Step Functions self-healing and auto-remediation, capacity and concurrency tuning, and setting up queue-based admission control for high-burst workloads. - Own infrastructure security: You will own least-privilege IAM, secret rotation, and infrastructure security, ensuring the overall security of the platform. - Carry on-call and drive incident response: You will be on-call, responsible for driving incident response and working towards reducing mean-time-to-resolve. Qualifications required for this role include: - Minimum 5 years of professional experience in DevOps, SRE, or platform engineering, with hands-on experience in AWS services such as Lambda, API Gateway, SQS, DynamoDB, IAM, VPC, CloudWatch, EventBridge, Step Functions, S3, CloudFront, Cognito, and Secrets Manager. - Experience in infrastructure-as-code tools like Terraform, CloudFormation, SAM, or CDK. - Proficiency in GitHub Actions CI/CD, including OIDC federation to cloud roles. - Strong knowledge in both NoSQL and relational databases, specifically DynamoDB (single-table design, capacity and throughput modeling) and PostgreSQL (RDS operations, tuning, backups). - Experience in managing REST API operations behind API Gateway, including authorizers, throttling, CORS, and staged deployments. - Proficiency in observability and incident response practices, including metrics, alarms, dashboards, alerting, SLOs, and mean-time-to-resolve discipline. - Strong skills in Python and Bash scripting, as well as Linux. Other skills that are preferred but not mandatory include capacity planning, queuing theory, chaos and load testing, high-volume messaging, multi-tenant SaaS operations, and experience with regulated-communications compliance. Please note that the location for this role is in the Aundh Area of Pune, and you should be able to commute to this location. As a DevOps / platform engineer at Voters AI, a silicon valley based start-up, your role will involve the following responsibilities: - Build and operate CI/CD: You will be responsible for building and operating CI/CD pipelines that deploy server-less services to AWS through GitHub Actions and OIDC, ensuring safe one-at-a-time deploys, verification, and fast rollback. - Operate and harden AWS estate: Your duties will include operating and hardening a production AWS estate consisting of various services such as Lambda, API Gateway, SQS with dead-letter queues, DynamoDB, RDS PostgreSQL, VPC networking, IAM, S3, CloudFront, Route 53, Cognito, Secrets Manager, and SES. - Build out observability: You will need to build out observability by setting up CloudWatch metrics, alarms, dashboards, structured logging, and alerting. You will also create an operations and monitoring surface for the team. - Engineer for resilience and scale: Your responsibilities will involve engineering for resilience and scale by conducting load and chaos testing, implementing Step Functions self-healing and auto-remediation, capacity and concurrency tuning, and setting up queue-based admission control for high-burst workloads. - Own infrastructure security: You will own least-privilege IAM, secret rotation, and infrastructure security, ensuring the overall security of the platform. - Carry on-call and drive incident response: You will be on-call, responsible for driving incident response and working towards reducing mean-time-to-resolve. Qualifications required for this role include: - Minimum 5 years of professional experience in DevOps, SRE, or platform engineering, with hands-on experience in AWS services such as Lambda, API Gateway, SQS, DynamoDB, IAM, VPC, CloudWatch, EventBridge, Step Functions, S3, CloudFront, Cognito, and Secrets Manager. - Experience in infrastructure-as-code tools like Terraform, CloudFormation, SAM, or CDK. - Proficiency in GitHub Actions CI/CD, including OIDC federatio

One address, no account. We’ll tell you when matching roles go live.

More at VotersAI

Related open roles

View all roles