Padmi
Zingtree logo
Zingtree

AI agents · customer support automation

Senior DevOps / Platform Reliability Engineer

Remote · United StatesPosted 2 months ago
InfrastructureSeniorFull Time Remote
Apply at Zingtree

Opens the source posting on jobs.lever.co

Source description

About the role

View original

Own and evolve CI/CD pipelines using GitHub Actions and OIDC-based authentication for microservices and agentic workloads, with safe, fast, and reversible deployments.

Automate infrastructure provisioning using Infrastructure as Code (IaC) tools such as Terraform and CloudFormation.

Operate and scale our Kubernetes platform (EKS + Argo CD), including autoscaling, ingress, external-dns, cert-manager, External Secrets Operator, backups, runtime guardrails, and multi-tenant isolation for enterprise customers.

Manage the edge and network perimeter, including Cloudflare (CDN, WAF, Bot Management, DDoS protection, Zero Trust / Access), CloudFront, API Gateway, ALB/NLB, Route 53, and network security controls.

Operate the data and event tier, including Aurora MySQL, ElastiCache/Redis, S3, and MSK (Kafka), with responsibility for backups, point-in-time recovery (PITR), and multi-AZ disaster recovery aligned to defined RTO/RPO objectives.

Build and maintain Lambda workloads where event-driven or serverless architectures are the right fit.

Build observability as a product using Prometheus, Grafana, and OpenTelemetry, including telemetry for LLM and agentic systems such as token cost, tool-call latency, evaluation signals, and prompt/version tracking.

Strengthen our security and compliance posture for SOC 2 and HIPAA, including least-privilege IAM, SCPs, secrets management, SAST/DAST, dependency and container scanning, image signing, AWS Config, Security Hub, GuardDuty, Inspector, and evidence automation.

Drive FinOps initiatives, including tagging standards, Savings Plans and Reserved Instances, per-tenant and per-workload cost attribution, and LLM cost controls.

Build and evolve our AI-native DevOps capabilities (see section below).

Partner with engineering teams to define platform standards, service templates, deployment best practices, and operational SLOs.

Monitor system performance and ensure reliability, scalability, and security across infrastructure and services.

Collaborate with software engineering teams to support continuous integration and continuous delivery best practices.

Document infrastructure, deployment processes, and operational standards to support knowledge sharing across the team.

More at Zingtree

Related open roles

View all roles