Source description
About the role
Total AI Systems, Inc. is looking for a serious DevOps Engineer—someone who understands infrastructure not as diagrams, but as systems that fail, recover, scale, and ship. We run on Amazon Web Services (AWS) and are evolving from single-tenant deployments toward multi-tenant architectures. You will own how code moves from GitHub CI/CD production, how environments are secured, monitored, scaled, and how incidents are handled when things go sideways. This is a hands-on, operator role, not a theoretical or mostly meetings position. If you enjoy building pipelines, hardening infrastructure, automating everything, and being the person engineering trusts at 3:00 AM—this role is for you. What You'll Be Responsible For Designing, building, and maintaining CI/CD pipelines from GitHub to AWS Managing AWS infrastructure using industry best practices (IaC, automation, security-first design) Supporting and improving single-tenant deployments while helping evolve toward multi-tenant architecture Ensuring high availability, scalability, security, and observability across environments Monitoring systems, responding to incidents, and performing root-cause analysis Implementing logging, alerting, backups, and disaster recovery strategies Partnering with engineering to enable fast, safe, repeatable deployments Owning production reliability—you build it, you keep it running Required Qualifications Bachelor's degree from an accredited college or university 3+ years of hands-on DevOps experience managing production systems Strong experience with Amazon Web Services (AWS) Experience building and maintaining CI/CD pipelines (GitHub-based workflows required) Experience working with single-tenant architectures, with exposure to multi-tenant systems Ability and willingness to work US business hours Fluent spoken and written English Extremely dependable—when issues arise (day or night), you respond and take ownership Preferred (Strongly) AWS Certification (Solutions Architect or higher preferred) Infrastructure-as-Code experience (Terraform, CloudFormation, etc.) Deep familiarity with: VPCs, IAM, EC2, ECS/EKS, RDS, S3, CloudWatch Networking, security groups, least-privilege access Zero-downtime deployments and rollback strategies Experience supporting SaaS platforms at scale Comfort being on-call and owning production health What We Mean by Nerdy We're looking for someone who: Knows why things break, not just that they broke Understands tradeoffs between cost, reliability, and speed Automates first and documents second Has strong opinions—but can back them up Treats uptime, security, and clean pipelines as personal pride