Source description
About the role
This software engineer-infrastructure contributes to the design, implementation, and operation of foundational infrastructure systems that power the company’s technology platform. This role supports compute, storage, networking, and container infrastructure used by enterprise applications, internal platforms, and hybrid cloud environments.
Software engineers at this level focus on building, operating, and improving infrastructure platforms using established patterns, automation, and infrastructure-as-code. They work collaboratively with platform and operations teams while continuing to build deep technical expertise in distributed systems and modern infrastructure practices. Key Responsibilities Infrastructure Engineering • Support the design, deployment, and operation of infrastructure platforms including compute, storage, networking, and container infrastructure • Build and maintain reliable infrastructure across on-premises data centers and cloud environments • Operate and support Kubernetes clusters and their underlying infrastructure components • Assist in ensuring availability, performance, and stability of infrastructure systems • Support hybrid infrastructure environments and platform services that run on top of them
Automation & Infrastructure as Code • Develop and maintain infrastructure automation using Go, Python, or Java • Implement infrastructure provisioning and configuration using infrastructure-as-code tools such as Terraform • Contribute to standardized infrastructure deployment and lifecycle management practices • Build tooling that reduces manual effort and improves operational reliability
Platform Integration • Support infrastructure dependencies for container platforms and distributed systems • Assist with deploying, upgrading, and maintaining Kubernetes clusters • Operate infrastructure services such as virtualization platforms and storage systems • Collaborate with platform engineering teams supporting CI/CD, messaging, observability, and developer platforms
Observability & Reliability • Implement monitoring and observability using Prometheus, Grafana, and OpenTelemetry • Participate in incident response and post-incident analysis • Contribute to reliability improvements and operational maturity
Security & Access Management • Apply infrastructure security best practices • Support identity, access management, and secrets management systems • Collaborate with security teams to ensure infrastructure resilience and compliance
More at Berkshire Hathaway Energy