Source description
About the role
About the Role:Grade Level (for internal use): 03Who We Are Kensho is S&P Globals hub for AI innovation and transformation. With expertise in Machine Learning and data discovery, we develop and deploy novel solutions for S&P Global and its customers worldwide. Our solutions help businesses harness the power of data and Artificial Intelligence to innovate and drive progress. Kensho's solutions and research focus on Generative AI, LLM Agents, speech recognition, entity linking, document extraction, text classification, natural language processing, and more. At Kensho, we hire talented people and give them the autonomy and support needed to build amazing technology and products. We collaborate using our teammates' diverse perspectives to solve hard problems. Our communication with one another is open, honest, and efficient. We dedicate time and resources to explore new ideas, but always rooted in engineering best practices. As a result, we can innovate rapidly to produce technology that is scalable, robust, and useful. About the Role As a Senior Site Reliability Engineer (SRE) at Kensho, you will be a hands-on technologist who combines strong infrastructure expertise with solid software engineering skills Python first. You will be responsible for ensuring the reliability, scalability, and security of both business-critical internal systems and external, customer facing services. You will work closely with Infrastructure, Application, and Security teams to design resilient systems, automate operations, and continuously improve platform stability. This role requires deep ownership of production systems, strong troubleshooting skills across infrastructure, Container orchestration systems, networking, and applications, and comfort operating in a 24/7 on call environment. What Youll Do Own and operate production services supporting critical financial applications with a strong focus on availability, performance, and reliabilityDesign, build, and manage AWS infrastructure, including EKS-based clusters, across lower and production environmentsProvision and manage infrastructure using Terraform (Infrastructure as Code) with a strong automation first mindsetDeploy, scale, and troubleshoot applications running on like Kubernetes, including cluster creation, upgrades, and lifecycle managementBuild and maintain automation frameworks and tooling Python based to reduce operational toil and prevent recurring incidentsMonitor system health using metrics, logs, and alerts; continuously tune alerts, dashboards, and runbooksTroubleshoot complex issues spanning clusters, networking, certificates, deployments, and application behaviorManage certificate lifecycle and expiration, ensuring secure and uninterrupted service operationCollaborate with InfoSec, Vulnerability Management, and Network Security teams (e.g., Zscaler) to maintain a strong security postureCollaborate with L1/L2 teams, helping them understand infrastructure and operational best practicesParticipate in on call and lead incident response, drive root cause analysis, and ensure effective post-incident remediation and learningsIdentify architectural anti-patterns and drive improvements by reviewing new services for production readiness, resiliency, and secure design prior to releaseEstablish and enforce production readiness standards, including deployment strategies, rollback plans, and observability requirementsOptimize infrastructure cost and resource utilization without compromising reliability and performanceWhat We Look For 6+ years of experience in SRE, DevOps, Platform, or Infrastructure Engineering rolesStrong software engineering background, with hands-on Python development used for automation, tooling, and system reliabilityExperience building or supporting scalable, distributed systems in productionDeep experience with AWS cloud environments, including AWS, IAM, networking, and access controlsStrong hands-on expertise with similar tools like Kubernetes (EKS preferred): cluster creation, deployments, scaling, and troubleshootingSolid understanding of networking fundamentals (VPCs, routing, DNS, load balancing, security groups)Experience with CI/CD pipelines, deployment tools, and infrastructure automationWorking knowledge of databases and query optimization, and understanding how applications behave under loadFamiliarity with similar tools like Kafka or other messaging systemsComfortable conducting code .
More at S&P Global