Padmi
Synchrony logo
Synchrony

credit cards · point-of-sale financing

Vice President - AI Reliability and Performance Architect

HyderabadPosted 3 months ago
Software engineeringStaff+Full Time; Regular
Apply at Synchrony

Opens the source posting on shine.com

Source description

About the role

View original

As the VP, AI Reliability & Performance Architect at Synchrony, you will play a crucial role in ensuring the production-grade reliability, accuracy, and performance of the AWS-based agentic AI ecosystem. Your responsibilities will include: - Lead investigations of complex agent/AI workflow failures using logs, metrics, and traces (CloudWatch, X-Ray, Splunk, New Relic or similar) and drive preventive actions through blameless post-mortems. - Improve the quality and performance of Retrieval-Augmented Generation (RAG) and agent workflows by tuning retrieval, ranking/re-ranking, prompt/tooling behavior, and data access patterns across stores such as PostgreSQL/Redshift. - Establish evaluation approaches for models, RAG, and agents to enhance fidelity and reduce regressions. - Collaborate with InfoSec/AppSec to review architectures and ensure designs follow enterprise security patterns, identity controls, and data residency requirements. - Implement and monitor guardrails and controls across the AI platform in collaboration with Governance teams. - Drive Design for Reliability patterns across both Platform and Agent Building teamsfault tolerance, graceful degradation, load/performance testing, incident readiness, and operational excellence. - Translate reliability risks, performance trends, and operational metrics into clear business language for senior leaders, risk, and product owners. - Coach DevLeads and architects on debugging agent behaviors, strengthening observability pipelines, improving orchestration, and hardening production deployments. Qualifications required for this role include: - Bachelor's degree in Computer Science, Engineering, Information Systems, or related field - 1014 years of IT experience including meaningful roles in application development, platform engineering, SRE/operations, and/or architecture or in lieu of a degree 1216 years of IT experience - Strong experience operating and improving reliability of cloud-native systems (AWS preferred) - Strong ability to script/build tooling in Python (or similar language) for reliability automation, analysis, testing, and operational workflows - Hands-on experience with observability practices and tools (CloudWatch/X-Ray/Splunk/New Relic or similar) - Working knowledge of identity and security patterns (OAuth2, SSO/federation, IAM roles/policies/SCP concepts) - Proven ability to lead through influence, drive standards/guardrails, and align multiple agile teams in a matrixed environment Desired skills and knowledge include experience with AWS Bedrock/AgentCore, agent frameworks, automated evaluation pipelines for LLM-based agents, LLM gateways, fintech/regulatory exposure, TOGAF/Zachman frameworks, and agent protocol standards. Join Synchrony's dynamic Engineering Team to contribute to cutting-edge tech solutions and shape the future of technology. Work in a collaborative environment that fosters creativity and career growth. As the VP, AI Reliability & Performance Architect at Synchrony, you will play a crucial role in ensuring the production-grade reliability, accuracy, and performance of the AWS-based agentic AI ecosystem. Your responsibilities will include: - Lead investigations of complex agent/AI workflow failures using logs, metrics, and traces (CloudWatch, X-Ray, Splunk, New Relic or similar) and drive preventive actions through blameless post-mortems. - Improve the quality and performance of Retrieval-Augmented Generation (RAG) and agent workflows by tuning retrieval, ranking/re-ranking, prompt/tooling behavior, and data access patterns across stores such as PostgreSQL/Redshift. - Establish evaluation approaches for models, RAG, and agents to enhance fidelity and reduce regressions. - Collaborate with InfoSec/AppSec to review architectures and ensure designs follow enterprise security patterns, identity controls, and data residency requirements. - Implement and monitor guardrails and controls across the AI platform in collaboration with Governance teams. - Drive Design for Reliability patterns across both Platform and Agent Building teamsfault tolerance, graceful degradation, load/performance testing, incident readiness, and operational excellence. - Translate reliability risks, performance trends, and operational metrics into clear business language for senior leaders, risk, and product owners. - Coach DevLeads and architects on debugging agent behaviors, strengthening observability pipelines, improving orchestration, and hardening production deployments. Qualifications required for this role include: - Bachelor's degree in Computer Science, Engineering, Information Systems, or related field - 1014 years of IT experience including meaningful roles in application development, platform engineering, SRE/operations, and/or architecture or in lieu of a degree 1216 years of IT experience - Strong experience operating and improving reliability of cloud-native systems (AWS preferred) - Strong ability to script/build tooling

One address, no account. We’ll tell you when matching roles go live.

More at Synchrony

Related open roles

View all roles