Source description
About the role
Key Responsibilities Design, develop, and optimize highly scalable AI inference infrastructure for production environments.
Build, deploy, and maintain AI-powered applications including Retrieval-Augmented Generation (RAG), autonomous agents, and other emerging AI technologies.
Lead technical initiatives across multiple engineering teams and serve as the primary technical authority for AI infrastructure projects.
Architect cloud-native solutions that prioritize scalability, performance, security, and reliability.
Establish best practices, technical standards, governance, and engineering policies for AI platform development.
Implement and manage observability solutions including monitoring, logging, alerting, and performance optimization.
Drive adoption of modern technologies, engineering practices, and automation across the organization.
Navigate complex technical challenges, define architecture for ambiguous requirements, and deliver scalable solutions.
Provide technical mentorship, coaching, and career development for a small engineering team through regular one-on-one meetings.
Serve as the primary point of contact for project coordination and contract-related activities.
Collaborate with cross-functional stakeholders and executive leadership to align technical strategy with business objectives.
Ensure AI platform components maintain high levels of availability, resiliency, security, and operational excellence.
Required Qualifications Active Top Secret/SCI Clearance with Full Scope Polygraph
Bachelor's degree in Computer Science, Engineering, or another technical discipline (or four additional years of equivalent professional experience)
12+ years of professional software engineering experience
Extensive experience designing, building, and operating large-scale production systems
Strong background in full-stack software development and cloud-native architectures
Expert-level experience with AWS
Advanced Kubernetes administration, orchestration, and deployment experience
Strong Python development skills
Experience integrating complex systems across multiple technologies and platforms
Hands-on experience implementing observability platforms using technologies such as: OpenTelemetry
Prometheus
Grafana
Application Performance Monitoring (APM)
Demonstrated success leading technical initiatives and influencing engineering organizations
Experience developing engineering standards, governance, and technical policies
Excellent communication, leadership, stakeholder management, and project coordination skills
Ability to balance strategic leadership with hands-on software engineering responsibilities
Preferred Qualifications
-
Experience with AI inference platforms such as vLLM , LiteLLM , or similar technologies
-
Experience building applications using agentic AI frameworks including LangChain
-
Knowledge of vector databases, embeddings, and semantic search technologies
-
Experience with distributed systems and high-performance computing environments
-
Proven track record driving technical innovation, modernization, and engineering culture improvements
More at W3Global