Padmi
Freya logo
Freya

voice AI agents · on-premise deployment

DevOps Enginer

Remote · Turkey$50k–$70k/yrPosted 7 days ago
InfrastructureMid-levelFull Time
Apply at Freya

Opens the source posting on ycombinator.com

Source description

About the role

View original

You will own the backbone that runs thousands of concurrent AI-powered phone calls inside bank-grade on-premise environments. Not just keeping pods alive, you architect distributed systems that handle real-time voice, scale STT / LLM / TTS inference across customer GPU clusters, integrate with enterprise telephony (Cisco CUBE, Genesys, Asterisk), and deploy behind the firewalls of largest financial institutions. Your work decides whether our platform answers a bank's rush-hour traffic or leaves customers on dead air.

What You'll Do

  • • Own on-prem deployments into OpenShift clusters inside banks. Helm charts, image registries, GPU allocation, CyberArk integration, SAML 2.0 / OIDC SSO.

  • • Scale GPU inference infrastructure for our STT, TTS, and LLM models across multiple customer environments (H100 / H200, NVLink, Triton or vLLM).

  • • Integrate with telephony : Asterisk, SIP trunks, Cisco CUBE, Genesys, WebRTC. SIP header parsing (X-Genesys-*), direction routing, warm transfers, DTMF.

  • • Own reliability : Splunk SIEM forwarding, Langfuse and Grafana observability, incident playbooks for bank-grade 24/7 SLAs.

  • • Security and compliance : RBAC, pentest remediation, KVKK and BDDK compliance patterns, pod security policies.

  • • Scale with growth : we onboard a new bank or insurer every quarter. Each is a new on-prem environment with its own constraints.

  • • Spot flaws early . We are building new architecture for a regulated industry. You help us see what needs to be solved next.

  • Interesting Problems to Own

  • • On-prem meets streaming . Most voice AI stacks assume cloud. We run the same stack inside banks with zero internet egress. Novel problems in image delivery, model updates, secrets rotation.

  • • Bank-scale concurrency . A single campaign can put millions of customers on the line the same afternoon. Queueing, graceful degradation, GPU-aware autoscaling are yours to design.

  • • Legacy-meets-new telephony . Cisco CUBE, Asterisk, Genesys, SIP, WebRTC. You wrangle old-school protocols alongside modern streaming stacks.

  • What Makes You a Great Fit

  • • 3+ years building and scaling distributed systems. Deep Kubernetes / OpenShift knowledge. AWS / GCP helpful.

  • • Fundamentals plus . You can sketch how a SIP INVITE flows through a proxy, explain a K8s GPU scheduler, or tell us the obscure thing you fell asleep reading last night.

  • • Real-time systems experience. Low-latency streaming or inference. Voice / video is a big plus.

  • • On-prem mentality . You have shipped software into environments you did not fully control: Turkish bank, European healthcare, US regulated finance, anything similar.

  • • Startup hats . You have worked where problems find you before process does.

  • • Opinionated without alienating . Opinions drive progress, but you find compromises with customers and teammates.

  • • Familiar with : Kubernetes / OpenShift, Helm, Docker, Terraform, NVIDIA GPU stack, Asterisk, SIP, Cisco CUBE, Genesys, Splunk, Grafana, CyberArk, Python or Go.

Bonus Points

  • • Telephony systems (SIP, VoIP, WebRTC).

  • • ML infrastructure, model serving, or GPU computing (Triton, vLLM, TensorRT).

  • • Real-time audio processing.

  • • Banking / fintech / BDDK and KVKK compliance familiarity.

  • • Fluent Turkish or comfortable working with Turkish customer teams daily.

  • Don't worry about the checklist. We hire for how you think, not how many boxes you tick. The work is hard, the hours

  • are long, and most of what we ship nobody has built before. If that sounds good instead of scary, apply.