Source description
About the role
What You Will Do
ML & AI Systems
Design, train, and own the full lifecycle of ML models for payment optimization — routing decisions, authorization rate improvement, cost reduction, and fraud signals — using PyTorch, TensorFlow, or XGBoost.
Build and operate LLM-powered workflows: LangGraph agent orchestration, RAG pipelines, and vector DB integrations (Pinecone, pgvector, or Weaviate).
Own the MLOps stack end-to-end: experiment tracking (MLflow / W&B), model registry, feature store, and automated retraining pipelines on AWS SageMaker.
Monitor model health continuously — drift, distribution shifts, retraining triggers — and define evaluation metrics tied directly to business outcomes.
Platform Engineering & Payments Integration
Build and maintain inference services in Go and Python integrated into live payment routing — strict latency SLAs (<100 ms), zero silent errors.
Own AWS infrastructure: ECS/EKS, Terraform IaC, SQS/SNS event streaming, RDS/Aurora, and S3 for model artifacts.
Design and ship on-premise and hybrid deployment architectures for enterprise clients requiring local data residency, including secure data sync pipelines.
Apply PCI-DSS standards across all components touching payment data; implement tokenization in ML pipelines; design for PSP-specific behavior (Cybersource, Worldpay, Prosa, Cielo, Pagbank, and others).
Build and maintain RESTful and gRPC APIs that expose AI platform capabilities to merchants and partners.
Technical Leadership
Own observability end-to-end: Prometheus/Grafana dashboards, OpenTelemetry tracing, model-specific monitors, and on-call runbooks.
Set the engineering bar for the team: architecture reviews, code standards, testing strategy (unit, integration, shadow mode), and CI/CD practices.
Mentor engineers, run design reviews, and translate product vision into executable technical roadmaps with clear timelines and trade-offs.
Technical Skills
Backend / Platform
Go (production services)
Python (ML + tooling)
gRPC & REST APIs
Event streaming (SQS/SNS)
Distributed systems
Cloud & Infra — AWS
ECS / EKS
Terraform / IaC
SageMaker or Vertex AI
RDS/Aurora, S3
Hybrid / on-prem deploy
AI / ML Stack
PyTorch or TensorFlow
XGBoost / scikit-learn
MLflow / W&B
Feature stores
Model monitoring & drift
LLMs & Agents
LangGraph / LangChain
RAG + vector DBs
Prompt engineering
LLM evaluation
Structured outputs
Payments Domain
PCI-DSS compliance
Tokenization patterns
PSP integrations
Auth rate optimization
Routing orchestration
Frontend
React / Next.js
TypeScript
Component systems
API integration
Observability
Prometheus / Grafana
OpenTelemetry
Structured logging
On-call runbooks
Data
SQL (analytical)
Airflow / dbt
Feature pipelines
Data quality & lineage
What We Are Looking For
-
8+ years in software engineering; 3+ at Staff, Principal, or Tech Lead level owning a production platform end-to-end.
-
Proven track record shipping ML/AI systems to production: training, serving, monitoring, and retraining — not just prototyping.
-
Hands-on LLM experience in production: agents, RAG pipelines, or AI workflow orchestration.
-
Payments or fintech background with practical knowledge of PSP behavior, PCI-DSS scope, authorization logic, and routing trade-offs.
-
Experience designing and deploying on-premise or hybrid enterprise infrastructure.
-
Bachelor's degree in Computer Science, Engineering, or equivalent demonstrated depth.
More at DEUNA
