Source description
About the role
ML Architecture at Scale: Demonstrated experience designing and shipping ML systems that serve multiple products or teams—not just models, but the platforms, contracts, and abstractions that make ML reusable and reliable at scale.
Technical Leadership Without Authority: Proven ability to drive technical decisions across teams you don't manage—through clear writing, credibility, and the ability to synthesize competing perspectives into a coherent path forward.
Deep Classical & Applied ML Mastery: Expert-level command of classical ML (XGBoost, LightGBM, calibration, cost-sensitive learning) with the judgment to know when—and when not—to reach for more complex approaches. You've operated beyond standard accuracy metrics and can design evaluation frameworks appropriate to the problem.
Production ML Engineering: Extensive experience taking models from experimentation to high-throughput, low-latency production environments. You've owned reliability, SLAs, and incident response for ML systems, and you've built MLOps tooling—not just consumed it.
Software Engineering Excellence: You write and review code at a senior+ level in Python, hold the team to high standards for testability and maintainability, and can credibly engage in systems design discussions with Principal and Staff engineers across Data and Platform.
Strategic Thinking & Business Acumen: You connect technical decisions to business outcomes—approval rate improvements to revenue, latency reductions to conversion, model drift to operational risk. You communicate clearly with non-technical stakeholders and can translate ambiguous business goals into concrete ML problems.
Payments Domain Depth: Strong understanding of the card payment lifecycle, issuer behavior, authorization codes, retry logic, network rules, and 3DS. You use domain knowledge to inform feature design, model architecture, and experimentation strategy—not just as background context.
Cloud Infrastructure Mastery: Deep experience designing and owning ML infrastructure on AWS or GCP at scale, including infrastructure-as-code, cost management, and the ability to make build-vs-buy decisions on platform components.
More at PPRO
