Source description
About the role
Own end-to-end delivery of significant data science projects — from problem scoping and approach design through to production deployment
Make sound, independently-reasoned decisions on methodology, model selection, and evaluation; document them clearly in technical solution documents covering problem statement, approach, metrics, and timeline
Lead solution design for your own initiatives; break down complex epics into well-scoped user stories with clear acceptance criteria, adopting DataOps and MLOps best practices throughout — experiment tracking, pipeline orchestration, model monitoring, and reproducibility
Build production-quality Python and PySpark code on Databricks — well-tested, documented, and reusable — and implement advanced ML and AI-powered workflows including entity resolution, probabilistic record linkage, embedding-based matching, semantic similarity, and LLM-augmented pipelines
Develop and maintain reusable tools, libraries, and documentation that improve team efficiency and technical standards; conduct code reviews with constructive, specific feedback that raises the bar
Mentor junior data scientists on technical execution, code quality, and career development; lead internal talks or workshops on ML topics
Collaborate cross-functionally with product, engineering, and operations — translate business requirements into technical specifications, partner with data engineering on scalable pipeline design, and participate in cross-functional design reviews and working groups
More at Samba
Related open roles
Senior Software Engineer, Full Stack
San Francisco Bay Area · Onsite
AI GTM Engineer
San Francisco Bay Area · Onsite
Integration Embedded Engineer
Taipei · Hybrid
Full Stack Engineer
Taipei · Hybrid
Senior Data Scientist
Amsterdam · Hybrid
Data Scientist
San Francisco Bay Area · Onsite
