Source description
About the role
Own end-to-end delivery of significant data science projects — from problem scoping and approach design through to production deployment
Make sound, independently-reasoned decisions on methodology, model selection, and evaluation; document them clearly in technical solution documents covering problem statement, approach, metrics, and timeline
Lead solution design for your own initiatives; break down complex epics into well-scoped user stories with clear acceptance criteria, adopting DataOps and MLOps best practices throughout — experiment tracking, pipeline orchestration, model monitoring, and reproducibility
Build production-quality Python and PySpark code on Databricks — well-tested, documented, and reusable — and implement advanced ML and AI-powered workflows including entity resolution, probabilistic record linkage, embedding-based matching, semantic similarity, and LLM-augmented pipelines
Develop and maintain reusable tools, libraries, and documentation that improve team efficiency and technical standards; conduct code reviews with constructive, specific feedback that raises the bar
Mentor junior data scientists on technical execution, code quality, and career development; lead internal talks or workshops on ML topics
Collaborate cross-functionally with product, engineering, and operations — translate business requirements into technical specifications, partner with data engineering on scalable pipeline design, and participate in cross-functional design reviews and working groups
More at Samba
Related open roles
Senior Software Engineer, Full Stack
United States · Onsite
Senior Data Scientist
Amsterdam · Hybrid
Data Scientist
San Francisco Bay Area · Onsite
Senior Data Engineer
Poland · Onsite
Ontology Engineer-Knowledge Graph & Identity
United States · Onsite
Senior Ontologist - Knowledge Graph & Identity
San Francisco Bay Area · Onsite
