Source description
About the role
As a Data Scientist at our company, your role will be to design, build, and maintain ML-powered systems that address core data quality and classification challenges. You will be responsible for the full lifecycle of these systems, from exploratory analysis and feature engineering to model training, deployment, and ongoing performance monitoring. The scope of your work will include entity resolution and multi-class classification models that drive decision-making in various business domains. Key Responsibilities: - Own the end-to-end model lifecycle, including problem framing, data exploration, feature engineering, model training, evaluation, deployment, and monitoring - Build and maintain entity resolution systems using supervised ML and string similarity techniques to detect duplicate records - Develop classification models for categorizing unstructured or semi-structured data into meaningful business categories - Engineer features from messy text data using string matching algorithms, phonetic encoding, n-grams, and other NLP techniques - Design candidate retrieval and indexing strategies to optimize model performance at scale - Tune thresholds, scoring logic, and rule-based overrides to balance precision and recall for production use cases - Maintain production model artifacts and data pipelines to ensure models remain current with evolving data - Collaborate with engineering and product teams to translate business requirements into well-scoped modeling tasks Qualifications: - 10+ years of experience in building and deploying ML models end-to-end - Strong Python skills, including proficiency in pandas, NumPy, scikit-learn, XGBoost, or similar gradient boosting frameworks - Hands-on experience with record linkage, entity resolution, or deduplication problems - Experience building classification models on structured and semi-structured data - Deep familiarity with string similarity algorithms such as edit distance, sequence matching, and phonetic encoding - Strong feature engineering instincts to extract signal from noisy, inconsistently formatted data - Comfort working with large serialized data structures and understanding memory/performance tradeoffs in production contexts - Experience with SQL and relational databases like PostgreSQL - Clear communication skills to explain model behavior and tradeoffs to non-technical stakeholders Nice to Have: - Experience with blocking and indexing strategies for scalable record linkage - Background in NLP, text normalization, or information extraction - Familiarity with model serving in API contexts using Flask, FastAPI, or similar frameworks - Experience in data quality, master data management, or marketplace domains - Exposure to deep learning frameworks like PyTorch or TensorFlow for text classification As a Data Scientist at our company, your role will be to design, build, and maintain ML-powered systems that address core data quality and classification challenges. You will be responsible for the full lifecycle of these systems, from exploratory analysis and feature engineering to model training, deployment, and ongoing performance monitoring. The scope of your work will include entity resolution and multi-class classification models that drive decision-making in various business domains. Key Responsibilities: - Own the end-to-end model lifecycle, including problem framing, data exploration, feature engineering, model training, evaluation, deployment, and monitoring - Build and maintain entity resolution systems using supervised ML and string similarity techniques to detect duplicate records - Develop classification models for categorizing unstructured or semi-structured data into meaningful business categories - Engineer features from messy text data using string matching algorithms, phonetic encoding, n-grams, and other NLP techniques - Design candidate retrieval and indexing strategies to optimize model performance at scale - Tune thresholds, scoring logic, and rule-based overrides to balance precision and recall for production use cases - Maintain production model artifacts and data pipelines to ensure models remain current with evolving data - Collaborate with engineering and product teams to translate business requirements into well-scoped modeling tasks Qualifications: - 10+ years of experience in building and deploying ML models end-to-end - Strong Python skills, including proficiency in pandas, NumPy, scikit-learn, XGBoost, or similar gradient boosting frameworks - Hands-on experience with record linkage, entity resolution, or deduplication problems - Experience building classification models on structured and semi-structured data - Deep familiarity with string similarity algorithms such as edit distance, sequence matching, and phonetic encoding - Strong feature engineering instincts to extract signal from noisy, inconsistently formatted data - Comfort working with large serialized data structures and understanding memory/perfor
More at REALPAGE INC