Padmi

Data Science Specialist

IndiaPosted 3 months ago
Data Science And StatisticsSeniorFull Time; Regular
Apply at REALPAGE INC

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Data Scientist, you will be responsible for designing, building, and maintaining ML-powered systems to address core data quality and classification issues within the business. Your role will involve owning the full lifecycle of models, from exploratory analysis and feature engineering to model training, deployment, and continuous performance monitoring. Key Responsibilities: - Own the end-to-end model lifecycle, including problem framing, data exploration, feature engineering, model training, evaluation, deployment, and monitoring - Build and manage entity resolution systems utilizing supervised ML and string similarity techniques to detect duplicate records - Develop classification models for categorizing unstructured or semi-structured data into meaningful business categories - Engineer features from messy text data using string matching algorithms, phonetic encoding, n-grams, and other NLP techniques - Design candidate retrieval and indexing strategies to enhance model performance at scale - Fine-tune thresholds, scoring logic, and rule-based overrides to balance precision and recall for production use cases - Maintain production model artifacts and data pipelines to ensure models remain up-to-date with evolving data - Collaborate with engineering and product teams to understand requirements and translate business problems into well-scoped modeling tasks Qualifications: - 10-15 years of experience in building and deploying ML models end-to-end - Proficiency in Python, specifically pandas, NumPy, scikit-learn, XGBoost, or similar gradient boosting frameworks - Hands-on experience with record linkage, entity resolution, or deduplication problems - Experience in building classification models (binary and multi-class) on structured and semi-structured data - Deep understanding of string similarity algorithms such as edit distance, sequence matching, and phonetic encoding - Strong feature engineering skills to extract signal from noisy and inconsistently formatted data - Comfort working with large serialized data structures and understanding memory/performance tradeoffs in production environments - Proficiency in SQL and relational databases like PostgreSQL - Excellent communication skills to explain model behavior and tradeoffs to non-technical stakeholders Nice to Have: - Experience with blocking and indexing strategies for scalable record linkage - Background in NLP, text normalization, or information extraction - Familiarity with model serving in API contexts using tools like Flask, FastAPI, or similar - Exposure to deep learning frameworks like PyTorch and TensorFlow for text classification As a Senior Data Scientist, you will be responsible for designing, building, and maintaining ML-powered systems to address core data quality and classification issues within the business. Your role will involve owning the full lifecycle of models, from exploratory analysis and feature engineering to model training, deployment, and continuous performance monitoring. Key Responsibilities: - Own the end-to-end model lifecycle, including problem framing, data exploration, feature engineering, model training, evaluation, deployment, and monitoring - Build and manage entity resolution systems utilizing supervised ML and string similarity techniques to detect duplicate records - Develop classification models for categorizing unstructured or semi-structured data into meaningful business categories - Engineer features from messy text data using string matching algorithms, phonetic encoding, n-grams, and other NLP techniques - Design candidate retrieval and indexing strategies to enhance model performance at scale - Fine-tune thresholds, scoring logic, and rule-based overrides to balance precision and recall for production use cases - Maintain production model artifacts and data pipelines to ensure models remain up-to-date with evolving data - Collaborate with engineering and product teams to understand requirements and translate business problems into well-scoped modeling tasks Qualifications: - 10-15 years of experience in building and deploying ML models end-to-end - Proficiency in Python, specifically pandas, NumPy, scikit-learn, XGBoost, or similar gradient boosting frameworks - Hands-on experience with record linkage, entity resolution, or deduplication problems - Experience in building classification models (binary and multi-class) on structured and semi-structured data - Deep understanding of string similarity algorithms such as edit distance, sequence matching, and phonetic encoding - Strong feature engineering skills to extract signal from noisy and inconsistently formatted data - Comfort working with large serialized data structures and understanding memory/performance tradeoffs in production environments - Proficiency in SQL and relational databases like PostgreSQL - Excellent communication skills to explain model behavior and tradeoffs to non-technical stakeholders Nice to Have: - Experience with

One address, no account. We’ll tell you when matching roles go live.

More at REALPAGE INC

Related open roles

View all roles