Padmi

Applied Data Scientist

HyderabadPosted 2 months ago
Data Science And StatisticsSeniorFull Time; Regular
Apply at REALPAGE INC

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior Data Scientist, your role will involve designing, building, and maintaining ML-powered systems to address core data quality and classification challenges across the business. You will be responsible for the entire lifecycle of models, from exploratory analysis and feature engineering to model training, deployment, and continuous performance monitoring. The focus will be on entity resolution and multi-class classification models that impact decision-making in various business domains. Key Responsibilities: - Own the complete model lifecycle, including problem framing, data exploration, feature engineering, model training, evaluation, deployment, and monitoring. - Develop and manage entity resolution systems to identify duplicate records using supervised ML and string similarity techniques. - Create classification models to categorize unstructured or semi-structured data into relevant business categories. - Engineer features from messy text data using string matching algorithms, phonetic encoding, n-grams, and other NLP techniques. - Design candidate retrieval and indexing strategies for scalable model performance. - Fine-tune thresholds, scoring logic, and rule-based overrides to optimize precision and recall for production use cases. - Maintain production model artifacts and data pipelines to ensure models remain relevant as underlying data changes. - Collaborate with engineering and product teams to understand requirements and translate business problems into well-scoped modeling tasks. Qualifications: - 10-15 years of experience in end-to-end ML model building and deployment. - Proficiency in Python, including pandas, NumPy, scikit-learn, XGBoost, or similar frameworks. - Hands-on experience with record linkage, entity resolution, or deduplication challenges. - Expertise in constructing classification models on structured and semi-structured data. - Deep knowledge of string similarity algorithms such as edit distance, sequence matching, and phonetic encoding. - Strong feature engineering skills to extract meaningful signals from noisy data. - Comfortable working with large data structures and understanding memory/performance tradeoffs in production settings. - Familiarity with SQL and relational databases like PostgreSQL. - Effective communication skills to explain model behavior and tradeoffs to non-technical stakeholders. Nice to Have: - Experience with blocking and indexing strategies for scalable record linkage. - Background in NLP, text normalization, or information extraction. - Familiarity with model serving in API contexts like Flask, FastAPI, or similar technologies. - Exposure to deep learning frameworks such as PyTorch or TensorFlow for text classification. As a Senior Data Scientist, your role will involve designing, building, and maintaining ML-powered systems to address core data quality and classification challenges across the business. You will be responsible for the entire lifecycle of models, from exploratory analysis and feature engineering to model training, deployment, and continuous performance monitoring. The focus will be on entity resolution and multi-class classification models that impact decision-making in various business domains. Key Responsibilities: - Own the complete model lifecycle, including problem framing, data exploration, feature engineering, model training, evaluation, deployment, and monitoring. - Develop and manage entity resolution systems to identify duplicate records using supervised ML and string similarity techniques. - Create classification models to categorize unstructured or semi-structured data into relevant business categories. - Engineer features from messy text data using string matching algorithms, phonetic encoding, n-grams, and other NLP techniques. - Design candidate retrieval and indexing strategies for scalable model performance. - Fine-tune thresholds, scoring logic, and rule-based overrides to optimize precision and recall for production use cases. - Maintain production model artifacts and data pipelines to ensure models remain relevant as underlying data changes. - Collaborate with engineering and product teams to understand requirements and translate business problems into well-scoped modeling tasks. Qualifications: - 10-15 years of experience in end-to-end ML model building and deployment. - Proficiency in Python, including pandas, NumPy, scikit-learn, XGBoost, or similar frameworks. - Hands-on experience with record linkage, entity resolution, or deduplication challenges. - Expertise in constructing classification models on structured and semi-structured data. - Deep knowledge of string similarity algorithms such as edit distance, sequence matching, and phonetic encoding. - Strong feature engineering skills to extract meaningful signals from noisy data. - Comfortable working with large data structures and understanding memory/performance tradeoffs in production settings. - Familiarity with SQL and relati

One address, no account. We’ll tell you when matching roles go live.

More at REALPAGE INC

Related open roles

View all roles