Padmi
Persistent Systems logo
Persistent Systems

Wave Relay MANET · mobile ad-hoc networking

Data Scientist Synthetic Data

IndiaPosted 3 months ago
Data Science And StatisticsSeniorFull Time; Regular
Apply at Persistent Systems

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: As a Data Scientist specialising in Generative AI and Synthetic Data Generation, your primary responsibility will be to design, develop, and deploy advanced data-driven solutions. You will focus on building and optimizing generative models capable of producing high-quality synthetic data that closely mimics real-world datasets across domains such as text, images, and structured data. Your role is crucial in leveraging machine learning and statistical techniques to enhance model performance, scalability, and reliability. You will collaborate closely with AI researchers, engineers, and domain experts to drive innovation in generative AI systems, particularly in data-constrained or regulated environments. Key Responsibilities: - Design and develop machine learning models for synthetic data generation using techniques such as GANs, VAEs, diffusion models, and other deep generative approaches. Ensure generated data maintains statistical fidelity, diversity, and privacy compliance. - Identify, acquire, and curate relevant datasets. Perform data cleansing, transformation, and structuring to ensure high-quality inputs for training generative models. - Build end-to-end data pipelines and workflows for training generative AI models using state-of-the-art architectures including GANs, Variational Autoencoders, Normalizing Flows, and Diffusion Networks. - Perform hyperparameter tuning and experiment with architectures to improve model accuracy, stability, and output quality. - Implement advanced data augmentation strategies to enhance dataset size, diversity, and model generalization. - Define, track, and improve model evaluation metrics to ensure objective assessment of generative model performance and synthetic data quality. - Analyse datasets and model outputs to identify biases and implement techniques to ensure fairness, robustness, and ethical AI practices. - Apply transfer learning approaches to fine-tune pre-trained models for new domains and specific use cases. - Work cross-functionally with AI researchers, software engineers, product teams, and domain SMEs to integrate generative AI and synthetic data solutions into production systems. - Maintain clear and comprehensive documentation of methodologies, experiments, and findings to ensure reproducibility and knowledge dissemination. Qualifications Required: - Masters or Ph.D. in Computer Science, Data Science, Machine Learning, or a related discipline with a focus on AI or generative modelling. - Mandatory experience in synthetic data generation using machine learning or deep learning models. - Strong proficiency in Python and leading AI frameworks such as TensorFlow or PyTorch. - In-depth understanding of generative modelling techniques including GANs, VAEs, Diffusion Models, and Normalizing Flows. - 10+ years of experience in data science, including large-scale data processing, feature engineering, and model training. - Strong foundation in statistics, probability, and quantitative analysis. Additional Details: Omit this section as there are no additional details of the company present in the job description. Role Overview: As a Data Scientist specialising in Generative AI and Synthetic Data Generation, your primary responsibility will be to design, develop, and deploy advanced data-driven solutions. You will focus on building and optimizing generative models capable of producing high-quality synthetic data that closely mimics real-world datasets across domains such as text, images, and structured data. Your role is crucial in leveraging machine learning and statistical techniques to enhance model performance, scalability, and reliability. You will collaborate closely with AI researchers, engineers, and domain experts to drive innovation in generative AI systems, particularly in data-constrained or regulated environments. Key Responsibilities: - Design and develop machine learning models for synthetic data generation using techniques such as GANs, VAEs, diffusion models, and other deep generative approaches. Ensure generated data maintains statistical fidelity, diversity, and privacy compliance. - Identify, acquire, and curate relevant datasets. Perform data cleansing, transformation, and structuring to ensure high-quality inputs for training generative models. - Build end-to-end data pipelines and workflows for training generative AI models using state-of-the-art architectures including GANs, Variational Autoencoders, Normalizing Flows, and Diffusion Networks. - Perform hyperparameter tuning and experiment with architectures to improve model accuracy, stability, and output quality. - Implement advanced data augmentation strategies to enhance dataset size, diversity, and model generalization. - Define, track, and improve model evaluation metrics to ensure objective assessment of generative model performance and synthetic data quality. - Analyse datasets and model outputs to identify biases and implement techniques to ensure fairness, robus

One address, no account. We’ll tell you when matching roles go live.

More at Persistent Systems

Related open roles

View all roles