Source description
About the role
As a Data Engineer at AiLogic Neural Network Pvt Ltd, you will play a crucial role in designing, developing, and maintaining scalable data pipelines for processing large volumes of structured and unstructured data. Your responsibilities will include: - Designing and developing scalable data pipelines for processing structured and unstructured data - Building document ingestion and processing workflows for various text sources like PDFs, scanned documents, and HTML pages - Implementing OCR, PDF parsing, HTML parsing, and text extraction pipelines - Developing document chunking and preprocessing frameworks for NLP and LLM-based applications - Working with Hugging Face models and NLP libraries for text processing tasks - Creating and optimizing data transformation workflows using Python, Apache Spark, and Spark SQL - Developing and managing Vector Database pipelines for embedding storage and retrieval - Implementing text normalization, sentence segmentation, deduplication, and data quality processes - Designing and implementing data masking, classification, and categorization solutions - Collaborating with AI/ML engineers to prepare datasets for model training and inference - Optimizing large-scale data processing workflows for performance, scalability, and cost efficiency - Maintaining CI/CD pipelines and following software engineering best practices - Monitoring, troubleshooting, and improving production data processing systems Mandatory Skills: - Strong experience with Python programming - Hands-on experience in NLP concepts such as Tokenization, Text Processing, and Hugging Face Transformers - Experience in PDF Parsing, OCR, HTML Parsing, Text Extraction, and Document Chunking - Experience with Apache Spark and Spark SQL - Working knowledge of Vector Databases - Good understanding of Git and CI/CD practices - Experience building data pipelines and ETL workflows - Strong debugging and problem-solving skills Preferred Skills: - Text Normalization - Sentence Segmentation - Exact Deduplication and Near Deduplication - Data Masking - Data Classification & Categorization - Embedding Generation and Retrieval Pipelines - Large-scale Document Processing Systems - RAG (Retrieval-Augmented Generation) Pipelines Performance Optimization Skills: - CPU Distribution and Parallel Processing - Pre-batch Generation Techniques - Chunking Optimization Strategies - Stream Processing vs Batch Processing - GPU and CPU Parallel Distribution - CUDA Optimization - PyTorch Performance Tuning - Spark Performance Optimization As a qualified candidate, you should possess: - Bachelor's or Master's degree in Computer Science, Data Science, Information Technology, or a related field - Minimum of 2 years of experience in Data Engineering, NLP Engineering, or AI Data Processing, working with large-scale datasets and distributed computing frameworks If you meet these qualifications and are looking for a challenging opportunity in a dynamic AI-driven product company, we encourage you to apply for this position. As a Data Engineer at AiLogic Neural Network Pvt Ltd, you will play a crucial role in designing, developing, and maintaining scalable data pipelines for processing large volumes of structured and unstructured data. Your responsibilities will include: - Designing and developing scalable data pipelines for processing structured and unstructured data - Building document ingestion and processing workflows for various text sources like PDFs, scanned documents, and HTML pages - Implementing OCR, PDF parsing, HTML parsing, and text extraction pipelines - Developing document chunking and preprocessing frameworks for NLP and LLM-based applications - Working with Hugging Face models and NLP libraries for text processing tasks - Creating and optimizing data transformation workflows using Python, Apache Spark, and Spark SQL - Developing and managing Vector Database pipelines for embedding storage and retrieval - Implementing text normalization, sentence segmentation, deduplication, and data quality processes - Designing and implementing data masking, classification, and categorization solutions - Collaborating with AI/ML engineers to prepare datasets for model training and inference - Optimizing large-scale data processing workflows for performance, scalability, and cost efficiency - Maintaining CI/CD pipelines and following software engineering best practices - Monitoring, troubleshooting, and improving production data processing systems Mandatory Skills: - Strong experience with Python programming - Hands-on experience in NLP concepts such as Tokenization, Text Processing, and Hugging Face Transformers - Experience in PDF Parsing, OCR, HTML Parsing, Text Extraction, and Document Chunking - Experience with Apache Spark and Spark SQL - Working knowledge of Vector Databases - Good understanding of Git and CI/CD practices - Experience building data pipelines and ETL workflows - Strong debugging and problem-solving skills **P