Padmi

Senior Generative AI Engineer - RAG, LLMs and MLOps | WFH | UK Client

Delhi NCRPosted 1 month ago
Software engineeringSeniorFull Time; Regular
Apply at itForte Staffing Services

Opens the source posting on shine.com

Source description

About the role

View original

Key Responsibilities - Design and develop enterprise Generative AI and LLM-based applications. - Build RAG pipelines connecting LLMs with private company data, documents, databases and knowledge repositories. - Develop document ingestion, cleaning, chunking, embedding, indexing and retrieval workflows. - Implement semantic search, hybrid search, metadata filtering, reranking and source attribution. - Design and manage vector databases such as Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS or pgvector. - Select and evaluate suitable embedding models, vector indexes and similarity-search techniques. - Fine-tune open-source models such as Llama, Mistral, Gemma or Qwen using domain-specific data. - Apply techniques such as LoRA, QLoRA, PEFT, quantisation and prompt tuning. - Develop scalable APIs and backend services using Python, FastAPI or similar frameworks. - Integrate LLMs with enterprise applications, APIs, databases and business workflows. - Deploy AI services using cloud platforms, Docker and Kubernetes. - Implement CI/CD pipelines for AI applications and model deployments. - Optimise AI systems for latency, throughput, reliability and inference cost. - Implement caching, asynchronous processing, rate limiting, load balancing and autoscaling. - Monitor model performance, token consumption, API errors, infrastructure utilisation and cloud costs. - Build evaluation frameworks to measure accuracy, relevance, groundedness and retrieval quality. - Identify and reduce hallucinations, irrelevant retrieval and inconsistent responses. - Implement AI security controls to prevent prompt injection, data leakage and unauthorised access. - Prepare technical documentation and participate in architecture and code reviews. - Collaborate with product managers, data engineers, software developers and UK-based stakeholders. - Mentor junior engineers and contribute to AI engineering standards and best practices. Required Skills - 3 to 8 years of experience in AI, machine learning, data science, backend engineering or software development. - Hands-on experience building production-level Generative AI or LLM applications. - Strong experience with Retrieval-Augmented Generation. - Experience with vector databases and semantic search. - Strong Python programming skills. - Experience with LangChain, LlamaIndex, Haystack, Semantic Kernel or similar frameworks. - Experience using OpenAI, Azure OpenAI, AWS Bedrock, Google Gemini, Anthropic or similar LLM APIs. - Experience with open-source models and Hugging Face Transformers. - Understanding of embeddings, transformers, tokenisation and model inference. - Experience developing REST APIs and backend services. - Knowledge of Git, testing and software-engineering best practices. - Experience with cloud deployment and production application monitoring. - Strong problem-solving and communication skills. - Ability to work remotely during UK business hours. Preferred Skills - Hands-on experience fine-tuning Large Language Models. - Knowledge of LoRA, QLoRA, PEFT and model quantisation. - Experience with PyTorch or TensorFlow. - Experience with Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, Elasticsearch, OpenSearch or pgvector. - Experience with vLLM, NVIDIA Triton or Hugging Face Text Generation Inference. - Knowledge of Docker, Kubernetes and CI/CD. - Experience with MLflow, LangSmith, Weights & Biases or similar tools. - Understanding of GPU deployment and inference optimisation. - Knowledge of Azure, AWS or Google Cloud AI services. - Experience with SQL, NoSQL databases and data pipelines. - Understanding of responsible AI, model safety and LLM security. - Experience in healthcare, legal, finance, engineering or another specialised domain will be an advantage. Qualifications Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Engineering or a related field. Candidates with strong practical AI engineering experience may also be considered. Why Apply - Permanent work-from-home opportunity. - Work directly with a UK-based client. - Regular full-time and freelance options available. - Work on advanced enterprise Generative AI projects. - Hands-on exposure to RAG, LLMs, vector databases and model fine-tuning. - Build and deploy AI solutions used by real business teams. - Chance to work on the complete AI application lifecycle. - Long-term career growth in Generative AI and machine learning engineering. .

One address, no account. We’ll tell you when matching roles go live.

More at itForte Staffing Services

Related open roles

View all roles