Padmi

Gen AI/LLM Engineer

MumbaiPosted 3 months ago
Software engineeringMid-levelFull Time; Regular
Apply at Zorba AI

Opens the source posting on shine.com

Source description

About the role

View original

As an experienced LLM Engineer at our leading consulting firm specializing in Enterprise Generative AI and Large Language Model (LLM) services, your primary responsibility will be to design, fine-tune, and deploy LLM-based solutions for various use-cases such as search, summarization, agents, and domain-specific assistants. Your role will involve the following key responsibilities: - Design, fine-tune, and validate LLMs for production use-cases, including instruction tuning, supervised fine-tuning, and parameter-efficient tuning (LoRA/adapters). - Implement retrieval-augmented generation (RAG) pipelines, which involve embeddings, vector search, chunking, and context assembly for high-recall responses. - Optimize inference for latency and cost by utilizing techniques such as quantization, model pruning, batching, and deployment with optimized runtimes (CUDA, Triton, bitsandbytes where applicable). - Build backend services and APIs to serve LLM inference and orchestration using containerized deployments (Docker/Kubernetes) and CI/CD pipelines. - Collaborate with product, data engineering, and ML teams to integrate LLMs into production flows, monitor model performance, and set up automated retraining/rollbacks. - Create reproducible training pipelines, implement evaluation suites, and produce documentation and runbooks for model governance and observability. In order to excel in this role, you must possess the following qualifications: Must-Have: - 4+ years of hands-on experience working with LLMs or advanced NLP models in production contexts. - Proficiency in Python for ML engineering and model development. - Experience with PyTorch and Hugging Face Transformers for training and fine-tuning. - Practical experience implementing RAG and vector search using tools like FAISS or similar vector databases. - Familiarity with LangChain (or equivalent orchestration) and integration with LLM APIs (OpenAI, Anthropic, etc.). - Experience containerizing and deploying ML services using Docker; familiarity with Kubernetes is a plus. Preferred: - Experience with inference optimizations such as quantization (bitsandbytes), Triton, or GPU-accelerated serving. - Exposure to distributed training frameworks (DeepSpeed) and cloud MLOps platforms (SageMaker, Azure ML, GCP AI Platform). - Knowledge of monitoring, logging, and model-evaluation frameworks for production LLMs (MLflow, Prometheus, Grafana). In addition to the exciting and challenging work you will be doing, you will also enjoy working in a collaborative, engineering-driven culture with a strong focus on ownership and rapid iteration. You will have the opportunity to build end-to-end LLM products for enterprise clients and influence architecture decisions while having hands-on access to GPU infrastructure and cross-functional product teams. Apply now and be part of our dynamic team where your skills in python, docker, llm, agentic, pytorch, and cuda will be put to great use! As an experienced LLM Engineer at our leading consulting firm specializing in Enterprise Generative AI and Large Language Model (LLM) services, your primary responsibility will be to design, fine-tune, and deploy LLM-based solutions for various use-cases such as search, summarization, agents, and domain-specific assistants. Your role will involve the following key responsibilities: - Design, fine-tune, and validate LLMs for production use-cases, including instruction tuning, supervised fine-tuning, and parameter-efficient tuning (LoRA/adapters). - Implement retrieval-augmented generation (RAG) pipelines, which involve embeddings, vector search, chunking, and context assembly for high-recall responses. - Optimize inference for latency and cost by utilizing techniques such as quantization, model pruning, batching, and deployment with optimized runtimes (CUDA, Triton, bitsandbytes where applicable). - Build backend services and APIs to serve LLM inference and orchestration using containerized deployments (Docker/Kubernetes) and CI/CD pipelines. - Collaborate with product, data engineering, and ML teams to integrate LLMs into production flows, monitor model performance, and set up automated retraining/rollbacks. - Create reproducible training pipelines, implement evaluation suites, and produce documentation and runbooks for model governance and observability. In order to excel in this role, you must possess the following qualifications: Must-Have: - 4+ years of hands-on experience working with LLMs or advanced NLP models in production contexts. - Proficiency in Python for ML engineering and model development. - Experience with PyTorch and Hugging Face Transformers for training and fine-tuning. - Practical experience implementing RAG and vector search using tools like FAISS or similar vector databases. - Familiarity with LangChain (or equivalent orchestration) and integration with LLM APIs (OpenAI, Anthropic, etc.). - Experience containerizing and deploying ML services using Docker; familiarity

One address, no account. We’ll tell you when matching roles go live.

More at Zorba AI

Related open roles

View all roles