Source description
About the role
About the internship We are looking for a Backend Developer Intern with hands-on experience integrating open-source LLMs (such as Mistral and LLaMA) and vector databases (such as FAISS, Qdrant, Weaviate, or Milvus) in offline or edge environments. You will help develop core APIs and infrastructure to support embedding pipelines, retrieval-augmented generation (RAG), and LLM inference on local or mobile hardware. Selected intern's day-to-day responsibilities include: 1. Design and implement backend services that interface with local LLMs and vector databases. 2. Develop APIs to support prompt engineering, retrieval, and LLM inference. 3. Integrate embedding models (such as BGE and MiniLM) and manage document chunking and embedding pipelines. 4. Optimize application performance for offline, mobile, and resource-constrained environments. 5. Package and deploy LLMs using frameworks such as Ollama, llama.cpp, GGUF, and on-device runtimes. 6. Build and maintain local file ingestion, metadata tagging, indexing, and document processing pipelines. 7. Develop retrieval-augmented generation (RAG) workflows using vector databases. 8. Create and deploy agentic AI frameworks with multi-agent architectures. 9. Debug, test, and optimize backend services for reliability and scalability. 10. Collaborate with the engineering team to build production-ready AI applications. Skill(s) required Artificial intelligence Machine Learning Natural Language Processing (NLP) Earn certifications in these skills Learn Artificial intelligence Learn Machine Learning Learn NLP Who can apply Only those candidates can apply who: 1. are available for full time (in-office) internship 2. can start the internship between 14th Jul'26 and 18th Aug'26 3. are available for duration of 5 months 4. have relevant skills and interests Other requirements Skill(s) required: 1. Python or TypeScript/Node.js 2. FastAPI or Flask 3. REST APIs 4. Open-source LLMs (Mistral, LLaMA) 5. RAG 6. Vector Databases (FAISS, Qdrant, Weaviate, Milvus) 7. LangChain 8. Ollama or llama.cpp 9. Git & GitHub 10. Docker 11. Prompt Engineering 12. Problem-Solving Only those candidates can apply who: 1. Are available for a full-time internship. 2. Can start immediately. 3. Are available for the internship duration. 4. Have strong backend development experience in Python or TypeScript/Node.js. 5. Have hands-on experience with vector databases and RAG pipelines. 6. Are familiar with running open-source LLMs locally using Ollama, llama.cpp, GGUF, or similar frameworks. 7. Have experience with FastAPI or similar backend frameworks. 8. Are proficient in Git and Docker. 9. Have a good understanding of prompt engineering and LLM inference. 10. Experience with edge deployments is a plus. Perks Certificate Letter of recommendation Number of openings 5 About Infoware Website Infoware is a process-driven software solutions provider specializing in bespoke software solutions. We work with several enterprises and startups and provide them with end-to-end solutions. You may visit the company website at https://www.infowareindia.com/ Activity on Internshala Hiring since March 2020 354 opportunities posted 188 candidates hired Apply now Additional Questions × Close Sign up to continue Candidate sign up Sign up/ Login with Google Sign up with Email OR Email Password First Name Last Name Sign up By continuing as a candidate, you agree to our T&C . Already registered? Login var to_show_recaptcha = 0; Sign up to continue Candidate sign up Sign up/ Login with Google Sign up with Email OR Email Password First Name Last Name Sign up By continuing as a candidate, you agree to our T&C . Already registered? Login var to_show_recaptcha = 0; var to_trigger_already_applied_btn_ga = 0;
More at Infoware