Source description
About the role
As a GenAI & Computer Vision Engineer at our company, you will be responsible for delivering production-grade AI solutions and tackling domain-specific challenges in generative AI, LLM engineering, computer vision development, and MLOps & deployment. Role Overview: In this role, you will fine-tune and evaluate LLMs for specialized tasks, deploy high-throughput inference pipelines, design agent-based workflows, build scalable inference APIs, develop and optimize CV models, implement real-time pipelines, handle data challenges, containerize models and services, define SLAs, evangelize best practices, and mentor junior engineers. Key Responsibilities: - Fine-tune and evaluate LLMs for specialized tasks - Deploy high-throughput inference pipelines - Design agent-based workflows integrating vector databases - Build scalable inference APIs - Develop and optimize CV models for detection, segmentation, classification, and tracking - Implement real-time pipelines - Containerize models and services with Docker - Define SLAs for latency, accuracy, and throughput - Mentor junior engineers on reproducible research and code reviews Qualification Required: - Proficiency in LLM Frameworks & Tooling: Hugging Face Transformers, Ollama, vLLM, or LLaMA - Proficiency in Agent & Retrieval Tools: LangChain or LangGraph; RAG with Pinecone, Weaviate, or Milvus - Proficiency in Inference Serving: Triton Inference Server; FastAPI or Flask - Proficiency in Computer Vision Frameworks & Libraries: PyTorch or TensorFlow; OpenCV (cv2) or NVIDIA DeepStream - Proficiency in Model Optimization: TensorRT; ONNX Runtime; Torch-TensorRT - Proficiency in MLOps & Versioning: Docker and Kubernetes (KServe, SageMaker); MLflow or DVC - Proficiency in Monitoring & Observability: Prometheus; Grafana - Proficiency in Cloud Platforms: AWS (SageMaker, EC2/EKS) or GCP (Vertex AI, AI Platform) or Azure ML (AKS, ML Studio) - Proficiency in Programming Languages: Python (required); C++ or Go (preferred) In addition to the technical qualifications, you should hold a Bachelors or Masters in Computer Science, Electrical Engineering, AI/ML, or a related field and have 35 years of professional experience in shipping generative and vision-based AI models in production. Strong problem-solving skills, excellent communication skills, and the ability to debug issues effectively are also essential. Let's work together to solve typical domain challenges such as LLM hallucination & safety, vector DB scaling, inference latency, concept & data drift, and multi-modal coordination in a collaborative environment. If you are looking to join a team that powers businesses with digital solutions and impactful experiences, come join Auriga IT. Visit our website at Auriga IT to learn more about us. As a GenAI & Computer Vision Engineer at our company, you will be responsible for delivering production-grade AI solutions and tackling domain-specific challenges in generative AI, LLM engineering, computer vision development, and MLOps & deployment. Role Overview: In this role, you will fine-tune and evaluate LLMs for specialized tasks, deploy high-throughput inference pipelines, design agent-based workflows, build scalable inference APIs, develop and optimize CV models, implement real-time pipelines, handle data challenges, containerize models and services, define SLAs, evangelize best practices, and mentor junior engineers. Key Responsibilities: - Fine-tune and evaluate LLMs for specialized tasks - Deploy high-throughput inference pipelines - Design agent-based workflows integrating vector databases - Build scalable inference APIs - Develop and optimize CV models for detection, segmentation, classification, and tracking - Implement real-time pipelines - Containerize models and services with Docker - Define SLAs for latency, accuracy, and throughput - Mentor junior engineers on reproducible research and code reviews Qualification Required: - Proficiency in LLM Frameworks & Tooling: Hugging Face Transformers, Ollama, vLLM, or LLaMA - Proficiency in Agent & Retrieval Tools: LangChain or LangGraph; RAG with Pinecone, Weaviate, or Milvus - Proficiency in Inference Serving: Triton Inference Server; FastAPI or Flask - Proficiency in Computer Vision Frameworks & Libraries: PyTorch or TensorFlow; OpenCV (cv2) or NVIDIA DeepStream - Proficiency in Model Optimization: TensorRT; ONNX Runtime; Torch-TensorRT - Proficiency in MLOps & Versioning: Docker and Kubernetes (KServe, SageMaker); MLflow or DVC - Proficiency in Monitoring & Observability: Prometheus; Grafana - Proficiency in Cloud Platforms: AWS (SageMaker, EC2/EKS) or GCP (Vertex AI, AI Platform) or Azure ML (AKS, ML Studio) - Proficiency in Programming Languages: Python (required); C++ or Go (preferred) In addition to the technical qualifications, you should hold a Bachelors or Masters in Computer Science, Electrical Engineering, AI/ML, or a related field and have 35 years o
More at aurigait