Padmi

Senior AI Research Engineer

BangalorePosted 2 months ago
Computer ResearchSeniorFull Time; Regular
Apply at Trianz

Opens the source posting on shine.com

Source description

About the role

View original

Company Overview Trianz is an applied AI solutions company that accelerates customer business transformation through AI powered "Transformation Services as a Software Model". With 25+ years of transforming enterprises, we've evolved to a product-led, platform-driven organization serving global enterprises across Financial Services, Insurance, Healthcare, Hi-Tech, Manufacturing, and other industries. With global presence across 4 continents, our platform portfolio under the unified Concierto brand delivers end-to-end transformations including solutions for Migrate, Manage, Maximize, Modernize, Insights & Agentic AI, and SecOps - delivered through strategic partnerships with leading hyperscalers. We're building the premier innovation-led organization in the digital transformation space through AI-first methodologies and data-driven excellence - RevolutionAIzing Transformations. Role Overview Every AI architectural decision at Trianz - which model to deploy, which hardware to run it on, whether a smaller fine-tuned model outperforms a larger API-consumed model on a specific enterprise task rests on evidence. You produce that evidence. As Senior AI Research Engineer, your work is not theoretical. It is the empirical foundation that justifies architectural bets with real infrastructure cost, customer-facing latency, and enterprise compliance consequences. When the team decides to deploy a 7B model on CPU rather than calling a 70B model via API, that decision rests on your benchmark data your evaluation harness, your reproducibility standards, your domain-specific task results, and your recommendation. This is not a role for someone who runs general benchmarks from leaderboards and presents the numbers. General benchmarks tell you how a model performs on academic tasks. Enterprise AI decisions require evaluation on the specific task types document extraction, policy reasoning, structured output generation, multi-step planning that actually run in production. You design the evaluation suite. You run the experiments. You produce the recommendation with the data behind it. You will also run fine-tuning experiments to test whether a smaller, domain-adapted model can match or exceed a larger general model on specific tasks a question with direct infrastructure cost and sovereignty implications. You will maintain the model leaderboard as new models are released and evaluate quantization strategies across the full accuracy-versus-performance tradeoff curve. If you consume existing benchmarks without questioning their applicability to enterprise tasks, this is not your role. If you design evaluation frameworks, produce reproducible results, and drive architectural decisions with data you are exactly who we are looking for. Key Responsibilities LLM Evaluation Framework Design Design and implement the LLM evaluation framework: benchmark suite, evaluation harness, and reproducibility standardsRun systematic evaluations across 5B, 10B, and 30B parameter models on enterprise-specific task typesBenchmark open-source models (Llama, Mistral, Qwen, Phi) vs closed models (GPT-4, Claude, Gemini) on accuracy, latency, and costEvaluate model performance on CPU-only inference vs GPU inference to validate hardware routing decisionsDesign domain-specific evaluation tasks: document extraction, policy reasoning, structured output generation, multi-step planningMaintain reproducibility standards across all evaluation runs: seed control, dataset versioning, harness version tracking Model Selection & Architectural Evidence Produce model selection recommendations with supporting data for every AI architectural decisionDesign the evidence pack that justifies open-source vs closed model decisions for specific enterprise use casesEvaluate cost-per-task across model sizes and inference modalities: API vs self-hosted vs CPU vs GPUMaintain the model leaderboard: continuously updated benchmarks as new models are releasedProduce rigorous, reproducible evaluation reports that architects and engineering teams act on Fine-Tuning Research Design and run fine-tuning experiments for smaller models on domain-specific enterprise tasksEvaluate whether a fine-tuned 7B model matches or exceeds a 70B general model on specific task typesApply parameter-efficient fine-tuning methods: LoRA, QLoRA, DoRA, PEFTDesign training data curation pipelines for domain-specific fine-tuningEvaluate fine-tuned model quality: accuracy, hallucination rate, instruction following, output format fidelity Quantization & Inference Research Evaluate quantization strategies (GPTQ, AWQ, GGUF) and their accuracy vs performance trade-offsBenchmark quantized models against full-precision baselines on enterprise task typesEvaluate throughput, latency, and memory footprint across quantization levelsProduce quantization selection recommendations for specific model sizes and hardware configurationsResearch emerging quantization and Company

One address, no account. We’ll tell you when matching roles go live.

More at Trianz

Related open roles

View all roles