Padmi

Senior Generative AI Engineer | 3+ YOE | MUMBAI

MumbaiPosted 1 month ago
Software engineeringSeniorFull Time; Regular
Apply at Square Yards

Opens the source posting on shine.com

Source description

About the role

View original

About the team We build and run the AI systems that Square Yards' sales, marketing and operations teams use every day about 2 billion LLM tokens a day in production. This is not a research lab and it is not a slide deck everything we ship talks to real customers, real agents, or real listings. Today that includes: A real-time AI voice calling platform that runs outbound cold-calling, lead qualification and inbound listener campaigns across multiple telephony providers, with per-campaign prompt logic, live audio streaming, barge-in handling and CRM write-back. A conversational intelligence stack that transcribes, diarizes and scores every call utterance-level sentiment, interruption detection, objection mining, QA scorecards, and post-call coaching advice exposed as a public API product. Retrieval-grounded assistants for property discovery, internal HR workflows, and in-product support widgets embedded directly in our dashboards. A generative media pipeline that turns property listings, floor plans and locality data into marketing videos and imagery at scale. Operational tooling health-check services, campaign managers, analytics dashboards that keeps all of the above observable and debuggable. We move fast, own our systems end to end, and are small enough that one engineer's work is visible across the business. What you'll do Own an AI product surface end to end from prompt and model design through API, deployment, monitoring and iteration based on production data. You will not be handed a spec; you'll help write it. Build low-latency real-time voice agents. Streaming ASR and TTS, voice activity detection, noise suppression, turn-taking and interruption handling, tool calling mid-conversation, and shaving hundreds of milliseconds off a response loop. Design LLM pipelines that produce reliable structured output. Schema-first extraction, grading and classification over long transcripts and multimodal inputs, with retries, validation, fallbacks and cost/latency budgets that hold up at volume. Build and improve retrieval systems chunking and embedding strategy, vector search, hybrid retrieval, reranking, and honest evaluation of whether retrieval actually improved the answer. Design agentic workflows tool use, MCP servers, multi-step planning, and the guardrails that keep an autonomous system from doing something expensive or embarrassing. Take models to production: containerized services, self-hosted inference for open-weight models on GPU, batching and concurrency tuning, and sensible tradeoffs between hosted APIs and self-hosting. Build evaluation into everything. Golden sets, regression suites, A/B comparisons across models and prompts, and dashboards that tell us when quality drifts before a stakeholder does. Work directly with sales, operations and product stakeholders to turn a vague business problem into a scoped AI system and to say no when AI is the wrong tool. Mentor engineers on the team, review code and prompts, and raise the bar on how we build. What we're looking for Required 3+ years building software, with 2+ years shipping ML or LLM-backed systems that real users depend on. Strong Python. You write services, not just notebooks FastAPI/Flask, async I/O, queues and background workers, clean error handling. Deep, practical experience with LLM APIs (OpenAI, Google Gemini, Anthropic Claude, Groq or similar): prompt design, function/tool calling, structured output, streaming, context management, token and cost control. Real production RAG experience you can explain what you tried, what failed, and how you measured it. Comfortable with databases and data plumbing: MongoDB or a comparable document store, SQL, and streaming/change-data patterns. Docker, Linux, cloud deployment (GCP preferred), and enough CI/ops sense to keep your own services alive. Git, code review, and the discipline to write code someone else can pick up. Strongly preferred Real-time audio or telephony experience streaming ASR/TTS, VAD, WebSockets, jitter and latency debugging, or integration with providers like Twilio, Exotel, Ozonetel. Speech and audio tooling: Whisper-class ASR, diarization, ElevenLabs / Azure Speech / Google TTS, librosa, ffmpeg. Self-hosted inference: vLLM or similar, GPU memory and throughput tuning, quantization. Orchestration and typed-prompt frameworks LangChain / LangGraph, BAML or equivalent. Computer vision or generative media: PyTorch, transformers, diffusion models, segmentation, CLIP-style embeddings, programmatic video generation. Basic front-end (React + TypeScript, Streamlit, Gradio) to ship a usable POC interfaces. MLOps tooling experiment tracking, model registry, observability for LLM systems. Nice to have Experience in real estate, fintech, or another high-volume outbound sales environment. Familiarity with Indian telephony, regulatory constraints around outbound calling, or multilingual/code-switched (HindiEnglish) conversational .

One address, no account. We’ll tell you when matching roles go live.

More at Square Yards

Related open roles

View all roles