Padmi

Ai Ml Engineer

ChennaiPosted 1 month ago
Software engineeringMid-level
Apply at Tata Consultancy Services

Opens the source posting on naukri.com

Source description

About the role

View original

TCS Hiring!!! Role & responsibilities Proficiency in Python/SQL, understanding of ML frameworks (PyTorch, TensorFlow), and data pre-processing logic. Core AI/ML Fundamentals knowledge Machine Learning Operations (MLOps) Should have good understanding about AI/ML Ecosystem Tools Strong understanding of GCP compute, storage, IAM, Vertex AI Exposure in managing GPU/TPU environments Debugging & Support Skills Ability to analyze logs, trace errors, and troubleshoot Knowledge of working with different APIs Identifying, resolving technical issues, and diagnosing the root cause of technical problems. Providing technical assistance to users, both internal and external, through various channels Experience working with Conversation Agents Communicating technical information clearly and concisely to users and stakeholders Support with 24x7 operations (Rotational Shifts) English language (verbal and written) Proficiency is must 1. Troubleshooting Model & Pipeline Issues The core of the role is diagnosing technical failures within the ML lifecycle. API & Integration Support: Debugging RESTful API integrations between the client's application and the AI model. Inference Failures: Investigating why a model is failing to provide predictions (e.g., timeout issues, memory overflows, or incorrect input data formatting). Environment Configuration: Assisting clients with setup issues related to Docker, Kubernetes, or cloud-specific ML environments (AWS SageMaker, Azure ML, etc.). 2. Data & Performance Monitoring AI products are only as good as the data fed into them. Data Quality Checks: Helping customers identify if their input data is the cause of poor model performance (e.g., missing values, incorrect data types, or schema mismatches). Monitoring Drift: Assisting in identifying "Model Drift"where the AIs performance degrades over time because real-world data has changed compared to the training data. Accuracy Inquiries: Explaining to customers why a model produced a specific result using interpretability tools (like SHAP or LIME) or logs. 3. Product Education & Technical Documentation Because AI is complex, the TSR serves as a technical teacher. Knowledge Base Authoring: Writing guides on "Best Practices for Prompt Engineering" or "How to Fine-tune Hyperparameters" for the specific platform. Customer Onboarding: Guiding new technical users through the initial setup of their ML experiments. Translating Documentation: Taking complex engineering release notes and making them understandable for the customer's IT team. 4. The "Feedback Bridge" to Engineering The TSR is the first to see patterns in how the product fails in the real world. Bug Reporting: Identifying and reproducing software bugs in the ML platform and escalating them to the ML Engineers or DevOps team. Feature Requests: Aggregating customer feedback regarding missing ML capabilities (e.g., "Customers are asking for support for PyTorch 2.0"). Edge Case Discovery: Documenting unique edge cases where the AI model consistently fails, which helps the data science team improve future training sets. Responsibility for maintaining SLA (Service Level Agreements) and identifying product bugs for the engineering team. Translating complex AI concepts for non-technical users and managing high-pressure customer interactions.

One address, no account. We’ll tell you when matching roles go live.

More at Tata Consultancy Services

Related open roles

View all roles