Source description
About the role
LNT/NSA/1803310 IU72-Larsen & Toubro Vyoma AMN Tower, Powai Posted On 20 Jul 2026 End Date 16 Jan 2027 Required Experience6 - 10 years Skills Knowledge & Posting Location Lustre GPFS storage admin AI GPGPU Cloud Minimum Qualification Bachelor of Technology (BTech) Job Description Job Purpose Provide high throughput, consistent storage tiers (Scratch/HPS + Object) for large scale training data ingest, checkpoints, and inference artifacts. Roles & Responsibilities Implementation Design/expand Lustre/BeeGFS HPS; NVMe oF and Object (S3) tiers; align with AI dataflow and GDS. Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets. Operations Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention. Proactive detection of hot spots and metadata contention; schema for small file handling. Performance & Optimization Tune RDMA paths, page cache, IO schedulers; validate end to end I/O profiles for LLM training/inference. Reliability & Incident Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA GPU clusters, ensuring rapid service restoration and minimal impact to business-critical workloads. Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms. Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures. Conduct comprehensive Root Cause Analysis (RCA) for storage-related outages affecting GPU training, inferencing, AI pipelines, and high-performance computing (HPC) workloads. Backup/DR for critical datasets; test restore time and RPO/RTO regularly; corruption and split brain handling. Security & Compliance Multi tenant encryption at rest and in transit, POSIX/ACLs, S3 policies, retention/legal holds. Experience & Educational Requirement BE/B-Tech or equivalent with Computer Science or Electronics & Communication Certification- SNIA; vendor (NetApp/Dell/VAST) preferred. RELEVANT EXPERIENCE 712 years distributed storage; hands on with Lustre/BeeGFS/Ceph, NVMe oF, and S3 in GPU environments. Tools / Tech Lustre/BeeGFS tooling, Ceph; GDS; Perf tools (fio, fs digests); Prometheus/Grafana; Ansible. LNT/NSA/1803310 IU72-Larsen & Toubro Vyoma AMN Tower, Powai Posted On 20 Jul 2026 End Date 16 Jan 2027 Required Experience6 - 10 years Skills Knowledge & Posting Location Lustre GPFS storage admin AI GPGPU Cloud Minimum Qualification Bachelor of Technology (BTech) Job Description Job Purpose Provide high throughput, consistent storage tiers (Scratch/HPS + Object) for large scale training data ingest, checkpoints, and inference artifacts. Roles & Responsibilities Implementation Design/expand Lustre/BeeGFS HPS; NVMe oF and Object (S3) tiers; align with AI dataflow and GDS. Establish namespace, OST/MDT layout, stripe/RAID policies; tiering for warm/cold datasets. Operations Capacity/performance planning; rebalance and failover testing; automate snapshots and checkpoint retention. Proactive detection of hot spots and metadata contention; schema for small file handling. Performance & Optimization Tune RDMA paths, page cache, IO schedulers; validate end to end I/O profiles for LLM training/inference. Reliability & Incident Lead P0/P1 critical incident response for enterprise-scale AI/ML storage infrastructure supporting NVIDIA GPU clusters, ensuring rapid service restoration and minimal impact to business-critical workloads. Act as the storage SME during major incidents involving BeeGFS, Lustre, GPFS (IBM Spectrum Scale), NFS, NVMe-oF, Parallel File Systems, and Object Storage platforms. Perform deep-dive troubleshooting and resolution of storage performance degradation, metadata bottlenecks, I/O latency spikes, filesystem corruption, capacity exhaustion, and hardware failures. Conduct comprehensive Root Cause Analysis (RCA) for storage-related outages affecting GPU training, inferencing, AI pipelines, and high-performance computing (HPC) workloads. Backup/DR for critical datasets; test restore time and RPO/RTO regularly; corruption and split brain handling. Security & Compliance Multi tenant encryption at rest and in transit, POSIX/ACLs, S3 policies, retention/legal holds. Experience & Educational Requirement BE/B-Tech or equivalent with Computer Science or Electronics & Communication Certification- SNIA; vendor (NetApp/Dell/VAST) preferred. RELEVANT EXPERIENCE 712 years distributed storage; hands on with Lustre/BeeGFS/Ceph, NVMe oF, and S3 in GPU environments. Tools / Tech Lustre/BeeGFS tooling, Ceph; GDS; Perf tools (fio, fs digests); Prometheus/Grafana; Ansible
More at Larsen & Toubro