Volver a ofertas

Director | AI / ML | Bengaluru | Engineering | Hybrid Cloud Engineering

Deloitte · Bengaluru, IN · India

Your work profile·      AI Data Center Architecture & Solution Design·      Design and implement AI-focused Data Center architectures aligned with Tier II, Tier III, and Tier IV standards.·      Develop end-to-end AI Data Center solutions, including retrofitting traditional CPU-based data centers into AI Factories.·      Create advisory...

Your work profile


·      AI Data Center Architecture & Solution Design

·      Design and implement AI-focused Data Center architectures aligned with Tier II, Tier III, and Tier IV standards.

·      Develop end-to-end AI Data Center solutions, including retrofitting traditional CPU-based data centers into AI Factories.

·      Create advisory documents, RFPs, technical proposals, and commercial proposals for AI Data Center engagements.

·      Design AI infrastructure solutions across hyperscalers (AWS, Azure, GCP, OCI) and NVIDIA Cloud Partners.

·      Prepare HLDs, LLDs, network diagrams, rack layouts, BOQs, and TCO models.

·      AI Networking & Fabric Architecture

·      Architect and deploy InfiniBand and NVIDIA Spectrum Ethernet fabrics for AI workloads.

·      Design and implement Spine-Leaf network architectures using EVPN-VXLAN overlays.

·      Configure and optimize BGP, ECMP, RoCE, and high-performance networking environments.

·      Lead Cumulus Linux-based deployments and network automation initiatives.

·      Optimize network performance, latency, throughput, and congestion management for AI environments.

·      AI Compute & GPU Infrastructure

·      Design and size GPU clusters using NVIDIA H100, H200, B200, B300, DGX, and AI Factory platforms.

·      Perform GPU capacity planning and workload profiling for AI and ML use cases.

·      Implement GPU virtualization and Multi-Instance GPU (MIG) architectures.

·      Support AI training and inference infrastructure deployments.

·      AI Storage & Platform Engineering

·      Design AI storage solutions utilizing NAS, SAN, NVMe, Object Storage, NFS, iSCSI, Fibre Channel, and parallel file systems.

·      Implement and manage Kubernetes-based AI platforms, including OpenShift and VMware Tanzu.

·      Deploy and integrate RUN and Slurm workload schedulers for GPU orchestration.

·      Ensure seamless integration of AI platforms with existing enterprise infrastructure.

·      Monitoring, Observability & Operations

·      Implement NVIDIA UFM, NVIDIA Mission Control, and NetQ for infrastructure monitoring and observability.

·      Configure telemetry, validation, troubleshooting, and fabric management workflows.

·      Drive infrastructure benchmarking, performance optimization, and capacity planning initiatives.

·      Support POCs, design validation exercises, production rollouts, and operational readiness activities.

·      Cloud & AI Services

·      Design AI infrastructure solutions across AWS, Azure, GCP, and OCI.

·      Enable AI services integration across hybrid and multi-cloud environments.

·      Provide guidance on AI platform adoption, scalability, and operational best practices.

 

Key skills required


Data Center Infrastructure

AI Networking

AI Compute & Platforms

AI Storage

Orchestration & Container Platforms

AI Software Stack

Cloud Technologies

Encuentra más ofertas como esta

Explora más ofertas activas de esta empresa o crea una cuenta en Insider Jobs para buscar, guardar y seguir oportunidades en todo el job board.