Design and oversee a holistic AI platform for large-scale LLM training and inference.
•Design and oversee a holistic AI rack-scale platform for trillion-parameter LLM training and high-throughput inference across hardware, software, compute, network, and storage.
•Key Responsibilities Define end-to-end architecture for highly clustered AI environments to eliminate data bottlenecks.
•Influence workload orchestration and scheduling using Kubernetes or Slurm for distributed training and inference.
•Profile and optimise full stack from DL frameworks to OS-level tuning and I/O scheduling.
•Collaborate with firmware, OS, and hardware teams to drive silicon-influencing strategy and requirements.
•Requirements Lead/principal-level experience in systems engineering, cloud architecture, HPC, or hardware for large-scale AI platforms (4+ years).
•Deep knowledge of distributed training (data/tensor/pipeline parallelism) and LLM infrastructure needs.
•Expertise with system interconnects (PCIe Gen5/6, NVMe, RDMA) and GPU-heavy bare-metal environments.
•Experience with container orchestration and infrastructure-as-code for GPU/bare-metal deployments.