About the role
TL;DR
Design and optimize high-performance networking infrastructure for AI/ML operations.
- •We're seeking an experienced Network Engineer to design, build, and optimize the high-performance networking infrastructure powering our AI/ML operations in Toronto.
- •Key Responsibilities Design, operate, and improve reliable infrastructure for AI training and inference workloads Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate Requirements 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role Strong hands-on expertise in at least one of the following: Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand Experience operating production systems with a focus on availability, performance, security, and automation
Required skills
KubernetesTerraformAnsiblePrometheusGrafanaLinux
Domain expertise
ai
Tech stack
KubernetesTerraformAnsiblePrometheusGrafanaLinux