Skip to content
Boson AI logo

Site Reliability Engineer

Boson AIVoice AI, company
TorontoSenior
Su Hua
Temasek logo
Temasek
Software Engineering

About the role

TL;DR

Design and optimize high-performance networking infrastructure for AI/ML operations.

  • We're seeking an experienced Network Engineer to design, build, and optimize the high-performance networking infrastructure powering our AI/ML operations in Toronto.
  • Key Responsibilities Design, operate, and improve reliable infrastructure for AI training and inference workloads Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate Requirements 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role Strong hands-on expertise in at least one of the following: Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand Experience operating production systems with a focus on availability, performance, security, and automation
View original posting →

Required skills

KubernetesTerraformAnsiblePrometheusGrafanaLinux

Domain expertise

ai

Tech stack

KubernetesTerraformAnsiblePrometheusGrafanaLinux

Similar jobs

S

Senior Red Team Operator

OpenAI logo

AI Support Engineer - Dublin (Weekend Shift)

OpenAI logo

Systems Software Engineer, Silicon Bringup