Lead a team ensuring reliability and operational excellence for AI development infrastructure.
•Tenstorrent is seeking a Principal Debug & Site Reliability Engineering Lead to guide a team responsible for ensuring the reliability, observability, and operational excellence of AI development environments.
•This role involves debugging complex hardware and software interactions and improving engineering workflows.
•Key Responsibilities Lead a team focused on the reliability and operational health of engineering infrastructure.
•Drive root-cause analysis for complex issues across hardware, software, and infrastructure.
•Build and improve debugging methodologies, monitoring systems, and automation.
•Partner with cross-functional teams to prioritize work and enhance platform reliability.
•Mentor engineers and establish technical direction for key initiatives.
•Requirements 10+ years of experience in building and operating complex software, infrastructure, SRE, or systems engineering environments.
•Strong Linux systems expertise with deep debugging experience across OS, networking, distributed services, hardware, and firmware.
•Proficient in automation and software development using Python, C++, Go, or Bash.
•Familiarity with observability platforms like Prometheus, Grafana, OpenTelemetry, or ELK.
•Passion for mentoring engineers and driving reliability through automation and collaboration.