Architect and own the systems that power Salient at scale.
•We're looking for a Staff Infrastructure Engineer to architect and own the systems that power Salient at scale.
•Key Responsibilities Lead architectural decisions and technical reviews for infrastructure-critical initiatives.
•Design, build, and own the cloud infrastructure (AWS/GCP) that runs Salient
•from compute and networking to storage and observability.
•Develop scalable harnesses that enable coding agents to operate reliably without compromising system stability or code quality.
•Partner closely with the modeling team to optimize the serving and performance of GPU-intensive workloads.
•Drive reliability and performance across the stack by defining SLOs, building robust monitoring and alerting, and leading incident response and postmortems.
•Requirements 5+ years of software engineering experience, with 2+ years at the senior or staff level in infrastructure/platform roles, working on large-scale distributed systems.
•Deep expertise in cloud platforms (AWS or GCP)
•compute, networking, storage, IAM, and cost optimization.
•Expert in infrastructure-as-code, with a strong track record of building scalable automation systems.
•Extensive experience owning and scaling Kubernetes and CI/CD systems in high-throughput, production environments.
•Track record of building and operating high-availability, high-throughput distributed systems with mature observability practices.
•Strong technical communication
•able to document architecture clearly and influence engineering decisions across teams.