Software Engineer, ML Serving - Rime Ai
UnusualAI Optimization company
San Francisco, United StatesMid
Software Engineering
About the role
TL;DR
Own the serving infrastructure connecting Rime's inference engines to the world.
- •We're hiring a Software Engineer to own the serving infrastructure that connects Rime's inference engines to the world.
- •This role sits at the intersection of ML systems and cloud infrastructure.
- •Key Responsibilities Architecture and implementation of Rime's TTS serving infrastructure.
- •Model optimization from single-node to disaggregated fleet serving.
- •Compatibility with different NVIDIA hardwares for on-prem and cloud deployments.
- •Continuous integration and deployment workflows for the model serving pipeline.
- •Site reliability: on-call rotation, monitoring, alerting, and observability.
- •Requirements Hands-on experience with real-time multinode ML serving infrastructure (NVIDIA Dynamo/Triton, vLLM, SGLang, or equivalent).
- •Experience with distributed or disaggregated model serving.
- •Strong cloud infrastructure fundamentals: Linux internals, networking, containerization (Docker, Kubernetes).
- •IaC experience (Terraform, Packer, or comparable).
- •On-call is part of the job.
Required skills
PythonLinuxDockerKubernetesTerraformgRPCWebSockets
Nice-to-have skills
AWSGoogle CloudAnsibleChefPuppet
Domain expertise
ai
Benefits & perks
equity, remote work
Tech stack
PythonLinuxDockerKubernetesTerraformgRPCWebSocketsAWSGoogle CloudAnsibleChefPuppet