Senior ML Engineer to lead serving engineering for Jockey Core LLM.
•Twelve Labs is building an intelligence layer for video data, enabling machines to understand video content like humans.
•This role focuses on the serving engineering for Jockey Core, the reasoning LLM powering their agentic system, optimizing for production scale and performance.
•Key Responsibilities Build benchmarks and load tests for agent traffic, measuring TTFT and inter-token latency.
•Apply inference optimization techniques (quantization, batching, speculative decoding) to meet cost and latency targets.
•Develop cost models and drive production hardening (autoscaling, capacity planning, observability).
•Collaborate with model teams to ensure efficiency gains translate to serving wins.
•Requirements Significant experience serving and optimizing large-scale LLM inference in production (vLLM, TensorRT-LLM, SGLang, or similar).
•Experience designing and operating large-scale distributed systems in high-performance GPU environments.
•Track record of driving technical decisions with measured latency/throughput/cost data.
•Experience building observability, SLOs, and failure-response for production services.
Equity, Health Insurance, Remote Work, Meal Allowance, Commuter Benefits, Professional Development, Paid Time Off, Stock Options, Gym Membership, Language Courses