ML Engineer, Inference & Optimization
PikaGenerative AI company
Palo Alto, United StatesSenior
Lightspeed Venture Partners
Spark Capital
Greycroft
SV Angel
Homebrew
Conviction
Data & AI
About the role
TL;DR
Accelerate the performance of Pika's AI-driven products through advanced inference acceleration.
- •We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products.
- •You will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.
- •Key Responsibilities Accelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
- •Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
- •Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.
- •Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.
- •Improve Training Efficiency: Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.
- •Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.
- •Requirements Experience: 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.
- •Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.
- •GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
- •AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models (LLMs).
- •Collaboration: Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.
- •Ownership Mindset: Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.
- •Bonus: Experience in enhancing training efficiency, stability, or resource optimization for large models.
Required skills
PyTorchTensorFlowHugging Face
Nice-to-have skills
TensorFlowHugging Face
Domain expertise
ai
Benefits & perks
Competitive salary in the AI industry, Equity in a fast-growing startup shaping the future of AI, Comprehensive health benefits, monthly stipends, company retreats, A supportive and collaborative office culture
Tech stack
TensorFlowPyTorchHugging Face