Develop and optimize large language models on custom AI accelerators.
•Tenstorrent is building next-generation AI systems that push the boundaries of model training, inference, and large-scale distributed compute.
•The ML Models team sits at the intersection of cutting-edge AI research and high-performance hardware, bringing state-of-the-art machine learning models to life on Tenstorrent’s custom AI accelerators.
•Key Responsibilities Lead research and development efforts focused on LLM training and inference optimization.
•Train, evaluate, and optimize state-of-the-art AI models on Tenstorrent hardware.
•Improve performance through techniques such as speculative decoding, quantization, kernel fusion, flash attention, and distributed training.
•Requirements Strong Python and PyTorch experience developing and training deep learning models.
•Deep understanding of ML architectures, LLM training, and inference optimization.
•Hands-on experience training large-scale machine learning models. 4+ years of industry and/or academic experience in ML research and LLM development.
•PhD, published research, or experience with speculative decoding is highly valued.