Designs and implements high-performance ML serving and inference infrastructure.
•Design and implement high-performance model serving and inference infrastructure to maximize throughput and minimize latency for generative media models, and build profiling and monitoring tools.