Peiheng Hu is an accomplished machine learning engineer with over a decade of industry experience and expertise in building large-scale AI systems. He currently works at NVIDIA, where he focuses on the cutting-edge distributed LLM inference, pushing the boundaries of high-performance inference engines on the latest NVIDIA GPUs.
He holds a Master of Science in Computational Science and Engineering from Harvard University and a Bachelor of Science in Industrial Engineering Operations Research from Georgia Institute of Technology.
Previously, Peiheng served as a Principal Member of Technical Staff at Salesforce, where he led the development of the company's only unified serving platform, handling thousands of per-tenant models and LLM optimizations for Agentforce that saved millions in AI infrastructure expenses. Prior to that, he was a senior ML engineer at Microsoft Azure, where he architected distributed ML processing solutions for cloud security detection and analytics, handling billions of transactions per hour.