Yuhao Ge
Software Engineer @ Google | LLM Serving & Performance | MSCS @ UIUC | ex-AWS Annapurna Labs
I am an LLM serving and performance engineer with cross-accelerator expertise (TPU, Trainium, GPU). My work spans framework, compiler, and kernel optimization for large-scale serving systems.
At Google, I lead development of the PyTorch-native vLLM TPU backend, making TPU a drop-in serving target with GPU-equivalent usability and performance. I work across the full TPU serving stack, including asynchronous scheduling, eager and compiled execution, static-shape bucketing, distributed serving, Pallas kernels, quantization, and TorchTPU. I also bring up and optimize open-source models such as GPT-OSS, Qwen3, and Qwen3.5 for production workloads on TPU.
Previously, I was a founding engineer of the NKI kernel language and compiler at AWS Annapurna Labs. I designed and optimized compiler passes, custom FlashAttention kernels, and autotuning infrastructure for Trainium.
I earned an M.S. in Computer Science and a B.S. in Computer Engineering from UIUC, graduating with Highest Honors and Bronze Tablet recognition. During my master’s research with Prof. Charith Mendis, I developed SPLAT, an optimized GPU code-generation framework for sparse attention accepted at OOPSLA 2025.
Earlier, I worked on ML-driven FPGA accelerator optimization at the UCLA VAST Lab with Prof. Jason Cong, and on real-time avatar animation systems at TikTok.
🎉 news
| Nov 01, 2025 | 🚀 I joined Google as a Software Engineer. I now lead development of the PyTorch-native vLLM TPU backend and work across the TPU serving stack, TorchTPU, Pallas kernels, model optimization, and quantization. |
|---|---|
| Apr 09, 2025 | Our paper, “SPLAT: A Framework for Optimised GPU Code-Generation for SParse reguLar ATtention”, has been accepted at OOPSLA 2025. |
| Mar 01, 2025 | 🚀 I joined Amazon AWS, Annapurna Labs, as a software engineer. |
| Aug 09, 2024 | I received a Certificate of Appreciation from the Neuron Compiler team at Amazon Annapurna Labs for developing autotuning infrastructure that significantly improved AWS Trainium performance. This project was recognized as one of the most impactful contributions during my internship. |
| May 20, 2024 | I joined Amazon AWS, Annapurna Labs, as a software engineer intern. I am working with the Compiler team on the AWS Neuron, which aims to optimize deep learning training and inference on AWS Trainium and Inferentia chips. |
💼 work
| Google Software Engineer | 2025.11 - Present |
| AWS, Annapurna Labs Software Engineer | 2025.3 - 2025.11 |
| AWS, Annapurna Labs Software Engineer Intern | 2024.5 - 2024.8 |
| TikTok Software Engineer Intern | 2022.5 - 2022.8 |
| University of California, Los Angeles Research Assistant | 2022.6 - 2022.11 |
🎓 education
| University of Illinois at Urbana-Champaign (UIUC) M.S. in Computer Science | 2023.8 - 2024.12 |
| University of Illinois at Urbana-Champaign (UIUC) B.S. in Computer Engineering | 2019.8 - 2023.5 |
📝 selected papers
- OOPSLA 2025
SPLAT: A Framework for Optimised GPU Code-Generation for SParse ReguLar ATtentionAhan Gupta, Yueming Yuan, Devansh Jain, Yuhao Ge, David Aponte, Yanqi Zhou, Charith MendisWe introduced the Affine Compressed Sparse-Row (ACSR) format and the SPLAT code-generation framework for regular sparse attention, achieving 2.05x and 4.05x speedups over Triton and TVM.
🥁 drumming
Drumming has been a consistent passion in my life, helping me maintain creativity and rhythm in everything I do. Check out my videos here, or visit my YouTube and Bilibili channels.