Yuhao Ge

Software Engineer @ Google | LLM Serving & Performance | MSCS @ UIUC | ex-AWS Annapurna Labs

I am an LLM serving and performance engineer with cross-accelerator expertise (TPU, Trainium, GPU). My work spans framework, compiler, and kernel optimization for large-scale serving systems.

At Google, I lead development of the PyTorch-native vLLM TPU backend, making TPU a drop-in serving target with GPU-equivalent usability and performance. I work across the full TPU serving stack, including asynchronous scheduling, eager and compiled execution, static-shape bucketing, distributed serving, Pallas kernels, quantization, and TorchTPU. I also bring up and optimize open-source models such as GPT-OSS, Qwen3, and Qwen3.5 for production workloads on TPU.

Previously, I was a founding engineer of the NKI kernel language and compiler at AWS Annapurna Labs. I designed and optimized compiler passes, custom FlashAttention kernels, and autotuning infrastructure for Trainium.

I earned an M.S. in Computer Science and a B.S. in Computer Engineering from UIUC, graduating with Highest Honors and Bronze Tablet recognition. During my master’s research with Prof. Charith Mendis, I developed SPLAT, an optimized GPU code-generation framework for sparse attention accepted at OOPSLA 2025.

Earlier, I worked on ML-driven FPGA accelerator optimization at the UCLA VAST Lab with Prof. Jason Cong, and on real-time avatar animation systems at TikTok.

🎉 news

Nov 01, 2025 🚀 I joined Google as a Software Engineer. I now lead development of the PyTorch-native vLLM TPU backend and work across the TPU serving stack, TorchTPU, Pallas kernels, model optimization, and quantization.
Apr 09, 2025 Our paper, “SPLAT: A Framework for Optimised GPU Code-Generation for SParse reguLar ATtention”, has been accepted at OOPSLA 2025.
Mar 01, 2025 🚀 I joined Amazon AWS, Annapurna Labs, as a software engineer.
Aug 09, 2024 I received a Certificate of Appreciation from the Neuron Compiler team at Amazon Annapurna Labs for developing autotuning infrastructure that significantly improved AWS Trainium performance. This project was recognized as one of the most impactful contributions during my internship.
May 20, 2024 I joined Amazon AWS, Annapurna Labs, as a software engineer intern. I am working with the Compiler team on the AWS Neuron, which aims to optimize deep learning training and inference on AWS Trainium and Inferentia chips.

💼 work

Google
Software Engineer
2025.11 - Present
AWS, Annapurna Labs
Software Engineer
2025.3 - 2025.11
AWS, Annapurna Labs
Software Engineer Intern
2024.5 - 2024.8
TikTok
Software Engineer Intern
2022.5 - 2022.8
University of California, Los Angeles
Research Assistant
2022.6 - 2022.11

🎓 education

University of Illinois at Urbana-Champaign (UIUC)
M.S. in Computer Science
2023.8 - 2024.12
University of Illinois at Urbana-Champaign (UIUC)
B.S. in Computer Engineering
2019.8 - 2023.5

📝 selected papers

  1. OOPSLA 2025
    splat.jpg
    SPLAT: A Framework for Optimised GPU Code-Generation for SParse ReguLar ATtention
    Ahan Gupta, Yueming Yuan, Devansh Jain, Yuhao Ge, David Aponte, Yanqi Zhou, Charith Mendis
    We introduced the Affine Compressed Sparse-Row (ACSR) format and the SPLAT code-generation framework for regular sparse attention, achieving 2.05x and 4.05x speedups over Triton and TVM.

🥁 drumming

Drumming has been a consistent passion in my life, helping me maintain creativity and rhythm in everything I do. Check out my videos here, or visit my YouTube and Bilibili channels.

Linkin Park - Numb
Bring Me The Horizon - Avalanche (Drum Cover)
LiSA - 紅蓮華 (Drum Cover)