cv
Basics
| Name | Yuhao Ge |
| Label | LLM Serving and Performance Engineer |
| geyuhao33@gmail.com | |
| Phone | (217) 926-3291 |
| Url | https://geyuhao.github.io |
| Summary | Cross-accelerator expertise spanning TPU, Trainium, and GPU systems, with work across framework, compiler, and kernel optimization for large-scale LLM serving. |
Work
-
2025.11 - Present Software Engineer
Google
TPU, vLLM, PyTorch, Pallas, Gemini
- Led development of the PyTorch-native vLLM TPU backend from 0 to 1, making TPU a drop-in target with GPU-equivalent usability and performance for production workloads.
- Designed and implemented the TPU serving stack: asynchronous scheduling, eager and compiled execution, static-shape bucketing and padding, distributed TP/EP/PP/DP serving, and Pallas kernel integration.
- Co-developed TorchTPU, the PyTorch-TPU bridge underlying the backend.
- Brought up and optimized GPT-OSS, Qwen3, and Qwen3.5 across framework, kernel, and compiler layers, reaching 80% of GB300 performance on TPUv7.
- Designed TPU-friendly hybrid-cache management and quantization schemes covering W4A8, W8A8, KV cache, NVFP4, INT4, FP8, and MXFP4 formats.
- Drove open-source co-development with Inferact and RadixArk through tiered CI/CD and accuracy and performance regression gates on TPU hardware.
-
2025.03 - 2025.11 Software Engineer
AWS, Annapurna Labs
ML Compiler, LLMs, LLM Systems, Architecture
- Founding engineer of the NKI kernel language and compiler for AWS Trainium, driving language and compiler design.
- Architected and optimized core compiler passes, including scheduling and allocation, with cycle-level optimization.
- Designed custom FlashAttention kernels for LLM workloads on Trainium.
- Led development of autotuning infrastructure for compiler and kernel performance.
-
2024.05 - 2024.08 Software Engineer Intern
AWS, Annapurna Labs
ML Compiler, LLM Systems
- Developed automated kernel generation and profiling infrastructure for micro-architectural performance data and ML-based instruction-latency cost models.
- Architected the first-generation autotuning infrastructure with modular design-space, search-strategy, and surrogate-cost-modeling components.
- Improved Llama 3.1 performance by 14.7% through Matmul Fusion Pass autotuning.
- Accelerated autotuning workflows by 8.6x with multi-process compilation and distributed profiling.
- Received a Certificate of Appreciation for high-impact contributions.
-
2022.06 - 2022.11 Research Assistant
University of California, Los Angeles
ML/RL, FPGA
- Developed a GNN-based surrogate cost model for HLS scheduling and resource estimation.
- Designed an ML-driven design-space exploration framework for automated FPGA accelerator optimization.
- Integrated reinforcement-learning and bandit strategies for adaptive search, improving exploration efficiency by 11%.
-
2022.05 - 2022.08 Software Engineer Intern
TikTok
ML, Graphics, Game Engine, AR/VR
- Developed TikTok's 3D Game Engine for interactive AR/VR stickers.
- Implemented a query-based Motion Matching system in C++ for real-time, low-latency avatar control.
- Developed an SDK for Skeleton Retargeting across character models.
Education
-
2023.08 - 2024.12 Urbana, IL
M.S.
University of Illinois at Urbana-Champaign
Computer Science
- Research with Prof. Charith Mendis in ML systems, compilers, and efficient sparse attention.
-
2019.08 - 2023.05 Urbana, IL
B.S.
University of Illinois at Urbana-Champaign
Computer Engineering
- Highest Honors, Bronze Tablet (top 3%), and Dean's List (2020 and 2022).
Awards
- 2023.05.01
Highest Honors
University of Illinois at Urbana-Champaign
- 2023.05.01
Bronze Tablet
University of Illinois at Urbana-Champaign
University-wide recognition awarded to the top 3% of graduating students.
- 2024.08.01
Certificate of Appreciation
AWS Neuron Compiler Team
Recognized for high-impact contributions to Trainium compiler and autotuning infrastructure.
Publications
-
2025 SPLAT: A Framework for Optimised GPU Code-Generation for SParse ReguLar ATtention
OOPSLA 2025
Introduced the ACSR sparse format and SPLAT code-generation framework, achieving 2.05x and 4.05x speedups over Triton and TVM.
Projects
- 2023.08 - 2024.07
Optimized GPU Code Generation for Sparse Regular Attention
Developed SPLAT, an optimized framework for efficient sparse-MHSA targeting moderate sparsity.
- Introduced the Affine Compressed Sparse-Row format for regular sparsity patterns.
- Engineered GPU code-generation algorithms for ACSR.
- Achieved 2.05x and 4.05x speedups over Triton and TVM; accepted at OOPSLA 2025.
- 2022.01 - 2022.05
Doodle Jump on FPGA
Implemented Doodle Jump efficiently on an FPGA board with SystemVerilog and a NIOS II SoC.
- Managed USB protocol and memory I/O in C.
- Used 400 KB of memory and 0.5 W while achieving 50 Hz.
- Won the Best Design Prize.
Skills
| ML Frameworks & Serving | |
| PyTorch | |
| JAX | |
| vLLM | |
| SGLang |
| Kernels & Compilers | |
| Pallas | |
| Triton | |
| NKI | |
| CUDA | |
| XLA | |
| MLIR | |
| TVM |
Languages
| Mandarin | |
| Native |
| English | |
| Fluent |