cv

Basics

Name Yuhao Ge
Label LLM Serving and Performance Engineer
Email geyuhao33@gmail.com
Phone (217) 926-3291
Url https://geyuhao.github.io
Summary Cross-accelerator expertise spanning TPU, Trainium, and GPU systems, with work across framework, compiler, and kernel optimization for large-scale LLM serving.

Work

  • 2025.11 - Present
    Software Engineer
    Google
    TPU, vLLM, PyTorch, Pallas, Gemini
    • Led development of the PyTorch-native vLLM TPU backend from 0 to 1, making TPU a drop-in target with GPU-equivalent usability and performance for production workloads.
    • Designed and implemented the TPU serving stack: asynchronous scheduling, eager and compiled execution, static-shape bucketing and padding, distributed TP/EP/PP/DP serving, and Pallas kernel integration.
    • Co-developed TorchTPU, the PyTorch-TPU bridge underlying the backend.
    • Brought up and optimized GPT-OSS, Qwen3, and Qwen3.5 across framework, kernel, and compiler layers, reaching 80% of GB300 performance on TPUv7.
    • Designed TPU-friendly hybrid-cache management and quantization schemes covering W4A8, W8A8, KV cache, NVFP4, INT4, FP8, and MXFP4 formats.
    • Drove open-source co-development with Inferact and RadixArk through tiered CI/CD and accuracy and performance regression gates on TPU hardware.
  • 2025.03 - 2025.11
    Software Engineer
    AWS, Annapurna Labs
    ML Compiler, LLMs, LLM Systems, Architecture
    • Founding engineer of the NKI kernel language and compiler for AWS Trainium, driving language and compiler design.
    • Architected and optimized core compiler passes, including scheduling and allocation, with cycle-level optimization.
    • Designed custom FlashAttention kernels for LLM workloads on Trainium.
    • Led development of autotuning infrastructure for compiler and kernel performance.
  • 2024.05 - 2024.08
    Software Engineer Intern
    AWS, Annapurna Labs
    ML Compiler, LLM Systems
    • Developed automated kernel generation and profiling infrastructure for micro-architectural performance data and ML-based instruction-latency cost models.
    • Architected the first-generation autotuning infrastructure with modular design-space, search-strategy, and surrogate-cost-modeling components.
    • Improved Llama 3.1 performance by 14.7% through Matmul Fusion Pass autotuning.
    • Accelerated autotuning workflows by 8.6x with multi-process compilation and distributed profiling.
    • Received a Certificate of Appreciation for high-impact contributions.
  • 2022.06 - 2022.11
    Research Assistant
    University of California, Los Angeles
    ML/RL, FPGA
    • Developed a GNN-based surrogate cost model for HLS scheduling and resource estimation.
    • Designed an ML-driven design-space exploration framework for automated FPGA accelerator optimization.
    • Integrated reinforcement-learning and bandit strategies for adaptive search, improving exploration efficiency by 11%.
  • 2022.05 - 2022.08
    Software Engineer Intern
    TikTok
    ML, Graphics, Game Engine, AR/VR
    • Developed TikTok's 3D Game Engine for interactive AR/VR stickers.
    • Implemented a query-based Motion Matching system in C++ for real-time, low-latency avatar control.
    • Developed an SDK for Skeleton Retargeting across character models.

Education

  • 2023.08 - 2024.12

    Urbana, IL

    M.S.
    University of Illinois at Urbana-Champaign
    Computer Science
    • Research with Prof. Charith Mendis in ML systems, compilers, and efficient sparse attention.
  • 2019.08 - 2023.05

    Urbana, IL

    B.S.
    University of Illinois at Urbana-Champaign
    Computer Engineering
    • Highest Honors, Bronze Tablet (top 3%), and Dean's List (2020 and 2022).

Awards

  • 2023.05.01
    Highest Honors
    University of Illinois at Urbana-Champaign
  • 2023.05.01
    Bronze Tablet
    University of Illinois at Urbana-Champaign
    University-wide recognition awarded to the top 3% of graduating students.
  • 2024.08.01
    Certificate of Appreciation
    AWS Neuron Compiler Team
    Recognized for high-impact contributions to Trainium compiler and autotuning infrastructure.

Publications

Projects

  • 2023.08 - 2024.07
    Optimized GPU Code Generation for Sparse Regular Attention
    Developed SPLAT, an optimized framework for efficient sparse-MHSA targeting moderate sparsity.
    • Introduced the Affine Compressed Sparse-Row format for regular sparsity patterns.
    • Engineered GPU code-generation algorithms for ACSR.
    • Achieved 2.05x and 4.05x speedups over Triton and TVM; accepted at OOPSLA 2025.
  • 2022.01 - 2022.05
    Doodle Jump on FPGA
    Implemented Doodle Jump efficiently on an FPGA board with SystemVerilog and a NIOS II SoC.
    • Managed USB protocol and memory I/O in C.
    • Used 400 KB of memory and 0.5 W while achieving 50 Hz.
    • Won the Best Design Prize.

Skills

ML Frameworks & Serving
PyTorch
JAX
vLLM
SGLang
Kernels & Compilers
Pallas
Triton
NKI
CUDA
XLA
MLIR
TVM

Languages

Mandarin
Native
English
Fluent