Ethan Zhen LI

About me

I am actively looking for academic collaboration. We are working with NVIDIA H- and B-series GPU platforms, including DGX Spark, B200/H200, GB200 and B300 systems. Feel free to contact me if you are interested in NVFP4/FP4/MXFP8 end-to-end training, low-bit RL, or efficient reasoning models.

Research Interests

  • AI Infrastructure: large-scale LLM training systems, distributed training, memory optimization, and system-algorithm co-design.
  • Efficient LLM Training and Inference: FP8/FP4/NVFP4 training, INT4 rollout, quantization-aware training, model compression, and inference acceleration.
  • Reasoning under Efficiency Constraints: preserving and improving mathematical and multimodal reasoning capabilities in low-bit LLMs.

Research Highlights

  • Ultra-Low-Bit LLM Pretraining in NVFP4 with System-Level Optimization: Designed end-to-end NVFP4 training pipelines with framework-level optimization, investigated per-token scaling, block-level scaling and module-sensitive precision allocation, and explored scalable low-bit training on large GPU clusters.
  • End-to-End FP8 Training for Reasoning-Capable LLMs at Scale: Developed FP8 training recipes covering continual pretraining, supervised fine-tuning, evaluation and deployment-oriented precision consistency, with hybrid-granularity FP8 quantization and BF16-level reasoning performance.
  • INT4 RL with Real W4A16 Rollout for Train-Inference Alignment: Built an INT4 QAT-RL pipeline that combines training-side fake quantization with serving-side real W4A16 rollout, improving rollout throughput on Qwen3-30B-A3B while keeping reasoning performance stable.

Publications and Manuscripts

(*Equal contribution)

Patents

  • InfiDeck: An AI Operating System with Integrated Hardware Extensions · China Patent
  • Reinforcement Learning Model Training Method and Device · U.S., China and Singapore Patent

Grants

  • Numerical Stability and End-to-End Performance Optimization of Low-Bit Mixed-Precision Training for Large Language Models · Tsinghua-PolyU Joint Research Initiative Fund
  • Scaling Low-Bit Training for Efficient Large Model Deployment · General Research Funding
  • Dynamic Precision-Aware Low-Bit Training for Multimodal LLMs on Blackwell GPUs · NVIDIA Academic Grant Program Award

Industrial Experiences

Project Experiences

ATorch GitHub stars ATorch: Make large model training more efficient and reproducible for everyone

DLRover GitHub stars DLRover: An automatic distributed deep learning system

Vitis-AI GitHub stars Vitis-AI: An Integrated Development Environment that can be leveraged to accelerate AI inference on AMD adaptable platforms

Lookahead GitHub stars Lookahead: A toolkit to accelerate LLM inference without headache (🤯) and tears (😭) .

Awards

  • PolyU Research Postgraduate Scholarship, 2025
  • Outstanding Postgraduate Student of USTC, 2024
  • 1st-Class Academic Scholarship of USTC, 2021-2024
  • Merit Student Award of the Province, 2021
  • Annual Outstanding Student of the Province, Province Level · 2019
  • Merit Student Award, 2019
  • Outstanding Student Leader Award, 2021

Teaching

  • Teaching Assistant, COMP6713 - Advanced Large Language Model and Beyond, PolyU · 2026 Spring
  • Teaching Assistant, COMP2021 – Object-oriented Programming, PolyU · 2025 Fall
  • Teaching Assistant, ML4432 – Machine Learning, PolyU · 2025 Spring
  • Teaching Assistant, CONT010177 – Modern Control Theory, USTC · 2022 Fall

Skills

  • Programming: Python, C++, CUDA, Golang
  • Systems and Frameworks: Megatron-LM, DeepSpeed, vLLM, Ray, Triton
  • Expertise: Ultra-low-bit LLM training, FP8/FP4/NVFP4, quantization-aware training, distributed training systems, inference optimization, profiling and memory optimization

Extracurricular Activities

  • President of the University Student Union

Statistics   [back top]