I am Ziren Wang, a first-year Ph.D. student in Computer Science and Engineering at the University of California, San Diego, working with Prof. Yiying Zhang.
Previously, I studied Computer Science in the Yao Class at Tsinghua University, where I was advised by Prof. Mingyu Gao. I was also a research intern at the SysLab at the University of Washington, advised by Prof. Baris Kasikci.
My research interests center around distributed systems and machine learning systems, with a focus on building efficient and scalable infrastructure for modern AI workloads. My work includes high-performance LLM inference and fine-grained GPU resource management.
If you are interested in my research, please feel free to email me at ziw205@ucsd.edu.
Publications
-
NanoFlow: Towards Optimal Large Language Model Serving Throughput.
Kan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Tian Tang, Qinyu Xu, Zihao Ye, Keisuke Kamahori, Chien-Yu Lin, Ziren Wang, Stephanie Wang, Arvind Krishnamurthy, and Baris Kasikci
19th USENIX Symposium on Operating Systems Design and Implementation (OSDI 25), Year 2025. [PDF] [Code]
Projects
-
Flexible LLM Infrastructure [Code]
I am leading the development of a flexible, high-performance Python framework for LLM inference with fine-grained GPU resource management.
We model SMs, memory bandwidth, and PCIe transfers as separable resources on independent streams, enabling kernel co-scheduling and overlap to maximize GPU utilization while minimizing cross-kernel interference. -
Reproducing GOAL [PDF] [Code]
I reproduced GOAL, a routing algorithm for torus networks, implemented different virtual channel (VC) control policies, and evaluated their performance. GOAL combines randomized routing directions in each dimension for global load balancing with adaptive routing for local load balancing. -
Parallel Decomposition with RChol [PDF]
We developed a parallel decomposition algorithm that combines RChol sampling with a multifrontal method and dynamically manages dependencies between threads and nodes. Experiments show improved decomposition speed for matrices with substantial parallelism, but no speedup for matrices with limited parallelism.