Xiaoyou WuHomeBlogProjects

Research Projects

BlockBatch: Multi-Scale Consensus Decoding for Efficient dLLM Inference

BlockBatch: Multi-Scale Consensus Decoding for Efficient dLLM Inference

AcceptedFindings of EMNLP 2026

Accepted to Findings of EMNLP 2026. BlockBatch is a training-free inference framework for diffusion language models that reduces denoising NFEs by 26.6% and achieves a 1.33× average speedup over Fast-dLLM.

May 2026
EMNLP 2026Diffusion LLMNLPInference OptimizationMachine LearningResearch
View Project
Hy²: Accelerating Hybrid Mamba-Transformer Serving Through Adaptive GPU–PIM Co-Design

Hy²: Accelerating Hybrid Mamba-Transformer Serving Through Adaptive GPU–PIM Co-Design

A co-design framework for serving hybrid Mamba-Transformer models on GPU–PIM heterogeneous hardware, achieving 1.51× throughput over prior art and 2.47× over GPU-only baselines while reducing tail latency to 0.61×.

May 2026
GPU-PIMLLM ServingMambaHardware ArchitectureSystemsResearch
View Project
HiDeS: Hierarchical Delta Sparsity for Efficient Diffusion LLM Inference

HiDeS: Hierarchical Delta Sparsity for Efficient Diffusion LLM Inference

An algorithm–system co-design that exploits cross-step temporal redundancy at three independent granularities — layers, tokens, and columns — achieving 2.66× and 2.30× throughput over dense baselines on A100 and H100 with negligible accuracy loss.

May 2026
Diffusion LLMSparse ExecutionGPU KernelsInference OptimizationSystemsResearch
View Project

Past Projects

08selected works
8 projects across hardware and software

Robotics, embedded systems, computer vision, AI, and product engineering.

Explore the archive →