Skip to content
Mingyu Lee

Mingyu Lee

Ph.D. Student in Electrical and Computer Engineering

Georgia Institute of Technology · Synergy Lab

Klaus Advanced Computing Building · Room 3306 · Atlanta, GA

I am a Ph.D. student at Georgia Tech, advised by Prof. Tushar Krishna. My research focuses on efficient AI systems through cross-layer hardware-software co-design.

I work across algorithms, systems, and hardware to make emerging models practical under real deployment constraints. My current interests include efficient generative and agentic AI, model compression, emerging data formats, and accelerator architecture.

Research Interests

My research asks how model structure, numerical representation, runtime systems, and hardware can be designed together rather than optimized in isolation.

Efficient Generative & Agentic AI

Efficient inference for diffusion language models, KV-cache optimization, and system support for emerging agentic workloads.

Polestar · Polestar-Cache

Model Compression

Data-free quantization, structured activation pruning, sparsity, and emerging low-precision data formats.

OuroMamba · RECAP

Hardware-Software Co-Design

GPU kernels, FPGA/ASIC accelerators, systolic arrays, RISC-V extensions, and memory-efficient architectures.

CUTLASS · FPGA · 65 nm ASIC

News

Started my Ph.D. in ECE at Georgia Tech 😊!

Released Polestar, our work on efficient diffusion LLM inference.

Presented MiniSA at ISPASS 2026 on behalf of Jianming.

Polestar-Cache appeared at the LIT Workshop at ICLR 2026.

Presented OuroMamba at ICCV 2025 in Honolulu.

Presented RECAP at MLArchSys, co-located with ISCA 2025.

Selected Publications

Explore by research area · select multiple areas to find work at their intersection. · * Equal contribution

Full list on Google Scholar

Showing all 4 publications

OuroMamba paper overview

OuroMamba: A Data-Free Quantization Framework for Vision Mamba Models

Akshat Ramachandran*, Mingyu Lee*, Huan Xu, Souvik Kundu, and Tushar Krishna

ICCVIEEE/CVF International Conference on Computer Vision2025Conference

Data-free calibration and dynamic mixed-precision quantization for Vision Mamba models, with up to 2.36× GPU speedup.

RECAP method overview

RECAP: Training-Free Compensation for Coarse Activation Channel Pruning in Compressed LLMs

Mingyu Lee*, Akshat Ramachandran*, and Tushar Krishna

MLArchSys @ ISCAWorkshop on ML for Computer Architecture and Systems2025Workshop

A lightweight compensation method for hardware-friendly channel pruning, improving BoolQ accuracy by 34% at 70% sparsity on LLaMA3-8B.

Polestar-Cache paper title and abstract

Polestar-Cache: Reconciling Parallel Decoding and Accuracy in Diffusion LLMs via Token Drift-Aware KV Cache Recalibration

Mingyu Lee*, Akshat Ramachandran*, Souvik Kundu, and Tushar Krishna

LIT @ ICLRLearning on Information and Tokens Workshop2026Workshop

A training-free, drift-aware KV-cache framework that selectively refreshes stale token representations during parallel diffusion decoding.

Polestar method overview

Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs

Mingyu Lee*, Akshat Ramachandran*, Souvik Kundu, and Tushar Krishna

arXivPreprint2026Preprint

A unified drift signal for selective KV-cache refresh and reliable token commitment, improving both throughput and generation accuracy.

Selected Projects

Implementation and hardware work, organized around my specific role and technical contribution.

65 nm Systolic GEMM Accelerator

Role

Designer · RTL to tapeout

TSMC 65 nm · SystemVerilog · Synopsys

My contribution

Designed and taped out a 5.56 TOPS/mW weight-stationary GEMM accelerator, including memory, external, and LFSR-based built-in self-test modes.

CXL-based LLM Accelerator

Role

FPGA Engineering Intern · HyperAccel

CXL · RISC-V · FPGA · FP8

My contribution

Extended the RISC-V ISA and designed adaptive processing elements for Qwen1.5-MoE, reducing model memory footprint by 3× through FP8 quantization.

Sparse ViT Accelerator

Role

Undergraduate Researcher · CAST Lab

FPGA · Sparse data formats

My contribution

Developed a compressed sparse diagonal format and its accelerator architecture, achieving 2× speedup and compression over conventional formats.

Talks & Presentations

Mingyu Lee presenting at ISPASS 2026

ISPASS 2026 · Seoul

Presented MiniSA at the main conference on behalf of Jianming.

Mingyu Lee and Akshat Ramachandran presenting OuroMamba at ICCV 2025

ICCV 2025 · Honolulu

Presented OuroMamba at the main conference poster session.

Mingyu Lee presenting RECAP at MLArchSys at ISCA 2025

MLArchSys @ ISCA 2025 · Tokyo

Oral and poster presentation of RECAP.

Education

2026 — Present

Georgia Institute of Technology

Ph.D. in Electrical and Computer Engineering

Advisor: Tushar Krishna

2019 — 2026

Korea Advanced Institute of Science and Technology

B.S. in Electrical Engineering

Experience

Synergy Lab · Graduate Research Assistant

2025 — Present

Georgia Tech

Synergy Lab · Undergraduate Researcher

2025 — 2026

Georgia Tech

FPGA Engineering Intern

2024

HyperAccel

Undergraduate Researcher

2024

CAST Lab, KAIST

Academic Service

2025 · Reviewer

MLArchSys @ ISCA

Workshop on ML for Computer Architecture and Systems.

2025 · Volunteer

CRNCH Summit

Center for Research into Novel Compute Hierarchies.

Honors

2026Georgia Tech ECE Fellowship

2025ICCV Broadening Participation Award

2025uArch Workshop / ISCA Grant

2025HPCA Student Travel Grant

2024Korea–U.S. Student Exchange Scholarship

Beyond Research

I recently started drawing manga, a hobby I hope to keep developing throughout my Ph.D. It has been a welcome way to slow down, observe details, and make something by hand.

I also love traveling. I enjoy making memories with friends, but I also like traveling alone with a camera. Recent trips include Tokyo, the Grand Canyon, and Yosemite in 2025, and Hiroshima in 2026.