As AI models continue to grow in size, they require vast amounts of energy, making sustainable AI unfeasible if current trends persist. Consequently, the importance of robust computing HW/SW infrastructure underpinning AI will become critical.

Our main research goal is to compute AI models in a faster and energy-efficient way through HW/SW co-design. Specifically, our research interests include:

  • Neural Processing Unit (NPU), domain-specific hardware, FPGA
  • Quantization, pruning, and knowledge distillation
  • Hardware-aware neural architecture search (HW-Aware NAS) and neural architecture accelerator search (NAAS)
  • Processing-in-memory (PIM)
  • Efficient LLM serving including KV Caching and other optimizations
  • On-Device AI

Latest Updates

  • Paper

    Our paper "CSP-KD: Clustering-Based Sentence Pruning for Multi-Teacher Knowledge Distillation" has been accepted at CCCI.

    2026-09-07

  • News

    Seungyeon Han has joined our group as a graduate student. Welcome!

    2026-09-01

  • News

    Jieui Kang has graduated. Congratulations!

    2026-08-28

  • Paper

    Our paper "RLE-Aware Weight Loader: A Lightweight Decompression Front-End for Fast LLM Cold Starts" has been accepted at ISOCC.

    2026-08-10

  • News

    Prof. Jaehyeong Sim has been promoted to Associate Professor, effective from the 2026-2 semester.

    2026-08-01

Recent Publications

  • RLE-Aware Weight Loader: A Lightweight Decompression Front-End for Fast LLM Cold Starts

    Jaeyoung Choi, Jaehyeong Sim

    ISOCCAccepted2026
  • CSP-KD: Clustering-Based Sentence Pruning for Multi-Teacher Knowledge Distillation

    Soeun Choi, Jieui Kang, Eunjoung Yoo, Yeonhee Kim, Jaehyeong Sim

    CCCIAccepted
  • Token-Based Task-Aware Knowledge Distillation for Encoder Adaptation

    Eunjoung Yoo, Jieui Kang, Soeun Choi, Yeonhee Kim, Jaehyeong Sim

    ACCESS2026
  • QubitCache: Quantum-Inspired Probabilistic Attention Preservation for KV-Cache Compression

    Jieui Kang, Jaeyoung Choi, Wonhui Roh, Jaehyeong Sim

    ACCESS2026
  • SHARP: Structured Hierarchical Attention Rank Projection for Efficient Language Model Distillation

    Jieui Kang, Eunjoung Yoo, Soeun Choi, Yeonhui Kim, Jaehyeong Sim

    ACCESS2026