Paper list of Generative Learned Compression (GLC). This includes topics such as Generative Image Compression (GIC), Generative Video Compression (GVC), Extreme Learned Image Compression (ELIC), and Extreme Learned Video Compression (ELVC) for human and machine vision perception. ELIC and ELVC, in particular, have emerged from advancements in generative models and could offer new insights into the relationship between generation ability and Shannon’s information theory. We also provide the test conditions for fair comparison.
Maintained by: Lingyu Zhu , Langxi Huang and Chengyan Jiang
Generative Image Compression Report: Chinese version
Invited Talk: Chinese version
- If you find papers relevant to this topic, please share them as a discussion post.
- Some papers may simultaneously belong to multiple subfields, and we categorize them accordingly to reflect these overlaps.
- Looking forward to your kind contributions and discussions! Many thanks!
This round prioritizes papers from TPAMI, TMM, TCSVT, IJCV, CVPR, ICCV, ECCV and closely related high-quality venues.
For very recent work, we also include arXiv / project pages when the official venue version is not yet fully indexed or the code release is only announced there.
Table of Contents
Some works naturally overlap multiple categories, and we intentionally duplicate them for easier lookup.
| Publish Date | Title | Authors (First Author) | Code | |
|---|---|---|---|---|
| 2026.04 | CoD-Lite: Real-Time Diffusion-Based Generative Image Compression | Bin Li et al. | 2604.12525 | GitHub |
| 2026.03 | DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression | Junqi Shi et al. | 2603.13162 | null |
| 2026.03 | ProGIC: Progressive and Lightweight Generative Image Compression with Residual Vector Quantization | Hao Cao et al. | 2603.02897 | null |
| 2026.02 | One-Step Diffusion for Perceptual Image Compression | Yiwen Jia et al. | 2602.01570 | GitHub |
| 2025.06 | Single-step Diffusion for Image Compression at Ultra-Low Bitrates | Chanung Park et al. | 2506.16572 | null |
| 2025.05 | Semantics-Guided Generative Image Compression | Cheng-Lin Wu et al. | 2505.24015 | null |
| 2025.05 | Generative Image Compression by Estimating Gradients of the Rate-variable Feature Distribution | Minghao Han et al. | 2505.20984 | null |
| 2025.06 | Bridging the Gap between Gaussian Diffusion Models and Universal Quantization for Image Compression | Lucas Relic et al. | CVPR 2025 | null |
| 2025.06 | Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image Compression | Jinchang Xu et al. | CVPR 2025 | null |
| 2025.04 | Once-for-All: Controllable Generative Image Compression with Dynamic Granularity Adaptation | Anqi Li et al. | 2406.00758 | GitHub |
| 2024.09 | Lossy Image Compression with Foundation Diffusion Models | Lucas Relic et al. | 2404.08580 | null |
| 2024.02 | MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model | Chunyi Li et al. | 2402.16749 | GitHub |
| 2023.10 | Towards Image Compression with Perfect Realism at Ultra-Low Bitrates | Marlene Careil et al. | 2310.10325 | null |
| Publish Date | Title | Authors (First Author) | Code | |
|---|---|---|---|---|
| 2026.03 | Generative Video Compression with One-Dimensional Latent Representation | Zihan Zheng et al. | 2603.15302 | Project |
| 2026.03 | ProGVC: Progressive-based Generative Video Compression via Auto-Regressive Context Modeling | Daowen Li et al. | 2603.17546 | null |
| 2026.03 | Generative video compression: towards 0.01% compression rate for video transmission | Xiangyu Chen et al. | Springer PDF | null |
| 2026.01 | YODA: Yet Another One-step Diffusion-based Video Compressor | Xingchen Li et al. | 2601.01141 | GitHub |
| 2026.01 | DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression | Wenzhuo Ma et al. | 2601.20564 | null |
| 2025.10 | GIViC: Generative Implicit Video Compression | Ge Gao et al. | ICCV 2025 | Project |
| 2025.10 | Diffusion Autoencoders are Foundation Video Compressors | Niccolo Niccoli et al. | ICCVW 2025 | null |
| 2025.05 | Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression | Linfeng Qi et al. | 2505.16177 | GitHub |
| 2025.06 | Towards Practical Real-Time Neural Video Compression | Zhaoyang Jia et al. | CVPR 2025 | GitHub |
| 2019.10 | Video Compression With Rate-Distortion Autoencoders | Amirhossein Habibian et al. | ICCV 2019 | null |
| 2019.10 | Learned Video Compression | Oren Rippel et al. | ICCV 2019 | null |
| 2018.09 | Video Compression through Image Interpolation | Chao-Yuan Wu et al. | ECCV 2018 | null |
| Publish Date | Title | Authors (First Author) | Code | |
|---|---|---|---|---|
| 2026.04 | CoD-Lite: Real-Time Diffusion-Based Generative Image Compression | Bin Li et al. | 2604.12525 | GitHub |
| 2026.02 | One-Step Diffusion for Perceptual Image Compression | Yiwen Jia et al. | 2602.01570 | GitHub |
| 2025.10 | StableCodec: Taming One-Step Diffusion for Extreme Image Compression | Tianyu Zhang et al. | ICCV 2025 | null |
| 2025.10 | DLF: Extreme Image Compression with Dual-generative Latent Fusion | Naifu Xue et al. | ICCV 2025 | Project |
| 2025.06 | Single-step Diffusion for Image Compression at Ultra-Low Bitrates | Chanung Park et al. | 2506.16572 | null |
| 2025.05 | Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression | Linfeng Qi et al. | 2505.16177 | GitHub |
| 2025.01 | Toward Extreme Image Compression with Latent Feature Guidance and Diffusion Prior | Zhiyuan Li et al. | 2404.18820 | GitHub |
| 2024.06 | Generative Latent Coding for Ultra-Low Bitrate Image Compression | Zhaoyang Jia et al. | CVPR 2024 | GitHub |
| 2024.06 | Once-for-All: Controllable Generative Image Compression with Dynamic Granularity Adaptation | Anqi Li et al. | 2406.00758 | GitHub |
| 2024.02 | MISC: Ultra-low Bitrate Image Semantic Compression Driven by Large Multimodal Model | Chunyi Li et al. | 2402.16749 | GitHub |
| 2023.10 | Towards Image Compression with Perfect Realism at Ultra-Low Bitrates | Marlene Careil et al. | 2310.10325 | null |
| 2023.07 | Extreme Image Compression using Fine-tuned VQGANs | Qi Mao et al. | 2307.08265 | GitHub |
| Publish Date | Title | Authors (First Author) | Code | |
|---|---|---|---|---|
| 2026.03 | Generative video compression: towards 0.01% compression rate for video transmission | Xiangyu Chen et al. | Springer PDF | null |
| 2026.03 | Generative Video Compression with One-Dimensional Latent Representation | Zihan Zheng et al. | 2603.15302 | Project |
| 2026.03 | ProGVC: Progressive-based Generative Video Compression via Auto-Regressive Context Modeling | Daowen Li et al. | 2603.17546 | null |
| 2026.01 | DiffVC-RT: Towards Practical Real-Time Diffusion-based Perceptual Neural Video Compression | Wenzhuo Ma et al. | 2601.20564 | null |
| 2026.01 | YODA: Yet Another One-step Diffusion-based Video Compressor | Xingchen Li et al. | 2601.01141 | GitHub |
| 2025.10 | Diffusion Autoencoders are Foundation Video Compressors | Niccolo Niccoli et al. | ICCVW 2025 | null |
| 2025.10 | GIViC: Generative Implicit Video Compression | Ge Gao et al. | ICCV 2025 | Project |
| 2025.05 | Generative Latent Coding for Ultra-Low Bitrate Image and Video Compression | Linfeng Qi et al. | 2505.16177 | GitHub |
| Publish Date | Title | Authors (First Author) | Code | |
|---|---|---|---|---|
| 2023.10 | EvalCrafter: Benchmarking and Evaluating Large Video Generation Models | Yaofang Liu et al. | 2310.11440 | GitHub |
| 2023.07 | AIGCIQA2023: A Large-scale Image Quality Assessment Database for AI Generated Images: from the Perspectives of Quality, Authenticity and Correspondence | Jiarui Wang et al. | 2307.00211 | GitHub |
| 2023.06 | AGIQA-3K: An Open Database for AI-Generated Image Quality Assessment | Chunyi Li et al. | 2306.04717 | GitHub |
| 2019.10 | KonIQ-10k: An ecologically valid database for deep learning of blind image quality assessment | Vlad Hosu et al. | 1910.06180 | Dataset |
| 2019.06 | KADID-10k: A Large-scale Artificially Distorted IQA Database | Hanhe Lin et al. | QoMEX 2019 PDF | Dataset |
| 2016.09 | MCL-JCV: A JND-based H.264/AVC Video Quality Assessment Dataset | Haiqiang Wang et al. | ICIP 2016 | Dataset |
| 2015.01 | Image database TID2013: Peculiarities, results and perspectives | Nikolay Ponomarenko et al. | Open PDF | Dataset |
| 2013.06 | Color image database TID2013: Peculiarities and preliminary results | Nikolay Ponomarenko et al. | EUVIP 2013 PDF | Dataset |
| Publish Date | Title | Authors (First Author) | Code | |
|---|---|---|---|---|
| 2025.03 | Image Quality Assessment: From Human to Machine Preference | Chunyi Li et al. | 2503.10078 | GitHub |
| 2017.06 | Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering | Yash Goyal et al. | CVPR 2017 | Dataset |
| 2017.07 | Scene Parsing through ADE20K Dataset | Bolei Zhou et al. | 1608.05442 | GitHub |
| 2017.01 | Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations | Ranjay Krishna et al. | 1602.07332 | Dataset |
| 2016.04 | The Cityscapes Dataset for Semantic Urban Scene Understanding | Marius Cordts et al. | 1604.01685 | Dataset |
| 2015.12 | ImageNet Large Scale Visual Recognition Challenge | Olga Russakovsky et al. | 1409.0575 | Dataset |
| 2018.11 | Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale | Alina Kuznetsova et al. | 1811.00982 | Dataset |
| 2014.05 | Microsoft COCO: Common Objects in Context | Tsung-Yi Lin et al. | 1405.0312 | Dataset |
- Image benchmark recommendation:
Kodak,CLIC2020/CLIC2021,DIV2K, and optionallyMS-COCO-30Kfor human preference or generative evaluation. - Video benchmark recommendation:
UVG,HEVC Class B/C/E,MCL-JCV, and clearly statelow-delayorrandom-accesssettings. - Report bitrate consistently: use
bppfor images andbpp / kbpsfor videos; specify whether entropy coding is included. - Report color/setup explicitly: resolution, crop policy, RGB vs. YUV,
4:4:4vs.4:2:0, GOP size, intra period, and frame count. - Report both fidelity and perceptual metrics:
PSNR,MS-SSIM,LPIPS,DISTS,FID,KID, and for video alsoFVDif applicable. - Human/machine dual evaluation: when the method targets joint perception, additionally report downstream task performance on
COCO,Cityscapes,ADE20K,VQA v2, or task-specific benchmarks. - Runtime reporting: GPU/CPU model, precision (
fp32/fp16/bf16), encoding speed, decoding speed, and whether the numbers include entropy coding.
