Skip to content

All

    Repositories list

    • [BMVC 2026] DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering
      Python
      0300Updated Sep 3, 2026Sep 3, 2026
    • [EMNLP 2026] Official implementation of CounterVid, a counterfactual video generation framework for mitigating action and temporal hallucinations in video-langu…
      0400Updated Aug 28, 2026Aug 28, 2026
    • Official repository for ShieldCLIP: selective safety alignment for harmful content mitigation in multimodal foundation models. Code, trained models, and dataset…
      0200Updated Aug 24, 2026Aug 24, 2026
    • GramSR

      Public
      Official implementation of "GramSR: Visual Feature Conditioning for Diffusion-Based Super-Resolution"
      Python
      1910Updated Aug 12, 2026Aug 12, 2026
    • LoT

      Public
      Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering
      HTML
      0401Updated Aug 7, 2026Aug 7, 2026
    • ReAG

      Public
      [CVPR 2026 Highlight] ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
      Python
      02920Updated Jul 13, 2026Jul 13, 2026
    • General Federated Continual Learning Framework
      Python
      MIT License
      32110Updated Jul 11, 2026Jul 11, 2026
    • HyperMIL

      Public
      [ICPR 2026] HyperMIL: Hypergraph-based channel reasoning for Multiple Instance Learning on Multivariate Time Series
      Python
      0100Updated Jul 1, 2026Jul 1, 2026
    • Dress-ED

      Public
      [ECCV 2026] "Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off"
      HTML
      01410Updated Jul 1, 2026Jul 1, 2026
    • HeRA

      Public
      Mind the Heads: Topological Representation Alignment for Multimodal LLMs
      HTML
      0510Updated Jun 30, 2026Jun 30, 2026
    • sva2021

      Public
      0000Updated Jun 26, 2026Jun 26, 2026
    • [ECCV 2026] Official implementation of "A Scalable Vector Graphics Latent Space"
      HTML
      0200Updated Jun 23, 2026Jun 23, 2026
    • [Under Review] Official implementation of cross-model safety steering, a framework that transfers safety directions from LLMs to heterogeneous image and video g…
      0710Updated Jun 5, 2026Jun 5, 2026
    • MAs-DiT

      Public
      Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers
      0900Updated May 21, 2026May 21, 2026
    • mammoth

      Public
      An Extendible (General) Continual Learning Framework based on Pytorch - official codebase of Dark Experience for General Continual Learning
      Python
      MIT License
      15883710Updated May 20, 2026May 20, 2026
    • cvcs2026

      Public
      0000Updated May 18, 2026May 18, 2026
    • ScanDiff

      Public
      This is the official repository for the paper "Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction". ICCV 2025
      Python
      62710Updated May 13, 2026May 13, 2026
    • MissRAG

      Public
      [ICCV 2025] MissRAG: Addressing the Missing Modality Challenge in Multimodal Large Language Models
      Python
      Apache License 2.0
      02610Updated May 12, 2026May 12, 2026
    • coldfront

      Public
      HPC Resource Allocation System
      Python
      GNU General Public License v3.0
      112000Updated Apr 19, 2026Apr 19, 2026
    • VHS

      Public
      [CVPR2026 Findings] VHS: Verifier on Hidden States, an efficient inference-time scaling verification framework for DiT-based image generation.
      Python
      Other
      01600Updated Mar 25, 2026Mar 25, 2026
    • Eruku

      Public
      Python
      0100Updated Mar 19, 2026Mar 19, 2026
    • JARVIS

      Public
      Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
      Python
      0900Updated Mar 14, 2026Mar 14, 2026
    • JavaScript
      0000Updated Feb 24, 2026Feb 24, 2026
    • IDAttn

      Public
      Implementation of IDAttn
      MIT License
      0000Updated Feb 23, 2026Feb 23, 2026
    • BFS-PO

      Public
      0000Updated Feb 15, 2026Feb 15, 2026
    • HEaD

      Public
      0000Updated Jan 7, 2026Jan 7, 2026
    • ReT-2

      Public
      Recurrence Meets Transformers for Universal Multimodal Retrieval
      Python
      Apache License 2.0
      11510Updated Dec 15, 2025Dec 15, 2025
    • CHAIR-DPO

      Public
      [BMVC 2025] Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
      Python
      01000Updated Nov 29, 2025Nov 29, 2025
    • [IJCAI 2025] Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
      Python
      13701Updated Nov 25, 2025Nov 25, 2025
    • DICE

      Public
      [ICCV 2025] What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
      Python
      01600Updated Nov 3, 2025Nov 3, 2025
    ProTip! When viewing an organization's repositories, you can use the props. filter to filter by custom property.