🌟A collection of papers, datasets, benchmarks, code, and pre-trained weights for Remote Sensing Foundation Models (RSFMs).
🔥🔥🔥 Last Updated on 2026.05.06 🔥🔥🔥
- Models
- Datasets & Benchmarks
- Others
| Abbreviation | Title | Publication | Paper | Code & Weights |
|---|---|---|---|---|
| GeoKR | Geographical Knowledge-Driven Representation Learning for Remote Sensing Images | TGRS2021 | GeoKR | link |
| - | Self-Supervised Learning of Remote Sensing Scene Representations Using Contrastive Multiview Coding | CVPRW2021 | Paper | link |
| GASSL | Geography-Aware Self-Supervised Learning | ICCV2021 | GASSL | link |
| SeCo | Seasonal Contrast: Unsupervised Pre-Training From Uncurated Remote Sensing Data | ICCV2021 | SeCo | link |
| RSP | An Empirical Study of Remote Sensing Pretraining | TGRS2022 | RSP | link |
| MATTER | Self-Supervised Material and Texture Representation Learning for Remote Sensing Tasks | CVPR2022 | MATTER | link |
| - | Self-supervised Vision Transformers for Land-cover Segmentation and Classification | CVPRW2022 | Paper | link |
| DINO-MM | Self-Supervised Vision Transformers for Joint SAR-Optical Representation Learning | IGARSS2022 | DINO-MM | link |
| GeCo | Geographical Supervision Correction for Remote Sensing Representation Learning | TGRS2022 | GeCo | link |
| RingMo | RingMo: A remote sensing foundation model with masked image modeling | TGRS2022 | RingMo | Code |
| RS-BYOL | Self-Supervised Learning for Invariant Representations From Multi-Spectral and SAR Images | JSTARS2022 | RS-BYOL | link |
| CSPT | Consecutive Pre-Training: A Knowledge Transfer Learning Strategy with Relevant Unlabeled Data for Remote Sensing Domain | RS2022 | CSPT | link |
| RVSA | Advancing plain vision transformer toward remote sensing foundation model | TGRS2022 | RVSA | link |
| SatMAE | SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery | NeurIPS2022 | SatMAE | link |
| DINO-MC | Extending Global-Local View Alignment for Self-Supervised Learning with Remote Sensing Imagery | Arxiv2023 | DINO-MC | link |
| CMID | CMID: A Unified Self-Supervised Learning Framework for Remote Sensing Image Understanding | TGRS2023 | CMID | link |
| Presto | Lightweight, Pre-trained Transformers for Remote Sensing Timeseries | Arxiv2023 | Presto | link |
| AST | AST: Adaptive Self-supervised Transformer for Optical Remote Sensing Representation | ISPRS JPRS2023 | AST | null |
| TOV | TOV: The original vision model for optical remote sensing image understanding via self-supervised learning | JSTARS2023 | TOV | link |
| CACo | Change-Aware Sampling and Contrastive Learning for Satellite Images | CVPR2023 | CACo | link |
| IaI-SimCLR | Multi-Modal Multi-Objective Contrastive Learning for Sentinel-1/2 Imagery | CVPRW2023 | IaI-SimCLR | null |
| - | A Self-Supervised Cross-Modal Remote Sensing Foundation Model with Multi-Domain Representation and Cross-Domain Fusion | IGARSS2023 | Paper | null |
| SatLas | SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image Understanding | ICCV2023 | SatLas | link |
| GFM | Towards Geospatial Foundation Models via Continual Pretraining | ICCV2023 | GFM | link |
| Scale-MAE | Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation Learning | ICCV2023 | Scale-MAE | link |
| Prithvi | Foundation Models for Generalist Geospatial Artificial Intelligence | Arxiv2023 | Prithvi | link |
| RingMo-Sense | RingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution Disentangling | TGRS2023 | RingMo-Sense | null |
| EarthPT | EarthPT: a time series foundation model for Earth Observation | NeurIPS2023 CCAI workshop | EarthPT | link |
| CROMA | CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders | NeurIPS2023 | CROMA | link |
| Cross-Scale MAE | Cross-Scale MAE: A Tale of Multiscale Exploitation in Remote Sensing | NeurIPS2023 | Cross-Scale MAE | link |
| USat | USat: A Unified Self-Supervised Encoder for Multi-Sensor Satellite Imagery | Arxiv2023 | USat | link |
| AIEarth | Analytical Insight of Earth: A Cloud-Platform of Intelligent Computing for Geospatial Big Data | Arxiv2023 | AIEarth | link |
| GeRSP | Generic Knowledge Boosted Pretraining for Remote Sensing Images | TGRS2024 | GeRSP | GeRSP |
| SMLFR | Generative ConvNet Foundation Model With Sparse Modeling and Low-Frequency Reconstruction for Remote Sensing Image Interpretation | TGRS2024 | SMLFR | link |
| RingMo-lite | RingMo-Lite: A Remote Sensing Lightweight Network With CNN-Transformer Hybrid Framework | IEEE TGRS2024 | RingMo-lite | null |
| U-BARN | Self-Supervised Spatio-Temporal Representation Learning of Satellite Image Time Series | JSTARS2024 | Paper | link |
| SpectralGPT | SpectralGPT: Spectral Remote Sensing Foundation Model | TPAMI2024 | SpectralGPT | link |
| SwiMDiff | SwiMDiff: Scene-Wide Matching Contrastive Learning With Diffusion Constraint for Remote Sensing Image | TGRS2024 | SwiMDiff | null |
| DOFA | Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation | Arxiv2024 | DOFA | link |
| - | Masked Feature Modeling for Generative Self-Supervised Representation Learning of High-Resolution Remote Sensing Images | IEEE JSTARS2024 | Paper | null |
| BFM | A Billion-scale Foundation Model for Remote Sensing Images | IEEE JSTARS2024 | BFM | null |
| Clay | Clay Foundation Model | Arxiv2024 | null | link |
| Hydro | Hydro--A Foundation Model for Water in Satellite Imagery | Arxiv2024 | null | link |
| S2MAE | S2MAE: A Spatial-Spectral Pretraining Foundation Model for Spectral Remote Sensing Data | CVPR2024 | S2MAE | null |
| SatMAE++ | Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery | CVPR2024 | SatMAE++ | link |
| msGFM | Bridging Remote Sensors with Multisensor Geospatial Foundation Models | CVPR2024 | msGFM | link |
| SkySense | SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery | CVPR2024 | SkySense | link |
| MTP | MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining | IEEE JSTARS2024 | MTP | link |
| RS-DFM | RS-DFM: A Remote Sensing Distributed Foundation Model for Diverse Downstream Tasks | Arxiv2024 | RS-DFM | null |
| OFA-Net | One for All: Toward Unified Foundation Models for Earth Vision | IGARSS2024 | OFA-Net | null |
| - | Lightweight and Efficient: A Family of Multimodal Earth Observation Foundation Models | IGARSS2024 | Paper | null |
| MM-VSF | Towards Knowledge Guided Pretraining Approaches for Multimodal Foundation Models: Applications in Remote Sensing | Arxiv2024 | MM-VSF | null |
| LeMeViT | LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation | IJCAI2024 | LeMeViT | link |
| SAR-JEPA | Predicting Gradient is Better: Exploring Self-Supervised Learning for SAR ATR with a Joint-Embedding Predictive Architecture | ISPRS JPRS2024 | SAR-JEPA | link |
| DeCUR | DeCUR: decoupling common & unique representations for multimodal self-supervision | ECCV2024 | DeCUR | link |
| MMEarth | MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning | ECCV2024 | MMEarth | link |
| OmniSat | OmniSat: Self-Supervised Modality Fusion for Earth Observation | ECCV2024 | OmniSat | link |
| MA3E | Masked Angle-Aware Autoencoder for Remote Sensing Images | ECCV2024 | MA3E | link |
| SoftCon | Multi-Label Guided Soft Contrastive Learning for Efficient Earth Observation Pretraining | TGRS2024 | SoftCon | link |
| PIS | Pretrain a Remote Sensing Foundation Model by Promoting Intra-instance Similarity | TGRS2024 | PIS | link |
| FG-MAE | Feature Guided Masked Autoencoder for Self-Supervised Learning in Remote Sensing | IEEE JSTARS2024 | FG-MAE | link |
| - | A Multimodal Unified Representation Learning Framework With Masked Image Modeling for Remote Sensing Images | IEEE TGRS2024 | Paper | null |
| OReole-FM | OReole-FM: successes and challenges toward billion-parameter foundation models for high-resolution satellite imagery | SIGSPATIAL2024 | OReole-FM | null |
| SatVision-TOA | SatVision-TOA: A Geospatial Foundation Model for Coarse-Resolution All-Sky Remote Sensing Imagery | Arxiv2024 | SatVision-TOA | link |
| Prithvi-EO-2.0 | Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications | Arxiv2024 | Prithvi-EO-2.0 | link |
| A2-MAE | A2-MAE: A spatial-temporal-spectral unified remote sensing pre-training method based on anchor-aware masked autoencoder | IEEE TGRS2025 | A2-MAE | null |
| FoMo | FoMo: Multi-Modal, Multi-Scale and Multi-Task Remote Sensing Foundation Models for Forest Monitoring | AAAI2025 | FoMo | link |
| PIEViT | Pattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote Sensing | IEEE TGRS2025 | PIEViT | null |
| SARATR-X | SARATR-X: Toward Building a Foundation Model for SAR Target Recognition | IEEE TIP2025 | SARATR-X | link |
| SatMamba | SatMamba: Development of Foundation Models for Remote Sensing Imagery Using State Space Models | Arxiv2025 | SatMamba | link |
| DynamicVis | DynamicVis: Dynamic Visual Perception for Efficient Remote Sensing Foundation Models | Arxiv2025 | DynamicVis | link |
| SUMMIT | SUMMIT: A SAR foundation model with multiple auxiliary tasks enhanced intrinsic characteristics | IJAEO2025 | SUMMIT | null |
| SenPa-MAE | SenPa-MAE: Sensor Parameter Aware Masked Autoencoder for Multi-Satellite Self-Supervised Pretraining | LNCS2025 | SenPa-MAE | link |
| HyperSIGMA | HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model | IEEE TPAMI2025 | HyperSIGMA | link |
| HyperSL | HyperSL: A Spectral Foundation Model for Hyperspectral Image Interpretation | IEEE TGRS2025 | HyperSL | link |
| TiMo | TiMo: Spatiotemporal Foundation Model for Satellite Image Time Series | Arxiv2025 | TiMo | link |
| Panopticon | Panopticon: Advancing Any-Sensor Foundation Models for Earth Observation | CVPRW2025 (EarthVision Best Paper) | Panopticon | link |
| HyperFree | HyperFree: A Channel-adaptive and Tuning-free Foundation Model for Hyperspectral Remote Sensing Imagery | CVPR2025 | HyperFree | link |
| SpectralEarth | SpectralEarth: Training Hyperspectral Foundation Models at Scale | IEEE JSTARS2025 | SpectralEarth | link |
| WV-Net | WV-Net: A foundation model for SAR WV-mode satellite imagery trained using contrastive self-supervised learning on 10 million images | AIES2025 | WV-Net | null |
| AnySat | AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities | CVPR2025 | AnySat | link |
| TerraFM | TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation | Arxiv2025 | TerraFM | link |
| Galileo | Galileo: Learning Global & Local Features of Many Remote Sensing Modalities | ICML2025 TerraBytes Workshop | Galileo | link |
| DeepAndes | DeepAndes: A Self-Supervised Vision Foundation Model for Multispectral Remote Sensing Imagery of the Andes | IEEE JSTARS2025 | DeepAndes | link |
| RingMamba | RingMamba: Remote Sensing Multisensor Pretraining With Visual State Space Model | IEEE TGRS2025 | RingMamba | link |
| RingMo-Aerial | RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning | IEEE TPAMI2025 | RingMo-Aerial | null |
| SeaMo | SeaMo: A Multi-Seasonal and Multimodal Remote Sensing Foundation Model | Information Fusion2025 | SeaMo | null |
| MoSAiC | MoSAiC: Multi-Modal Multi-Label Supervision-Aware Contrastive Learning for Remote Sensing | IEEE Sensors Journal 2025 | MoSAiC | null |
| CGEarthEye | CGEarthEye: A High-Resolution Remote Sensing Vision Foundation Model Based on the Jilin-1 Satellite Constellation | Arxiv2025 | CGEarthEye | null |
| AlphaEarth | AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data | Arxiv2025 | AlphaEarth | link |
| SkySense++ | A semantic-enhanced multi-modal remote sensing foundation model for Earth observation | Nature Machine Intelligence 2025 | SkySense++ | link |
| CtxMIM | CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding | ACM TOMM2025 | CtxMIM | null |
| ViTP | Visual Instruction Pretraining for Domain-Specific Foundation Models | Arxiv2025 | ViTP | link |
| SatDiFuser | Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models? | ICCV2025 | SatDiFuser | link |
| FedSense | Towards Privacy-preserved Pre-training of Remote Sensing Foundation Models with Federated Mutual-guidance Learning | ICCV2025 | FedSense | null |
| Copernicus-FM | Towards a Unified Copernicus Foundation Model for Earth Vision | ICCV2025 | Copernicus-FM | link |
| TerraMind | TerraMind: Large-Scale Generative Multimodality for Earth Observation | ICCV2025 | TerraMind | link |
| SelectiveMAE | Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling | ICCV2025 | SelectiveMAE | link |
| SMARTIES | SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images | ICCV2025 | SMARTIES | link |
| SkySense V2 | SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing | ICCV2025 | SkySense V2 | null |
| RS-vHeat | RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation Model | ICCV2025 | RS-vHeat | null |
| RoMA | RoMA: Scaling up Mamba-based Foundation Models for Remote Sensing | NeurIPS2025 | RoMA | link |
| GeoLink | GeoLink: Empowering Remote Sensing Foundation Model with OpenStreetMap Data | NeurIPS2025 | GeoLink | link |
| CrossEarth | CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation | IEEE TPAMI2025 | CrossEarth | link |
| PhySwin | PhySwin: An Efficient and Physically-Informed Foundation Model for Multispectral Earth Observation | NeurIPS2025 | PhySwin | null |
| FlexiMo | FlexiMo: A Flexible Remote Sensing Foundation Model | IEEE TGRS2026 | FlexiMo | null |
| MAPEX | MAPEX: Modality-Aware Pruning of Experts for Remote Sensing Foundation Models | IEEE TGRS2026 | MAPEX | link |
| - | A Complex-Valued SAR Foundation Model Based on Physically Inspired Representation Learning | IEEE TIP2026 | Paper | null |
| MAESTRO | MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data | WACV2026 | MAESTRO | link |
| RingMoE | RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation | IEEE TPAMI2026 | RingMoE | null |
| Alliance | Alliance: All-in-One Spectral-Spatial-Frequency Awareness Foundation Model | IEEE TPAMI2026 | Alliance | null |
| THOR | THOR: A Versatile Foundation Model for Earth Observation Climate and Society Applications | Arxiv2026 | THOR | link |
| AgriFM | AgriFM: A multi-source temporal remote sensing foundation model for Agriculture mapping | RSE2026 | Paper | link |
| SIGMAE | SIGMAE: A Spectral-Index-Guided Foundation Model for Multispectral Remote Sensing | Arxiv2026 | SIGMAE | link |
| CrossEarth-SAR | CrossEarth-SAR: A SAR-Centric and Billion-Scale Geospatial Foundation Model for Domain Generalizable Semantic Segmentation | Arxiv2026 | CrossEarth-SAR | link |
| NeighborMAE | NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining | CVPR2026 | NeighborMAE | null |
| MOMO | MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications | CVPR2026 | MOMO | link |
| TESSERA | TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis | CVPR2026 | TESSERA | link |
| OlmoEarth | OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation | CVPR2026 | OlmoEarth | link |
| RAMEN | RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation | CVPR2026 | RAMEN | link |
| SARMAE | SARMAE: Masked Autoencoder for SAR Representation Learning | CVPR2026 | SARMAE | link |
| Abbreviation | Title | Publication | Paper | Code & Weights |
|---|---|---|---|---|
| - | Charting New Territories: Exploring the Geographic and Geospatial Capabilities of Multimodal LLMs | Arxiv2023 | Paper | link |
| - | Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models | Arxiv2024 | Paper | link |
| SkyCLIP | SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing | AAAI2024 | SkyCLIP | link |
| GeoRSCLIP | RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing | IEEE TGRS2024 | GeoRSCLIP | link |
| RemoteCLIP | RemoteCLIP: A Vision Language Foundation Model for Remote Sensing | IEEE TGRS2024 | RemoteCLIP | link |
| EarthGPT | EarthGPT: A Universal Multimodal Large Language Model for Multisensor Image Comprehension in Remote Sensing Domain | IEEE TGRS2024 | EarthGPT | link |
| GRAFT | Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment | ICLR2024 | GRAFT | null |
| EarthMarker | EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing | IEEE TGRS2024 | EarthMarker | link |
| RingMoGPT | RingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and Grounded Tasks | IEEE TGRS2024 | RingMoGPT | null |
| RS-LLaVA | RS-LLaVA: Large Vision Language Model for Joint Captioning and Question Answering in Remote Sensing Imagery | RS2024 | RS-LLaVA | link |
| GeoChat | GeoChat: Grounded Large Vision-Language Model for Remote Sensing | CVPR2024 | GeoChat | link |
| SkySenseGPT | SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding | Arxiv2024 | SkySenseGPT | link |
| RSCLIP | Pushing the Limits of Vision-Language Models in Remote Sensing without Human Annotations | Arxiv2024 | RSCLIP | null |
| GeoText | Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching | ECCV2024 | GeoText | link |
| LHRS-Bot | LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model | ECCV2024 | LHRS-Bot | link |
| GeoGround | GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding | Arxiv2024 | GeoGround | link |
| RSUniVLM | RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts | Arxiv2024 | RSUniVLM | null |
| REO-VLM | REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation | Arxiv2024 | REO-VLM | null |
| UniRS | UniRS: Unifying Multi-temporal Remote Sensing Tasks through Vision Language Models | Arxiv2024 | UniRS | null |
| VHM | VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis | AAAI2025 | VHM | link |
| - | Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models | Arxiv2025 | Paper | null |
| DOFA-CLIP | DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation | Arxiv2025 | DOFA-CLIP | link |
| Falcon | Falcon: A Remote Sensing Vision-Language Foundation Model | Arxiv2025 | Falcon | link |
| GeoRSMLLM | GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing | Arxiv2025 | GeoRSMLLM | null |
| OmniGeo | OmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial Intelligence | Arxiv2025 | OmniGeo | null |
| DGTRS-CLIP | DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment | Arxiv2025 | DGTRS-CLIP | link |
| EagleVision | EagleVision: Object-level Attribute Multimodal LLM for Remote Sensing | Arxiv2025 | EagleVision | link |
| SkyEyeGPT | SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model | ISPRS JPRS2025 | SkyEyeGPT | link |
| SegEarth-R1 | SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model | Arxiv2025 | SegEarth-R1 | link |
| EarthGPT-X | EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting | Arxiv2025 | EarthGPT-X | link |
| TEOChat | TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data | ICLR2025 | TEOChat | link |
| LISAt | LISAt: Language-Instructed Segmentation Assistant for Satellite Imagery | Arxiv2025 | LISAt | null |
| DynamicVL | DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding | Arxiv2025 | DynamicVL | null |
| RSGPT | RSGPT: A Remote Sensing Vision Language Model and Benchmark | ISPRS JPRS2025 | RSGPT | link |
| LHRS-Bot-Nova | LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation | ISPRS JPRS2025 | LHRS-Bot-Nova | link |
| EarthDial | EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues | CVPR2025 | EarthDial | link |
| SkySense-O | SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling | CVPR2025 | SkySense-O | link |
| XLRS-Bench | XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery? | CVPR2025 | XLRS-Bench | null |
| GeoPix | GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing | IEEE GRSM2025 | GeoPix | link |
| Co-LLaVA | Co-LLaVA: Efficient Remote Sensing Visual Question Answering via Model Collaboration | RS2025 | Co-LLaVA | null |
| RLita | RLita: A Region-Level Image-Text Alignment Method for Remote Sensing Foundation Model | RS2025 | RLita | null |
| EarthMind | EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM | Arxiv2025 | EarthMind | link |
| - | Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling | Arxiv2025 | Paper | null |
| GeoPixel | GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing | ICML2025 | GeoPixel | link |
| RingMo-Agent | RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning | Arxiv2025 | RingMo-Agent | null |
| Geo-R1 | Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning | Arxiv2025 | Geo-R1 | null |
| RSThinker | Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models | Arxiv2025 | RSThinker | null |
| FUSAR-KLIP | FUSAR-KLIP: Towards Multimodal Foundation Models for Remote Sensing | Arxiv2025 | FUSAR-KLIP | link |
| GeoMag | GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing | ACMMM2025 | GeoMag | null |
| RemoteSAM | RemoteSAM: Towards Segment Anything for Earth Observation | ACMMM2025 | RemoteSAM | link |
| LRS-VQA | When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning | ICCV2025 | LRS-VQA | link |
| UrbanLLaVA | UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding | ICCV2025 | UrbanLLaVA | link |
| SARCLIP | SARCLIP: a multimodal foundation framework for SAR imagery via contrastive language-image pre-training | ISPRS JPRS2025 | SARCLIP | link |
| FUSE-RSVLM | FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing | Arxiv2025 | FUSE-RSVLM | link |
| Aquila | Aquila: A Hierarchically Aligned Vision-Language Model for Enhanced Remote Sensing Image Comprehension | Journal of Remote Sensing 2026 | Aquila | null |
| RSCoVLM | Co-Training Vision-Language Models for Remote Sensing Multi-Task Learning | RS2026 | RSCoVLM | link |
| GeoReason | GeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement Learning | Arxiv2026 | GeoReason | link |
| RemoteReasoner | RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow | AAAI2026 | RemoteReasoner | null |
| SkyMoE | SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts | AAAI2026 | SkyMoE | null |
| GeoEyes | GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery | Arxiv2026 | GeoEyes | link |
| GeoSolver | GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision | Arxiv2026 | GeoSolver | null |
| GeoAlignCLIP | GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning | Arxiv2026 | GeoAlignCLIP | null |
| RS-WorldModel | RS-WorldModel: a Unified Model for Remote Sensing Understanding and Future Sense Forecasting | Arxiv2026 | RS-WorldModel | null |
| Decoding-the-Delta | Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models | Arxiv2026 | Paper | null |
| RemoteShield | RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation | Arxiv2026 | RemoteShield | null |
| UniChange | UniChange: Unifying Change Detection with Multimodal Large Language Model | CVPR2026 | UniChange | link |
| ZoomEarth | ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks | CVPR2026 | ZoomEarth | null |
| UniGeoSeg | UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes | CVPR2026 | UniGeoSeg | link |
| SegEarth-R2 | SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images | CVPR2026 | SegEarth-R2 | link |
| FUSAR-GPT | FUSAR-GPT: A Spatiotemporal Feature-Embedded and Two-Stage Decoupled Visual Language Model for SAR Imagery | CVPR2026 | FUSAR-GPT | null |
| SATtxt | Spectrally Distilled Representations Aligned with Instruction-Augmented LLMs for Satellite Imagery | CVPR2026 | SATtxt | link |
| TerraScope | TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation | CVPR2026 | TerraScope | null |
| Abbreviation | Title | Publication | Paper | Code & Weights |
|---|---|---|---|---|
| GeoRSSD | RS5M: A Large Scale Vision-Language Dataset for Remote Sensing Vision-Language Foundation Model | Arxiv2023 | Paper | link |
| - | Generate Your Own Scotland: Satellite Image Generation Conditioned on Maps | NeurIPSW2023 | Paper | link |
| CRS-Diff | CRS-Diff: Controllable Remote Sensing Image Generation with Diffusion Model | Arxiv2024 | Paper | link |
| DiffusionSat | DiffusionSat: A Generative Foundation Model for Satellite Imagery | ICLR2024 | DiffusionSat | link |
| HSIGene | HSIGene: A Foundation Model For Hyperspectral Image Generation | Arxiv2024 | Paper | link |
| Text2Earth | Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model | GRSM2025 | Text2Earth | link |
| MetaEarth | MetaEarth: A Generative Foundation Model for Global-Scale Remote Sensing Image Generation | TPAMI2025 | MetaEarth | link |
| EcoMapper | EcoMapper: Generative Modeling for Climate-Aware Satellite Imagery | ICML2025 | EcoMapper | link |
| OSMGen | OSMGen: Highly Controllable Satellite Image Synthesis using OpenStreetMap Data | NeurIPSW2025 | OSMGen | link |
| Any2RSI | Any2RSI: Controllable Remote Sensing Text-to-Image Generation via Any Control and Enriched Description | AAAI2026 | Any2RSI | link |
| GeoDiT | GeoDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis | Arxiv2026 | GeoDiT | null |
| MetaEarth3D | MetaEarth3D: Unlocking World-scale 3D Generation with Spatially Scalable Generative Modeling | Arxiv2026 | MetaEarth3D | null |
| Abbreviation | Title | Publication | Paper | Code & Weights |
|---|---|---|---|---|
| CSP | CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual Representations | ICML2023 | CSP | link |
| GeoCLIP | GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localization | NeurIPS2023 | GeoCLIP | link |
| SatCLIP | SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery | AAAI2025 | SatCLIP | link |
| GAIR | GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations | Arxiv2025 | GAIR | link |
| RANGE | RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings | CVPR2025 | RANGE | null |
| Geo² | Geo²: Geometry-Guided Cross-view Geo-Localization and Image Synthesis | CVPR2026 | Geo² | null |
| GeoBridge | GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization | CVPR2026 | GeoBridge | link |
| Abbreviation | Title | Publication | Paper | Code & Weights |
|---|---|---|---|---|
| - | Self-supervised audiovisual representation learning for remote sensing data | JAG2022 | Paper | link |
| GeoBind | GeoBind: Binding Text, Image, and Audio through Satellite Images | IGARSS2024 | GeoBind | null |
| PSM | PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping | ACM MM 2024 | PSM | link |
| Sat2Sound | Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping | Arxiv2025 | Sat2Sound | null |
| Abbreviation | Title | Publication | Paper | Code & Weights |
|---|---|---|---|---|
| GeoLLM-QA | Evaluating Tool-Augmented Agents in Remote Sensing Platforms | ICLR 2024 ML4RS Workshop | Paper | null |
| GeoLLM-Engine | GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots. | CVPRW2024 | Paper | null |
| RS-Agent | RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent | Arxiv2024 | Paper | null |
| Change-Agent | Change-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and Analysis | TGRS2024 | Paper | link |
| GeoLLM-Squad | Multi-Agent Geospatial Copilots for Remote Sensing Workflows | Arxiv2025 | Paper | null |
| - | Towards LLM Agents for Earth Observation: The UnivEARTH Dataset | Arxiv2025 | Paper | null |
| AirSpatialBot | AirSpatialBot: A Spatially Aware Aerial Agent for Fine-Grained Vehicle Attribute Recognition and Retrieval | IEEE TGRS2025 | Paper | link |
| ThinkGeo | ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks | Arxiv2025 | Paper | link |
| PEACE | PEACE: Empowering Geologic Map Holistic Understanding with MLLMs | CVPR2025 | Paper | link |
| Geo-OLM | Geo-OLM: Enabling Sustainable Earth Observation Studies with Cost-Efficient Open Language Models & State-Driven Workflows | COMPASS'2025 | Paper | link |
| REMSA | REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent | Arxiv2025 | Paper | link |
| OpenEarthAgent | OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents | Arxiv2026 | Paper | link |
| OpenEarth-Agent | OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation | Arxiv2026 | Paper | link |
| Earth-Agent | Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents | ICLR2026 | Paper | link |
| GeoMMAgent | GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing | CVPR2026 | Paper | null |
| IMAIA | IMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial Intelligence | CVPR2026 | Paper | null |
| Abbreviation | Title | Publication | Paper | Link | Downstream Tasks |
|---|---|---|---|---|---|
| - | Revisiting pre-trained remote sensing model benchmarks: resizing and normalization matters | Arxiv2023 | Paper | link | Classification |
| GEO-Bench | GEO-Bench: Toward Foundation Models for Earth Monitoring | Arxiv2023 | Paper | link | Classification & Segmentation |
| FoMo-Bench | FoMo: Multi-Modal, Multi-Scale and Multi-Task Remote Sensing Foundation Models for Forest Monitoring | Arxiv2023 | FoMo-Bench | Coming soon | Classification & Segmentation & Detection for forest monitoring |
| PhilEO | PhilEO Bench: Evaluating Geo-Spatial Foundation Models | Arxiv2024 | Paper | link | Segmentation & Regression estimation |
| SkySense | SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery | CVPR2024 | SkySense | Targeted open-source | Classification & Segmentation & Detection & Change detection & Multi-Modal Segmentation: Time-insensitive LandCover Mapping & Multi-Modal Segmentation: Time-sensitive Crop Mapping & Multi-Modal Scene Classification |
| VLEO-Bench | Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data | Arxiv2024 | VLEO-bench | link | Location Recognition & Captioning & Scene Classification & Counting & Detection & Change detection |
| VRSBench | VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding | NeurIPS2024 | VRSBench | link | Image Captioning & Object Referring & Visual Question Answering |
| UrBench | UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios | AAAI2025 | UrBench | link | Object Referring & Visual Question Answering & Counting & Scene Classification & Location Recognition & Geolocalization |
| PANGAEA | PANGAEA: A Global and Inclusive Benchmark for Geospatial Foundation Models | Arxiv2024 | PANGAEA | link | Segmentation & Change detection & Regression |
| CHOICE | CHOICE: Evaluating and Understanding Vision-Language Model Choices in Remote Sensing | NeurIPS2025 | CHOICE | link | Perception & Reasoning |
| GEO-Bench-VLM | GEO-Bench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks | ICCV2025 | GEO-Bench-VLM | link | Scene Understanding & Counting & Object Classification & Event Detection & Spatial Relations |
| Copernicus-Bench | Towards a Unified Copernicus Foundation Model for Earth Vision | ICCV2025 | Copernicus-Bench | link | Segmentation & Classification & Change detection & Regression |
| REOBench | REOBench: Benchmarking Robustness of Earth Observation Foundation Models | Arxiv2025 | REOBench | link | Robustness across 6 Earth observation tasks |
| Plantation Bench | Plantation Bench: A Multiscale, Multimodal Remote Sensing Benchmark for Plantation Mapping Under Distribution Shift | ICCVW2025 | Plantation Bench | null | Plantation Mapping under Distribution Shift |
| ChatEarthBench | ChatEarthBench: Benchmarking multimodal large language models for Earth observation | IEEE GRSM2026 | ChatEarthBench | null | Benchmarking EO multimodal large language models |
| GeoReason-Bench | GeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement Learning | Arxiv2026 | GeoReason-Bench | link | Logical consistency & multi-step reasoning |
| Earth-Bench | Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents (Earth-Bench is the benchmark introduced in this paper) | ICLR2026 | Earth-Agent | link | Tool-augmented EO reasoning & multi-step planning & quantitative spatiotemporal analysis |
| OmniEarth | OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks | Arxiv2026 | OmniEarth | link | Perception & Reasoning & Robustness across geospatial tasks |
| SpatialSky-Bench | Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation | CVPR2026 | SpatialSky-Bench | null | UAV spatial intelligence evaluation for Vision-Language Models |
| GeoMMBench | GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing | CVPR2026 | GeoMMBench | null | Expert-level multimodal QA across RS disciplines, sensors & tasks |
| RSVLM-QA | RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering | ACM MM2025 | RSVLM-QA | link | Visual Question Answering & Image Captioning & Counting |
| Geo3DVQA | Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery | WACV2026 | Geo3DVQA | link | 3D geospatial reasoning & height-aware spatial analysis |
| OpenEarth-Bench | OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation (OpenEarth-Bench is the benchmark introduced in this paper) | Arxiv2026 | OpenEarth-Agent | link | Full-pipeline EO across 7 application domains |
| GeoAgentBench | GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis | Arxiv2026 | GeoAgentBench | null | Dynamic execution evaluation for tool-augmented GIS agents |
| Abbreviation | Title | Publication | Paper | Attribute | Link |
|---|---|---|---|---|---|
| fMoW | Functional Map of the World | CVPR2018 | fMoW | Vision | link |
| SEN12MS | SEN12MS -- A Curated Dataset of Georeferenced Multi-Spectral Sentinel-1/2 Imagery for Deep Learning and Data Fusion | - | SEN12MS | Vision | link |
| BEN-MM | BigEarthNet-MM: A Large Scale Multi-Modal Multi-Label Benchmark Archive for Remote Sensing Image Classification and Retrieval | GRSM2021 | BEN-MM | Vision | link |
| MillionAID | On Creating Benchmark Dataset for Aerial Image Interpretation: Reviews, Guidances, and Million-AID | JSTARS2021 | MillionAID | Vision | link |
| SeCo | Seasonal Contrast: Unsupervised Pre-Training From Uncurated Remote Sensing Data | ICCV2021 | SeCo | Vision | link |
| fMoW-S2 | SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery | NeurIPS2022 | fMoW-S2 | Vision | link |
| TOV-RS-Balanced | TOV: The original vision model for optical remote sensing image understanding via self-supervised learning | JSTARS2023 | TOV | Vision | link |
| SSL4EO-S12 | SSL4EO-S12: A Large-Scale Multi-Modal, Multi-Temporal Dataset for Self-Supervised Learning in Earth Observation | GRSM2023 | SSL4EO-S12 | Vision | link |
| SSL4EO-L | SSL4EO-L: Datasets and Foundation Models for Landsat Imagery | Arxiv2023 | SSL4EO-L | Vision | link |
| SatlasPretrain | SatlasPretrain: A Large-Scale Dataset for Remote Sensing Image Understanding | ICCV2023 | SatlasPretrain | Vision (Supervised) | link |
| CACo | Change-Aware Sampling and Contrastive Learning for Satellite Images | CVPR2023 | CACo | Vision | Coming soon |
| SAMRS | SAMRS: Scaling-up Remote Sensing Segmentation Dataset with Segment Anything Model | NeurIPS2023 | SAMRS | Vision | link |
| RSVG | RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing Data | TGRS2023 | RSVG | Vision-Language | link |
| RS5M | RS5M: A Large Scale Vision-Language Dataset for Remote Sensing Vision-Language Foundation Model | Arxiv2023 | RS5M | Vision-Language | link |
| GEO-Bench | GEO-Bench: Toward Foundation Models for Earth Monitoring | Arxiv2023 | GEO-Bench | Vision (Evaluation) | link |
| RSICap & RSIEval | RSGPT: A Remote Sensing Vision Language Model and Benchmark | Arxiv2023 | RSGPT | Vision-Language | Coming soon |
| Clay | Clay Foundation Model | - | null | Vision | link |
| SATIN | SATIN: A Multi-Task Metadataset for Classifying Satellite Imagery using Vision-Language Models | ICCVW2023 | SATIN | Vision-Language | link |
| SkyScript | SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing | AAAI2024 | SkyScript | Vision-Language | link |
| ChatEarthNet | ChatEarthNet: a global-scale image-text dataset empowering vision-language geo-foundation models | ESSD2025 | ChatEarthNet | Vision-Language | link |
| LuoJiaHOG | LuoJiaHOG: A hierarchy oriented geo-aware image caption dataset for remote sensing image-text retrieval | ISPRS JPRS2025 | LuoJiaHOG | Vision-Language | null |
| MMEarth | MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning | Arxiv2024 | MMEarth | Vision | link |
| SeeFar | SeeFar: Satellite Agnostic Multi-Resolution Dataset for Geospatial Foundation Models | Arxiv2024 | SeeFar | Vision | link |
| FIT-RS | SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding | Arxiv2024 | Paper | Vision-Language | link |
| RS-GPT4V | RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding | Arxiv2024 | Paper | Vision-Language | link |
| RS-4M | Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling | ICCV2025 | RS-4M | Vision | link |
| Major TOM | Major TOM: Expandable Datasets for Earth Observation | Arxiv2024 | Major TOM | Vision | link |
| VRSBench | VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding | NeurIPS2024 | VRSBench | Vision-Language | link |
| MMM-RS | MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation | Arxiv2024 | MMM-RS | Vision-Language | link |
| DDFAV | DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark | RS2025 | DDFAV | Vision-Language | link |
| M3LEO | A Multi-Modal, Multi-Label Earth Observation Dataset Integrating Interferometric SAR and Multispectral Data | NeurIPS2024 | M3LEO | Vision | link |
| Copernicus-Pretrain | Towards a Unified Copernicus Foundation Model for Earth Vision | ICCV2025 | Copernicus-Pretrain | Vision | link |
| DGTRSD | DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment | Arxiv2025 | Paper | Vision-Language | link |
| EarthDial-Instruct | EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues | CVPR2025 | Paper | Vision-Language | link |
| GeoPixelD | GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing | ICML2025 | Paper | Vision-Language | link |
| GeoPixInstruct | GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing | IEEE GRSM2025 | Paper | Vision-Language | link |
| GeoLangBind-2M | Rethinking Remote Sensing CLIP: Leveraging Multimodal Large Language Models for High-Quality Vision-Language Dataset | ICONIP2024 | Paper | Vision-Language | link |
| Falcon_SFT | Falcon: A Remote Sensing Vision-Language Foundation Model | Arxiv2025 | Paper | Vision-Language | link |
| UnivEARTH | Towards LLM Agents for Earth Observation: The UnivEARTH Dataset | Arxiv2025 | Paper | Vision-Language & Agents | null |
| RemoteSAM-270K | RemoteSAM: Towards Segment Anything for Earth Observation | ACMMM2025 | Paper | Vision-Language | link |
| OpenEarthAgent Dataset | OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents | Arxiv2026 | Paper | Vision-Language & Agents | link |
| UHR-CoZ | GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery | Arxiv2026 | Paper | Vision-Language | link |
| SOMA-1M | SOMA-1M: A Large-Scale SAR-Optical Multi-resolution Alignment Dataset for Multi-Task Remote Sensing | Arxiv2026 | SOMA-1M | Vision (SAR-Optical) | link |
| Abbreviation | Title | Publication | Paper | Code | Dataset / Product |
|---|---|---|---|---|---|
| CLAY Embeddings | Clay Model v0 Embeddings | Source Cooperative2024 | null | link | link |
| Major TOM Embeddings | Global and Dense Embeddings of Earth: Major TOM Floating in the Latent Space | Arxiv2024 | Paper | link | link |
| Earth Genome Embeddings | Embeddings for all | Medium2025 | Paper | null | link |
| TESSERA | TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis | CVPR2026 | Paper | link | link |
| AlphaEarth | AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data | Arxiv2025 | Paper | null | link |
| ESD | Democratizing planetary-scale analysis: An ultra-lightweight Earth embedding database for accurate and flexible global land monitoring | Arxiv2026 | Paper | link | link |
| Copernicus-Embed | Towards a Unified Copernicus Foundation Model for Earth Vision | ICCV2025 Oral | Paper | link | link |
| Title | Link | Brief Introduction |
|---|---|---|
| RSFMs (Remote Sensing Foundation Models) Playground | link | An open-source playground to streamline the evaluation and fine-tuning of RSFMs on various datasets. |
| PANGAEA | link | A Global and Inclusive Benchmark for Geospatial Foundation Models. |
| GeoFM | link | Evaluation of Foundation Models for Earth Observation. |
| rs-embed | link | One line code to get Any Remote Sensing Foundation Model (RSFM) embeddings for Any Place and Any Time. |
| TerraTorch | link | A PyTorch toolkit for fine-tuning Geospatial Foundation Models, supporting Prithvi, TerraMind, SatMAE, ScaleMAE, DOFA, Clay, and more. |
| TorchGeo | link | A PyTorch domain library for geospatial data, providing datasets, samplers, transforms, and pre-trained models. |
| Awesome-Geospatial-Embeddings | link | A curated list of papers on how to represent Earth data in embedding space — spatial, temporal, or semantic. |
| Title | Publication | Paper | Attribute |
|---|---|---|---|
| The Potential of Visual ChatGPT For Remote Sensing | Arxiv2023 | Paper | Vision-Language |
| Self-Supervised Remote Sensing Feature Learning: Learning Paradigms, Challenges, and Future Works | TGRS2023 | Paper | Vision & Vision-Language |
| Revisiting pre-trained remote sensing model benchmarks: resizing and normalization matters | Arxiv2023 | Paper | Vision |
| An Agenda for Multimodal Foundation Models for Earth Observation | IGARSS2023 | Paper | Vision |
| 遥感大模型:进展与前瞻 | 武汉大学学报 (信息科学版) 2023 | Paper | Vision & Vision-Language |
| 地理人工智能样本:模型、质量与服务 | 武汉大学学报 (信息科学版) 2023 | Paper | - |
| Brain-Inspired Remote Sensing Foundation Models and Open Problems: A Comprehensive Survey | JSTARS2023 | Paper | Vision & Vision-Language |
| 遥感基础模型发展综述与未来设想 | 遥感学报2023 | Paper | - |
| On the Promises and Challenges of Multimodal Foundation Models for Geographical, Environmental, Agricultural, and Urban Planning Applications | Arxiv2023 | Paper | Vision-Language |
| Transfer learning in environmental remote sensing | RSE2024 | Paper | Transfer learning |
| Vision-Language Models in Remote Sensing: Current Progress and Future Trends | IEEE GRSM2024 | Paper | Vision-Language |
| On the Foundations of Earth and Climate Foundation Models | Arxiv2024 | Paper | Vision & Vision-Language |
| 多模态遥感基础大模型:研究现状与未来展望 | 测绘学报2024 | Paper | Vision & Vision-Language & Generative & Vision-Location |
| Towards Vision-Language Geo-Foundation Model: A Survey | Arxiv2024 | Paper | Vision-Language |
| Vision Foundation Models in Remote Sensing: A Survey | Arxiv2024 | Paper | Vision |
| Foundation model for generalist remote sensing intelligence: Potentials and prospects | Science Bulletin2024 | Paper | - |
| Advancements in Visual Language Models for Remote Sensing: Datasets, Capabilities, and Enhancement Techniques | Arxiv2024 | Paper | Vision-Language |
| When Geoscience Meets Foundation Models: Toward a general geoscience artificial intelligence system | IEEE GRSM2024 | Paper | Vision & Vision-Language |
| Towards the next generation of Geospatial Artificial Intelligence | JAG2025 | Paper | - |
| When Remote Sensing Meets Foundation Model: A Survey and Beyond | RS2025 | Paper | Vision & Vision-Language & Generative & Agents |
| Vision Foundation Models in Remote Sensing: A survey | IEEE GRSM2025 | Paper | Vision |
| Remote Sensing Tuning: A Survey | CVM2025 | Paper | Vision & Vision-Language |
| Unleashing the potential of remote sensing foundation models via bridging data and computility islands | The Innovation2025 | Paper | - |
| A Survey on Remote Sensing Foundation Models: From Vision to Multimodality | Arxiv2025 | Paper | - |
| Foundation Models for Remote Sensing and Earth Observation: A survey | IEEE GRSM2025 | Paper | Vision & Vision-Language |
| Vision-Language Modeling Meets Remote Sensing: Models, datasets, and perspectives | IEEE GRSM2025 | Paper | Vision-Language |
| MIMRS: A Survey on Masked Image Modeling in Remote Sensing | IGARSS2025 | Paper | Vision |
| A Review of Challenges and Applications in Remote Sensing Foundation Models | IGARSS2025 | Paper | Vision & Vision-Language |
| On the Status of Foundation Models for SAR Imagery | Arxiv2025 | Paper | Vision (SAR) |
| Advances on Multimodal Remote Sensing Foundation Models for Earth Observation Downstream Tasks: A Survey | RS2025 | Paper | Vision & Vision-Language |
| Agentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging Systems | WACVW2026 | Paper | Agents |
| Onboard Deployment of Remote Sensing Foundation Models: A Comprehensive Review of Architecture, Optimization, and Hardware | RS2026 | Paper | Vision & Vision-Language |
| On the foundations of Earth foundation models | Communications Earth & Environment 2026 | Paper | Vision & Vision-Language |
| A Genealogy of Foundation Models in Remote Sensing | ACM TSAS2026 | Paper | Vision & Vision-Language |
| Foundation Models in Remote Sensing: Evolving from Unimodality to Multimodality | IEEE GRSM2026 | Paper | Vision & Vision-Language |
If you find this repository useful, please consider giving a star ⭐ and citation:
@inproceedings{guo2024skysense,
title={Skysense: A multi-modal remote sensing foundation model towards universal interpretation for earth observation imagery},
author={Guo, Xin and Lao, Jiangwei and Dang, Bo and Zhang, Yingying and Yu, Lei and Ru, Lixiang and Zhong, Liheng and Huang, Ziyuan and Wu, Kang and Hu, Dingxiang and others},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
pages={27672--27683},
year={2024}
}
@article{li2025unleashing,
title={Unleashing the potential of remote sensing foundation models via bridging data and computility islands},
author={Li, Yansheng and Tan, Jieyi and Dang, Bo and Ye, Mang and Bartalev, Sergey A and Shinkarenko, Stanislav and Wang, Linlin and Zhang, Yingying and Ru, Lixiang and Guo, Xin and others},
journal={The Innovation},
year={2025},
publisher={Elsevier}
}
@article{wu2025semantic,
author = {Wu, Kang and Zhang, Yingying and Ru, Lixiang and Dang, Bo and Lao, Jiangwei and Yu, Lei and Luo, Junwei and Zhu, Zifan and Sun, Yue and Zhang, Jiahao and Zhu, Qi and Wang, Jian and Yang, Ming and Chen, Jingdong and Zhang, Yongjun and Li, Yansheng},
title= {A semantic‑enhanced multi‑modal remote sensing foundation model for Earth observation},
journal= {Nature Machine Intelligence},
year= {2025},
doi= {10.1038/s42256-025-01078-8},
url= {https://doi.org/10.1038/s42256-025-01078-8}
}
@inproceedings{zhu2025skysense,
title={Skysense-o: Towards open-world remote sensing interpretation with vision-centric visual-language modeling},
author={Zhu, Qi and Lao, Jiangwei and Ji, Deyi and Luo, Junwei and Wu, Kang and Zhang, Yingying and Ru, Lixiang and Wang, Jian and Chen, Jingdong and Yang, Ming and others},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
pages={14733--14744},
year={2025}
}
@article{luo2024skysensegpt,
title={Skysensegpt: A fine-grained instruction tuning dataset and model for remote sensing vision-language understanding},
author={Luo, Junwei and Pang, Zhen and Zhang, Yongjun and Wang, Tingzhu and Wang, Linlin and Dang, Bo and Lao, Jiangwei and Wang, Jian and Chen, Jingdong and Tan, Yihua and others},
journal={arXiv preprint arXiv:2406.10100},
year={2024}
}