Data-Centric Land Cover Classification Challenge
In remote sensing image classification, training state-of-the-art deep learning models typically requires large Earth observation datasets containing thousands of high-resolution satellite and aerial images. However, storing, transferring, and training models on these large-scale geospatial datasets incurs immense computational costs, energy consumption, and storage overhead. Dataset distillation offers a promising solution by synthesizing a small, highly informative set of images that captures the essential spatial patterns of the full dataset. When a model is trained on this compact synthetic set, it achieves competitive performance in a fraction of the time and compute.
This Challenge invites you to develop novel dataset distillation algorithms tailored for remote sensing imagery. Your challenge is to condense large-scale satellite data into a minimal set of synthetic, semantically rich visual representations. By advancing dataset distillation for Earth observation, your solutions will help enable resource-efficient planetary-scale mapping, on-board satellite edge computing, and rapid model retraining.
Challenge Summary
In this competition, you will develop machine learning methods to perform dataset distillation (also known as dataset synthesis) for remote sensing image classification. The primary objective is to condense a large training dataset into a tiny, highly concentrated set of synthetic images while retaining model performance when training from scratch on the reduced set. To this end, you are provided with a training dataset composed of 18,900 64x64 images spanning 10 classes, alongside their corresponding ground truth labels.
Success is evaluated by comparing your submitted test predictions with the undisclosed ground truth and calculating a composite score that balances overall accuracy against a log-normalized efficiency score based on the number of synthetic images used during training. To ensure fair comparison, all submissions must evaluate their distilled datasets using the same model architecture - specifically, a Vision Transformer (ViT). Starter baseline code is provided here - note that hyperparameters (such as learning rate, etc.) must not be altered in the evaluation code.
Top-performing teams will be required to submit their distilled training images for verification after the competition concludes.
Please, check for more info on Kaggle.
Dataset
The challenge will use this dataset, which is composed of 18,900 RGB satellite/aerial training images evenly split across 10 anonymized land-cover classes. Validation and test sets are also available (with 4,050 samples each). Participants are tasked with synthesizing a minimal, highly informative subset of training images while maintaining model classification performance on the test set.
More info on Kaggle.
Submission
Your submission must be a single .csv file containing predictions for all test images PLUS exactly one sentinel row declaring how many synthetic images were used to train your final model. These are the submission file requirements:
- Header: Must contain
filenameandlabelcolumns. - Sentinel Row: Must include the exact row
filename = "__num_train_images__"wherelabelis an integer indicating number of training samples. - Prediction Rows: One row for every test image ID with its predicted class label.
More information can be found here. After creating this file, please submit it on Kaggle.
Timeline
Participants may submit multiple solutions. All submitted solutions will be listed on the leaderboard.
| Subject | Date |
|---|---|
| Submission Deadline | Sunday, 15 November 2026, 23:59 (UK time) |
| Workshop | Thursday, 26 November 2026 |
Results, Presentation, Awards, and Prizes
The final results of this challenge will be presented during the Workshop. The authors of the top-ranked methods will be invited to present their approaches at the Workshop in Lancaster/UK, on 26 November 2026. These authors will also be invited to co-author a journal paper which will summarize the outcome of this challenge and will be submitted with open access to IEEE JSTARS.
Organizers
Keiller Nogueira, University of Liverpool, UK
Ronny Hänsch, German Aerospace Center (DLR), Germany
Wanli Ma, University of Cambridge, UK
Vahid Akbari, University of Stirling, UK