首页|A Dataset is Worth 1 MB

A Dataset is Worth 1 MB

Elad Kimchi Shoshani Leeyam Gabay Yedid Hoshen

来源：

Arxiv

A Dataset is Worth 1 MB

Elad Kimchi Shoshani Leeyam Gabay Yedid Hoshen

作者信息

Abstract

A dataset server must often distribute the same large payload to many clients, incurring massive communication costs. Since clients frequently operate on diverse hardware and software frameworks, transmitting a pre-trained model is often infeasible; instead, agents require raw data to train their own task-specific models locally. While dataset distillation attempts to compress training signals, current methods struggle to scale to high-resolution data and rarely achieve sufficiently small files. In this paper, we propose Pseudo-Labels as Data (PLADA), a method that completely eliminates pixel transmission. We assume agents are preloaded with a large, generic, unlabeled reference dataset (e.g., ImageNet-1K, ImageNet-21K) and communicate a new task by transmitting only the class labels for specific images. To address the distribution mismatch between the reference and target datasets, we introduce a pruning mechanism that filters the reference dataset to retain only the labels of the most semantically relevant images for the target task. This selection process simultaneously maximizes training efficiency and minimizes transmission payload. Experiments on 10 diverse datasets demonstrate that our approach can transfer task knowledge with a payload of less than 1 MB while retaining high classification accuracy, offering a promising solution for efficient dataset serving.

引用本文复制引用

Elad Kimchi Shoshani,Leeyam Gabay,Yedid Hoshen.A Dataset is Worth 1 MB[EB/OL].(2026-02-26)[2026-02-28].https://arxiv.org/abs/2602.23358.

学科分类

计算技术、计算机技术

首发时间： 2026-02-26

下载量：0

点击量：4

段落导航

A Dataset is Worth 1 MB

A Dataset is Worth 1 MB

Abstract

引用本文复制引用

学科分类

评论