跨场景适配的多模态科学数据可复用性智能评价方法研究
Research on an Intelligent Evaluation Method for the Reusability of Multimodal Scientific Data with Cross-Scenario Adaptability
闫琦 1季亚楠 1汤廷楷1
作者信息
- 1. 天津师范大学管理学院 天津 300387
- 折叠
摘要
【目的/意义】随着以数据为中心的人工智能范式快速发展,构建面向 AI 可理解与场景适配的多模态科学数据可复用性评价框架,对提升科学数据的实际复用能力具有重要意义。【方法/过程】本文提出 FAIR C-S 三层可复用性评估框架,在 FAIR 原则基础上引入可理解性(C)维度与场景适配性(S)维度,构建涵盖5个评价维度、18个可机器执行指标的综合量化评价方法,并形成规则计算与 AI 模型协同的智能评价机制。基于 Kaggle、Figshare 与 Zenodo 三大平台的300个跨学科数据集开展实证研究,采用非参数检验进行平台间与跨学科差异分析,并与现有 FAIR 评价工具进行一致性比较。【结果/结论】实证结果表明,三平台在 F、A、I、C 和 S 五个维度上均存在显著差异(p < 0.01);跨学科比较中,F、A、I 和 S 四个维度呈现显著差异(p < 0.01),C 维度未达到显著水平;与现有 FAIR 工具的比较显示,F 和 I 维度总体评价结果呈显著正相关,而 A 维度未表现出显著一致性;同时识别出61个多模态数据集。研究结果验证了 FAIR C-S 框架的跨平台、跨学科差异识别能力,拓展了 FAIR 框架的评价边界,为跨平台跨场景的多模态科学数据可复用性评价提供了理论依据与方法参考。
Abstract
[Purpose/Significance] With the rapid development of data-centric artificial intelligence, AI is becoming increasingly involved in scientific data analysis and reuse, making machine recognition and understanding of data semantics and contextual information an important condition for effective reuse. Meanwhile, the growing importance of multimodal scientific data in data-intensive research has increased the complexity of data comprehension and cross-scenario reuse. Scientific data reuse is also constrained by specific application scenarios and task requirements. Existing FAIR-based assessments, however, mainly focus on compliance with technical metadata requirements and generally design indicators according to FAIR sub-principles, with limited attention to AI-oriented comprehensibility and the fit between data and specific reuse scenarios. This study therefore develops a multimodal scientific data reusability evaluation framework incorporating both AI comprehensibility and scenario adaptability. [Methods/Process] This study proposes the three-tier FAIR C-S reusability evaluation framework, repositioning reusability as the top-level evaluation objective. Building on the FAIR principles, a Comprehensibility (C) dimension is introduced to assess the data conditions supporting AI understanding, while a Scenario Adaptability (S) dimension is developed based on Task-Technology Fit (TTF) theory. The framework comprises five evaluation dimensions and 18 machine-executable indicators and incorporates an intelligent evaluation mechanism integrating rule-based computation with AI models. An empirical study was conducted on 300 datasets from Kaggle, Figshare, and Zenodo across six disciplines: life sciences, chemical sciences, computer science, earth and space sciences, psychology, and economics. Kruskal-Wallis H tests and Dunns post-hoc tests were used to examine cross-platform and cross-disciplinary differences, while a modality recognition agent assisted in identifying multimodal datasets. An existing automated FAIR assessment tool was further introduced for comparison on Findability (F), Accessibility (A), and Interoperability (I), with Kendalls -b used to test consistency between the two assessment results. [Result/Conclusion] Of the 300 datasets, 61 were identified as multimodal. Significant differences were found among Kaggle, Figshare, and Zenodo across all five dimensions, F, A, I, C, and S (p < 0.01). Kaggle performed relatively better in Interoperability, Figshare in Accessibility and Scenario Adaptability, and Zenodo in Comprehensibility. Across disciplines, significant differences were observed in F, A, I, and S (p < 0.01), whereas C showed no significant disciplinary difference. Life sciences, chemical sciences, computer science, and earth and space sciences generally performed better in F and I, while psychology and economics showed stronger Scenario Adaptability, indicating that general FAIR conditions and scenario-oriented reuse support do not exhibit identical distribution patterns across disciplines. Comparison with the existing automated FAIR assessment tool showed significant positive correlations for F and I, but no significant consistency for A. Overall, the FAIR C-S framework can distinguish differences in basic FAIR conditions, AI comprehensibility, and scenario adaptability across platforms and disciplines, while its fine-grained indicators reveal internal characteristics that may be obscured by composite scores. [Innovation/Value] The study contributes at three levels. Theoretically, it reorganizes the hierarchical relationships among FAIR elements by positioning reusability as the top-level objective and introducing Comprehensibility and Scenario Adaptability based on cognitive load theory and TTF theory, respectively. This extends the evaluation scope from general FAIR conditions to AI-oriented understanding and task-specific adaptability. Methodologically, machine-executable quantitative indicators are integrated with AI-based semantic processing, extending scientific data reusability assessment from rule-driven automated measurement toward intelligent evaluation of both structured attributes and complex information. Its operability and reproducibility in cross-platform assessment are validated using 300 cross-disciplinary datasets. Practically, the study identifies weaknesses in scientific data reuse conditions from the perspectives of basic data standards, support for AI understanding, and scenario-oriented reuse requirements, providing empirical evidence for improving data organization, description, and reuse-support mechanisms. [Limitations/Improvement] The main limitation of this study is that the sample is concentrated on Kaggle, Figshare, and Zenodo. The generalizability of the findings requires further validation across a broader range of data repositories and disciplinary fields.关键词
多模态科学数据/可复用性/可理解性/场景适配性Key words
Multimodal scientific data/Reusability/Comprehensibility/Scenario-adaptability引用本文复制引用
闫琦,季亚楠,汤廷楷.跨场景适配的多模态科学数据可复用性智能评价方法研究[EB/OL].(2026-09-24)[2026-09-29].https://chinaxiv.org/abs/202609.00315.学科分类
信息科学、信息技术