首页|Looking for a reference for large datasets: relative reliability of visual and automatic sleep scoring

Looking for a reference for large datasets: relative reliability of visual and automatic sleep scoring

来源：

英文摘要

Abstract Study ObjectivesNew challenges in sleep science require to describe fine grain phenomena or to deal with large datasets. Beside the human resource challenge of scoring huge datasets, the inter- and intra-expert variability may also reduce the sensitivity of such studies. Searching for a way to disentangle the variability induced by the scoring method from the actual variability in the data, visual and automatic sleep scorings of healthy individuals were examined. MethodsA first dataset (DS1, 4 recordings) scored by 6 experts plus an autoscoring algorithm was used to characterize inter-scoring variability. A second dataset (DS2, 88 recordings) scored a few weeks later was used to investigate intra-expert variability. Percentage agreements and Conger’s kappa were derived from epoch-by-epoch comparisons on pairwise, consensus and majority scorings. ResultsOn DS1 the number of epochs of agreement decreased when the number of expert increased, in both majority and consensus scoring, where agreement ranged from 86% (pairwise) to 69% (all experts). Adding autoscoring to visual scorings changed the kappa value from 0.81 to 0.79. Agreement between expert consensus and autoscoring was 93%. On DS2 intra-expert variability was evidenced by the kappa systematic decrease between autoscoring and each single expert between datasets (0.75 to 0.70). ConclusionsVisual scoring induces inter- and intra-expert variability, which is difficult to address especially in big data studies. When proven to be reliable and if perfectly reproducible, autoscoring methods can cope with intra-scorer variability making them a sensible option when dealing with large datasets. Statement of SignificanceWe confirmed and extended previous findings highlighting the intra- and inter-expert variability in visual sleep scoring. On large datasets those variability issues cannot be completely addressed by neither practical nor statistical solutions such as group training, majority or consensus scoring.When an automated scoring method can be proven to be as reasonably imperfect as visual scoring but perfectly reproducible, it can serve as a reliable scoring reference for sleep studies.

作者：Jaspar M.、Benoit O.、Maquet J.、Vandewalle G.、Meyer C.、Phillips C.、Berthomier P.、Berthomier C.、Gaggioni G.、Chellappa S. L.、Salmon E.、Devillers J.、Schmidt C.、Brandewinder M.、Prado J.、Mattout J.、Muto V.

作者单位：GIGA-Cyclotron Research Centre-In vivo Imaging||Walloon Excellence in Life sciences and Biotechnology (WELBIO)||Psychology and Cognitive Neuroscience Research UnitPHYSIPGIGA-Cyclotron Research Centre-In vivo Imaging||Walloon Excellence in Life sciences and Biotechnology (WELBIO)||Department of NeurologyGIGA-Cyclotron Research Centre-In vivo Imaging||Walloon Excellence in Life sciences and Biotechnology (WELBIO)GIGA-Cyclotron Research Centre-In vivo Imaging||Walloon Excellence in Life sciences and Biotechnology (WELBIO)GIGA-Cyclotron Research Centre-In vivo Imaging||Department of Electrical Engineering and Computer SciencePHYSIPPHYSIPGIGA-Cyclotron Research Centre-In vivo ImagingGIGA-Cyclotron Research Centre-In vivo Imaging||Walloon Excellence in Life sciences and Biotechnology (WELBIO)GIGA-Cyclotron Research Centre-In vivo ImagingGIGA-Cyclotron Research Centre-In vivo Imaging||Walloon Excellence in Life sciences and Biotechnology (WELBIO)GIGA-Cyclotron Research Centre-In vivo Imaging||Psychology and Cognitive Neuroscience Research UnitPHYSIPPHYSIPLyon Neuroscience Research CenterGIGA-Cyclotron Research Centre-In vivo Imaging||Walloon Excellence in Life sciences and Biotechnology (WELBIO)||Psychology and Cognitive Neuroscience Research Unit

DOI：10.1101/576090

学科分类：医学研究方法基础医学神经病学、精神病学

英文关键词：Autoscoringscoring variabilitylarge datasets

推荐引用：Jaspar M.,Benoit O.,Maquet J.,Vandewalle G.,Meyer C.,Phillips C.,Berthomier P.,Berthomier C.,Gaggioni G.,Chellappa S. L.,Salmon E.,Devillers J.,Schmidt C.,Brandewinder M.,Prado J.,Mattout J.,Muto V..Looking for a reference for large datasets: relative reliability of visual and automatic sleep scoring[EB/OL].(2025-03-28)[2025-04-26].https://www.biorxiv.org/content/10.1101/576090.点此复制

Looking for a reference for large datasets: relative reliability of visual and automatic sleep scoring

Looking for a reference for large datasets: relative reliability of visual and automatic sleep scoring

评论