Aller directement à la PREreview
PREreview demandée

PREreview de Digital Acoustic Phenotypes of Overnight Respiratory Activity

Publié
DOI
10.5281/zenodo.22789060
Licence
CC BY 4.0

Summary of main findings

This exploratory study reanalyzes 32 overnight home recordings (audio + respiratory polygraphy) from the APSAA dataset for obstructive sleep apnea. The premise: a respiratory-event rate (AHI-like) summarizes an entire night as one number but discards how respiratory-related acoustic activity is distributed over time. The author built a "digital acoustic phenotype" from five acoustic features (episode density, relative intensity, duration, late long-event occurrence, temporal distribution) combined into a composite score classifying each recording as High/Intermediate/Low.

Core findings: (1) The score was computationally reproduced with near-perfect fidelity in 28/28 recordings. (2) This acoustic representation does not simply mirror annotated respiratory-event rate — two recordings with nearly identical annotated rates (30.45 vs 33.38 events/h) contained 382 vs 5 acoustic episodes. (3) Load-adjusted "temporal compaction" differed across phenotypes (p=0.0166), but the effect was concentrated almost entirely in the Low-score group.

Field contribution: the core idea — that event frequency alone hides temporal organization — is methodologically reasonable and worth pursuing, but this is a very preliminary, small-sample exploratory analysis, not a demonstration of real biological phenotypes, as the author himself repeatedly and explicitly states throughout.

Major issues

  • Very small sample, especially the Low-score group: only 4 recordings drive the paper's headline statistical result (temporal compaction). Leave-one-out analyses show the Low-vs-High comparison is unstable (significant in only 1 of 4 omissions), while Low-vs-Intermediate is comparatively more robust.

  • Arbitrary, non-biological thresholds: the High/Intermediate boundary separation is just 0.004 around the 0.60 cutoff — smaller than the reconstruction error itself (max 0.0046). This means the two categories are essentially indistinguishable at their boundary and the threshold has no independent justification.

  • One component saturated in 75% of the sample: the most heavily weighted component (density, 30% weight) hit its ceiling in 21/28 recordings, meaning it contributed zero discriminating power across three-quarters of the data — undermining the composite score's claimed multidimensionality.

  • Detector not validated against ground truth: the acoustic segmentation rule was "inherited" from a prior exploratory analysis and never optimized or validated against the respiratory annotations. Four recordings yielded zero acoustic episodes despite substantial annotated respiratory event burden (one had 537 annotated events, second-highest in the dataset) — a fundamental detector failure mode that the paper acknowledges but doesn't resolve.

  • No clinical validation whatsoever: the "phenotypes" are explicitly stated to not represent sleep stages, OSA severity, or any clinical construct — this is appropriately cautious, but it also means the paper demonstrates a signal-processing pipeline is reproducible, not that it measures anything clinically meaningful.

  • Single night per participant: no within-person reproducibility, stability, or night-to-night variability can be assessed — a critical gap given documented substantial night-to-night variability in OSA itself (cited by the author).

  • Pending patent application disclosed at the end — worth noting as a potential conflict-of-interest consideration even though formally declared.

Minor issues

  • The Methods section is extremely dense with normalization formulas and rounding rules; a summary table of the scoring pipeline early on would aid readability.

  • Repeated caveating throughout ("this is not a clinical phenotype...") is appropriate but becomes somewhat repetitive — could be consolidated into one clear framing statement plus targeted reminders.

  • Figure 1B is visually informative but a supplementary table of per-component saturation rates would make the "component D is capped in 75%" point more quantitatively transparent.

  • The comparison between annotation-derived rate and acoustic score (Section 3.3) would benefit from a scatter plot matrix rather than prose-heavy pairwise anchor-recording narration.

  • AI-assistance disclosure is present and appropriately scoped (workflow/scripting/figures only, not analysis/interpretation) — good practice, no issue there.

Competing interests

The author declares that they have no competing interests.

Use of Artificial Intelligence (AI)

The author declares that they did not use generative AI to come up with new ideas for their review.