Saltar al contenido principal

Escribir un comentario

PREreview del Refining Cell-free DNA Metagenomics for Sepsis in Acute Leukemia: Towards Standardized Computational Strategies for Host Depletion and Microbial Detection from Shallow-Depth, Low-Coverage Profiles

Publicado
DOI
10.5281/zenodo.22992744
Licencia
CC0 1.0

Overall assessment

This manuscript addresses an important problem in plasma cell-free DNA metagenomics: microbial reads are present at very low abundance against an overwhelming background of human DNA, and computational host depletion can therefore have a major effect on downstream microbial detection. The authors compare BWA-MEM2, Minimap2, several Bowtie2 configurations, and Kraken2 using simulated data, plasma cfDNA from patients with acute leukemia and sepsis, and additional public sepsis datasets.

The work is potentially useful because host-read removal is often treated as a routine preprocessing step despite its ability to substantially alter downstream results. The combination of simulation and clinical material is also a strength.

However, the manuscript currently supports a narrower conclusion than the one presented. The data demonstrate that different computational approaches generate substantially different outputs. They do not yet establish that BWA-MEM2 and Bowtie2 very-sensitive-local constitute generally optimal solutions for clinical cfDNA metagenomics.

The major concern is the lack of an independent biological ground truth for most of the patient-level analyses. Additional concerns include relatively low absolute taxonomic classification performance, favourable simulation conditions, the restricted microbial reference database, inconsistent comparison of fundamentally different algorithms, and clinical claims regarding antimicrobial resistance that extend beyond what sequence detection alone can demonstrate.

Major comments

1. “Most stringent” is not equivalent to “most accurate”

BWA-MEM2 consistently leaves fewer reads classified as human after the host-depletion procedure and performs similarly in additional public sepsis datasets. The authors therefore identify it as the preferred host-removal approach.

However, greater depletion is not automatically evidence of greater accuracy. The manuscript itself shows that BWA-MEM2 removes more genuine bacterial reads than the more permissive Bowtie2 modes. Thus, the key question is not simply which method removes the greatest number of putative human reads, but which method provides the optimal balance between removal of human contamination and preservation of genuine microbial signal.

That trade-off is especially important in cfDNA metagenomics, where true pathogen-derived reads may already be extremely rare.

The manuscript would be substantially stronger if host-depletion performance were expressed using formal sensitivity and specificity measures against datasets in which the true origin of every read is known. Precision-recall or receiver-operating-characteristic analyses across mapping-quality thresholds would also be preferable to selecting a single “optimal” method primarily on the basis of stringency.

2. The taxonomic classification results are weak in absolute terms

The authors describe BWA-MEM2 as the strongest method for bacterial taxonomic identification. However, after exclusion of five species, the reported aggregate performance across 18 species is approximately:

precision = 0.498, recall = 0.227, F1 = 0.273, mapping efficiency = 35.2%.

These values indicate that BWA-MEM2 performs better than the evaluated alternatives under this particular benchmark, but they do not indicate high absolute classification accuracy.

In particular, a recall of approximately 23% means that a large proportion of expected microbial signal is not recovered.

The distinction between “best among the tested methods” and “accurate for taxonomic classification” should therefore be made much clearer throughout the manuscript.

The relatively low F1 score also raises the question of whether optimization of the alignment algorithm alone is sufficient to produce clinically reliable taxonomic calls in shallow cfDNA data.

3. The simulation design appears favourable to the classification pipeline

The taxonomic benchmark uses simulated cfDNA-like reads generated from known bacterial reference genomes and then aligns these reads to a curated bacterial reference database.

This is a relatively favourable situation because clinical organisms may differ substantially from the exact reference genomes represented in a database. Real samples additionally contain strain variation, mobile genetic elements, repetitive regions, sequencing errors, and closely related species sharing large genomic regions.

A more demanding benchmark would remove the exact source genome or strain from the reference database before classification and test whether the correct species can still be identified.

Even better, the authors could use independently sequenced clinical isolates that are absent from the reference database.

Such experiments would more realistically evaluate the generalisability of the proposed pipeline.

4. Exclusion of species from the benchmark requires stronger justification

Five organisms were removed before the final comparison of taxonomic classification performance, leaving 18 species in the analysis.

The rationale and timing of these exclusions are important. If species were removed because they performed poorly or because their reference sequences produced ambiguous results, their exclusion could make the final performance estimates more favourable.

The manuscript should clearly state predefined exclusion criteria and provide performance results both before and after exclusions.

An important benchmarking principle is that difficult organisms are themselves informative. Poor performance for particular species may reveal limitations of the method rather than observations that should be omitted from the principal analysis.

5. Validation in 202 external samples demonstrates host depletion, not taxonomic correctness

The authors report that BWA-MEM2 was the most stringent method in 201 of 202 public sepsis samples and the most specific in 188 samples.

This is useful evidence that the computational behaviour observed in the main cohort is reproducible across datasets. However, the terminology “specificity” requires caution.

If there is no independent knowledge of which residual reads are genuinely human and which are microbial, the number of reads removed cannot establish biological specificity.

These external datasets therefore validate consistency of host-depletion behaviour, rather than demonstrating that BWA-MEM2 produces the most biologically accurate microbiome profile.

The manuscript should distinguish these concepts explicitly.

6. The patient cohort lacks an adequate microbiological reference standard

The inclusion of cfDNA from 102 patients with acute leukemia and sepsis is potentially a major strength. Yet these clinical samples do not appear to provide an independent ground truth for microbial identification.

The manuscript reports substantial changes in bacterial profiles after host filtering, but the fact that a microbial profile changes after filtering does not establish that the new profile is correct.

The clinically important question is whether the organisms detected after filtering correspond to the organisms actually causing infection.

The authors should therefore compare cfDNA findings, whenever possible, against blood culture, targeted PCR, clinically validated microbiological assays, or an adjudicated diagnosis of infection.

For culture-positive episodes, species-level concordance would provide a particularly informative test.

Without such validation, apparent removal of false positives and appearance of new bacterial signals remain computational observations rather than demonstrated improvements in diagnostic accuracy.

7. The AMR analysis does not yet justify clinical treatment claims

The study concludes that Bowtie2 very-sensitive-local is preferable for antimicrobial resistance gene detection and suggests that the workflow could assist antibiotic treatment decisions.

This interpretation needs substantial qualification.

Detection of an AMR-associated sequence in plasma does not establish:

that the sequence originates from the organism causing the current infection;

that the resistance gene is expressed;

that the gene produces phenotypic resistance;

or that an antibiotic should consequently be avoided or selected.

This problem is especially important in metagenomic cfDNA because microbial DNA fragments are generally not physically linked in a way that reliably assigns every AMR gene to a particular detected organism.

Clinical claims would require comparison with isolate-level antimicrobial susceptibility testing or another validated resistance assay.

The current results support detection of AMR-associated DNA sequences, not clinical prediction of antimicrobial susceptibility.

8. The comparison combines fundamentally different classes of algorithms

BWA-MEM2, Minimap2, and Bowtie2 are sequence aligners, whereas Kraken2 is primarily a k-mer-based taxonomic classification system.

Treating all six configurations as equivalent “alignment methods” obscures important differences in their intended operation and scoring systems. The manuscript's abstract currently groups them together.

Similarly, MAPQ thresholds used for alignment-based methods are not directly equivalent to Kraken2 confidence thresholds.

The comparison would be more rigorous if performance were evaluated across multiple thresholds for each tool and compared at matched sensitivity or false-positive rates.

The terminology should also distinguish “alignment-based methods” from “k-mer-based classification”.

9. The restricted sepsis database may inflate apparent performance

The authors intentionally construct a curated microbial reference database containing organisms considered relevant to sepsis, partly to reduce misallocation of reads to large public databases.

This is understandable operationally, but it changes the scientific question.

A limited reference database will generally reduce the number of competing genomic sequences and may therefore improve apparent precision. At the same time, it risks missing unusual, emerging, or unexpected pathogens.

This is particularly relevant for immunocompromised patients, in whom uncommon organisms may be clinically important.

The manuscript should compare performance using both the curated database and a broader microbial reference database. This would clarify how much of the reported performance derives from the alignment algorithm and how much derives from restricting the search space.

10. Human reference choice may influence the apparent superiority of an aligner

The workflow uses T2T-CHM13 for human-read depletion. The authors acknowledge that a human pangenome such as the HPRC resource contains greater population and structural diversity.

This is not a minor detail. Human sequences poorly represented in a single reference are exactly the sequences that may remain after host depletion and later be misassigned to microbial genomes.

The apparent performance of a host-filtering algorithm therefore reflects both the algorithm and the host reference against which it operates.

A useful sensitivity experiment would compare GRCh38, T2T-CHM13, and a pangenome-based reference while keeping the alignment algorithm constant.

The current recommendation should consequently be framed as applying to BWA-MEM2 with the specific reference and filtering strategy evaluated here, rather than BWA-MEM2 universally.

11. Low-biomass sequencing requires stronger treatment of contamination

Plasma microbial cfDNA is an extremely low-biomass setting, making reagent and environmental contamination a major analytical concern.

The manuscript uses negative controls and filtering rules, which is appropriate. Nevertheless, conclusions about uncommon bacterial signals should demonstrate robustness to different contamination-removal criteria.

The authors should report the number and type of negative controls per batch, organisms identified in controls, batch-level contamination patterns, and the effect of alternative contamination thresholds.

Ideally, microbial abundance should also be evaluated relative to extraction blanks and library-preparation controls rather than relying primarily on presence/absence criteria.

12. Statistical uncertainty around performance estimates is insufficiently developed

Much of the manuscript relies on ranking algorithms by average read retention, residual human reads, precision, recall, F1 score, or number of samples in which a method ranks first.

Formal estimates of uncertainty would improve these comparisons.

Confidence intervals should be reported for the main performance measures. For comparisons based on the same samples, paired statistical analyses or bootstrap confidence intervals would be appropriate.

This is particularly important because the stated differences between algorithms can otherwise appear more definitive than the sample size supports.

13. Clinical claims should be separated from computational benchmarking

The manuscript concludes that the proposed guidelines may help clinicians diagnose sepsis and make informed antibiotic-treatment decisions.

That is a reasonable long-term objective, but the present study is primarily a bioinformatic benchmarking study.

It does not demonstrate that using BWA-MEM2 or Bowtie2 improves time to diagnosis, antimicrobial selection, patient outcomes, mortality, or antibiotic stewardship.

The Discussion and Conclusion should distinguish:

analytical performance of the computational pipeline;

diagnostic accuracy for infection;

prediction of antimicrobial resistance;

and clinical utility.

These are sequential but distinct levels of evidence.

Minor comments

The term “optimal” appears repeatedly and should generally be replaced by “best-performing among the evaluated methods under the tested conditions.”

Runtime and memory consumption should also be reported. A bioinformatic workflow proposed for clinical implementation should be judged not only by classification performance but also by computational requirements and turnaround time.

The complete pipeline, command-line parameters, database construction scripts, software versions, simulation parameters, filtering thresholds, and code used for performance calculations should be made publicly available. The raw FASTQ files are currently described as available upon reasonable request rather than deposited in a repository.

Finally, the claim that this is the first systematic evaluation of these approaches in plasma cfDNA should either be supported by a clearly described literature search or softened.

Conclusion

The manuscript addresses an important methodological problem and convincingly demonstrates that host-depletion strategy can substantially influence downstream cfDNA metagenomic results.

However, the current evidence is stronger for the conclusion that different pipelines produce different trade-offs than for the conclusion that BWA-MEM2 and Bowtie2 very-sensitive-local represent generally optimal solutions.

The most important limitation is the lack of independent biological ground truth in the clinical samples. In addition, the absolute taxonomic performance remains modest, the simulation framework is relatively favourable, the reference database is deliberately restricted, and AMR sequence detection is interpreted more clinically than the current validation supports.

The manuscript would be considerably strengthened by independent microbiological validation, more realistic out-of-reference simulations, formal sensitivity-specificity analyses, broader database testing, confidence intervals for performance estimates, and a more conservative distinction between analytical detection and clinical utility.

With these revisions, the study could provide a useful framework for standardising host depletion and downstream analysis in plasma cfDNA metagenomics, while more accurately defining the limits of what the current benchmarking data demonstrate.

Competing interests

The authors declare that they have no competing interests.

Use of Artificial Intelligence (AI)

The authors declare that they used generative AI to come up with new ideas for their review.

Puedes escribir un comentario en esta PREreview de Refining Cell-free DNA Metagenomics for Sepsis in Acute Leukemia: Towards Standardized Computational Strategies for Host Depletion and Microbial Detection from Shallow-Depth, Low-Coverage Profiles.

Antes de comenzar

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Comenzar ahora