When Does Higher-Order Representation Improve Protein-Complex Recovery? A Leakage-Controlled, Matched-k Comparison of Graph and Penalized Hypergraph Spectral Clustering
- Posted
- Server
- Preprints.org
- DOI
- 10.20944/preprints202608.0397.v1
Group-level molecular-interaction evidence is often reduced to pairwise protein-protein interaction graphs before protein-complex detection, potentially discarding experiment membership and over-weighting large groups. Yet prior higher-order studies typically changed the representation, objective, output granularity, and evaluation protocol simultaneously. We isolate the operator-level contribution of higher-order representation by comparing an inverse-size-weighted clique graph, the normalized Zhou hypergraph operator, and a degree-aware penalized hypergraph operator on identical experiment-derived groups. Hyperedges were split by size into training, validation, and test sets; the training split alone defined the protein universe and operators; regularization was selected on validation data; and recovery was evaluated against held-out test hyperedges. Across nine IntAct-derived thematic PSI–MITAB collections and 25 paired initializations per dataset, the penalized hypergraph increased mean held-out symmetric best-match F1 on six datasets after within-dataset Benjamini–Hochberg correction, was indistinguishable on two, and was lower by 0.0015 on Cancer. Positive mean differences ranged from 0.0098 to 0.0356. However, the across-dataset Friedman test did not reach significance (χ2 = 5.20, p = 0.074), and the graph–penalized-hypergraph Nemenyi comparison was non-significant (p = 0.111). The penalized operator had the best mean rank (1.39), but secondary metrics and eigengap-selected cluster counts exposed dataset-dependent trade-offs. Observed recovery exceeded degree-preserving and hyperedge-size-preserving nulls on eight of nine datasets. Gene Ontology coherence was high for both representations, with no aspect-level difference surviving multiplicity correction and substantial pooled term overlap (Jaccard 0.78–0.86). Higher-order modeling is therefore conditionally beneficial rather than universally superior: it is most defensible when group membership is retained, regularization is validation-supported, and output granularity is controlled