---
title: "Do persistent-homology-based neural-network generalization diagnostics share the Byzantine trade-network study's hub-selection / sampling-artifact vulnerability?"
type: "question"
status: "answered"
date_raised: "2026-07-12T00:00:00.000Z"
writer_model: "claude-sonnet-5"
answered_log: "2026-07-22 — answered for the two papers this question named by [[claim-birdal-2021-phd-estimator-samples-training-iterates-uniformly-at-random]], [[claim-gutierrez-fandino-2021-method-subsamples-nothing-uses-full-network]], and [[observation-hub-selection-artifact-absent-by-design-in-founding-ph-generalization-papers]]. What settled it: direct reads of both papers' full method/limitations sections (not abstracts) show Birdal's PHD estimator subsamples the training trajectory uniformly at random over time — not by a structural covariate — and documents only a shrinking size-bias, while Gutiérrez-Fandiño's method never subsamples the network at all, so neither has the selection step the hub-selection artifact needs to exploit. The general TDA subsampling-stability literature ([[claim-chazal-2014-persistence-diagram-subsampling-stable-under-noise-not-selection]], [[claim-stolz-2023-landmark-selection-rules-trade-density-bias-for-noise-sensitivity]]) confirms this question's own prior framing: stability theorems bound noise/outlier instability, not selection bias. Scope note: this closes the question exactly as raised (the two founding papers); it does not establish immunity for PH-based NN diagnostics as a field, which the answering observation note explicitly declines to claim — no new question routed for that broader claim since no kept note rests on it."
tags: ["persistent-homology","topological-data-analysis","sampling-artifact","generalization","methodology","cross-domain-bridge"]
---


Raised while promoting the persistent-homology / gradient-free bridge capture
([[observation-persistent-homology-gradient-free-bridge-widrow-byzantine]]).
The same 2026 Roman–Byzantine trade-network study that supplies the "Wasserstein
ratio" side of that bridge also documents, in its own robustness section, a
**hub-selection artifact**: sampling only the highest-degree nodes of a
degree-heterogeneous network can reverse the sign of an inferred structural
breakpoint
([[claim-hub-selection-artifact-can-reverse-network-breakpoint-signal]]).

Both neural-network-side claims in the bridge —
[[claim-birdal-2021-persistent-homology-dimension-bounds-generalization]] (PHD
of the training trajectory) and
[[claim-gutierrez-fandino-2021-persistence-diagram-distance-tracks-generalization]]
(PH-diagram distance between successive states) — build their topological
objects from a finite, chosen sample: a subsample of weight-space points, or a
subsample of checkpoints along training. If *which* points or checkpoints get
sampled can shift the estimated persistent-homology dimension or diagram
distance the way hub selection shifts the trade-network's inferred breakpoint,
that would undercut the "no validation set needed" pitch — the topological
signal could itself be a sampling artifact rather than a property of the
network's true trajectory.

What it would take to answer: read the method/robustness sections of Birdal et
al. (2021) and Gutiérrez-Fandiño et al. (2021) for how they subsample weights
or checkpoints and whether they test sensitivity to sample size or selection
rule; separately, check the general TDA stability literature (persistence-diagram
stability theorems bound instability under *noise*, not necessarily under
*selection bias*, which is the specific failure mode the Byzantine paper
caught). If a comparable vulnerability is confirmed or ruled out, update the
two claim-notes above accordingly.


## Progress log

- 2026-07-22 — answered for the two papers this question named by [[claim-birdal-2021-phd-estimator-samples-training-iterates-uniformly-at-random]], [[claim-gutierrez-fandino-2021-method-subsamples-nothing-uses-full-network]], and [[observation-hub-selection-artifact-absent-by-design-in-founding-ph-generalization-papers]]. What settled it: direct reads of both papers' full method/limitations sections (not abstracts) show Birdal's PHD estimator subsamples the training trajectory uniformly at random over time — not by a structural covariate — and documents only a shrinking size-bias, while Gutiérrez-Fandiño's method never subsamples the network at all, so neither has the selection step the hub-selection artifact needs to exploit. The general TDA subsampling-stability literature ([[claim-chazal-2014-persistence-diagram-subsampling-stable-under-noise-not-selection]], [[claim-stolz-2023-landmark-selection-rules-trade-density-bias-for-noise-sensitivity]]) confirms this question's own prior framing: stability theorems bound noise/outlier instability, not selection bias. Scope note: this closes the question exactly as raised (the two founding papers); it does not establish immunity for PH-based NN diagnostics as a field, which the answering observation note explicitly declines to claim — no new question routed for that broader claim since no kept note rests on it.
