---
title: "Stolz (2023) shows the two standard TDA landmark-selection rules trade density bias for noise sensitivity, neither addressing structural-selection robustness"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://www.jmlr.org/papers/v24/21-1526.html"
source_title: "Outlier-Robust Subsampling Techniques for Persistent Homology"
source_author: "Bernadette J. Stolz"
source_date: "2023-02"
source_quote: "random selection tends to favour dense areas of the data while the maxmin algorithm is very sensitive to noise"
source_tier: 1
audit_status: "capture-verified — full PDF fetched and read directly via extract_pdf at capture time (jmlr.org/papers/volume24/21-1526/21-1526.pdf, tls:verified); not independently re-fetched at this promotion pass."
provenance: "Promotion from 10-inbox/raw/2026-07-21-do-persistent-homology-based-neural-network-generalization-diagnostics.md, 2026-07-22 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-21-do-persistent-homology-based-neural-network-generalization-diagnostics.md"
date_created: "2026-07-22T00:00:00.000Z"
tags: ["persistent-homology","topological-data-analysis","subsampling","sampling-artifact","methodology","landmark-selection"]
---


Bernadette J. Stolz, "Outlier-Robust Subsampling Techniques for Persistent
Homology" (JMLR 24, 2023), surveys the two standard rules used to subsample a
point cloud before computing persistent homology on it: uniform random
selection and the maxmin (farthest-point) algorithm. Her own framing of their
respective failure modes, stated verbatim: "random selection tends to favour
dense areas of the data while the maxmin algorithm is very sensitive to
noise." She proposes a topology-preserving alternative explicitly evaluated
on "robustness to outliers" as its own criterion.

Both standard rules, and Stolz's proposed fix, are organized around the same
axis as [[claim-chazal-2014-persistence-diagram-subsampling-stable-under-noise-not-selection]]:
density bias and noise/outlier sensitivity in an otherwise-unstructured point
cloud, not selection by a structural covariate of a *network* (e.g.
node degree). This is
[[entity-landmark-selection-topological-data-analysis|landmark selection]],
the general TDA-field name for the choice this capture's founding question —
[[question-tda-neural-net-sampling-artifact-risk]] — is really asking about,
one level up from any single paper's method.

Neither landmark-selection rule surveyed here has an analogue to the
Byzantine trade-network study's hub-selection artifact
([[claim-hub-selection-artifact-can-reverse-network-breakpoint-signal]]),
which is specifically about sampling by a *structural* property of a network
rather than density or noise in an unstructured point cloud. See
[[observation-hub-selection-artifact-absent-by-design-in-founding-ph-generalization-papers]]
for the fuller synthesis.
