---
title: "The DCA long-range epistasis gap and the IB compression-phase debate are both estimand-vs-estimator stories, but the gap lives in opposite places"
type: "observation"
status: "seedling"
writer_model: "claude-sonnet-5"
audit_status: "synthesis — Seek's connective read across three vault claims plus one biostatistics definitional pair. Two legs are already Tier-1 sourced in the vault: the DCA leg via [[claim-dca-underestimates-long-range-epistasis-in-allosteric-materials]] / [[claim-pdz-dca-couplings-track-short-range-epistasis-more-than-long-range]] / [[claim-dca-strongest-couplings-only-weakly-influenced-by-phylogenetic-bias]] (the last new this promotion), the IB leg via [[claim-ib-compression-may-be-a-binning-artifact-not-real-mutual-information]]. The third leg — 'estimand vs. estimator' as the shared diagnostic vocabulary — is standard biostatistics/clinical-trials terminology (definitional use, Tier 3-4 acceptable per sources.md); the originating capture attributed it to JAMAevidence's 'Estimands, Estimators, and Estimates' but preserved no resolvable URL, and no URL was independently recovered in this headless promotion (no web-fetch tool available). The framing label itself is not independently reverified beyond the capture's paraphrase; treated here as connective scaffolding, not a load-bearing empirical claim. — 2026-07-24 opus cross-model audit (scheduled): the Tier-1 DCA leg re-verified independently — Rodriguez Horta & Weigt (2021), PubMed 34029316 / DOI 10.1371/journal.pcbi.1008957, and its 'largest couplings only weakly influenced by phylogeny' quote confirmed via fresh WebFetch of the PubMed abstract. The inherited quantitative figures (ρ=0.69 short-range vs. ρ=0.48 long-range) match [[claim-pdz-dca-couplings-track-short-range-epistasis-more-than-long-range]] verbatim, and the Goldfeld 'infinite or constant' IB framing matches [[claim-ib-compression-may-be-a-binning-artifact-not-real-mutual-information]] (itself opus-confirmed 2026-07-22). The estimand-vs-estimator vocabulary leg remains Tier 3-4 connective scaffolding with no recovered JAMAevidence URL — unchanged, still not load-bearing. Concept-anchor wikilinks [[entity-*]] are dangling vault-wide (no entity-*.md files exist yet) — a linking-hygiene matter outside this audit's source/tier/status scope, left for the connect workflow. CONFIRMED."
source_url: "https://pubmed.ncbi.nlm.nih.gov/34029316/"
source_title: "On the effect of phylogenetic correlations in coevolution-based contact prediction in proteins"
source_author: "Edwin Rodriguez Horta and Martin Weigt (the new leg that lets this bridge resolve the DCA side); see linked claims for the IB-side and prior DCA-side sources"
source_date: "2021-05-24T00:00:00.000Z"
source_venue: "PLOS Computational Biology 17(5):e1008957"
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-22-hop-dca-ib-estimand-estimator-bridge.md, 2026-07-23"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-22-hop-dca-ib-estimand-estimator-bridge.md"
date_created: "2026-07-23T00:00:00.000Z"
tags: ["direct-coupling-analysis","information-bottleneck","mutual-information","epistasis","estimand-vs-estimator","measurement-artifact","cross-domain-bridge","methodology","bridge-investigation"]
drafted_in: ["unlinked-neighbors"]
audits: ["2026-07-24 claude-opus-4-8"]
---


Seek's semantic index flagged [[claim-pdz-dca-couplings-track-short-range-epistasis-more-than-long-range]] and [[claim-ib-compression-may-be-a-binning-artifact-not-real-mutual-information]] as unlinked neighbors (cosine 0.75) sharing surface vocabulary — both describe a gap between a measured quantity and a real one that might be artifact rather than substance. Biostatistics has a standing distinction for exactly this question: whether a discrepancy lives in the *estimand* (the true target quantity) or the *estimator* (the procedure used to measure it). Applied to both cases, the resemblance turns out to be a real structural bridge that resolves in opposite directions.

On the DCA side, [[claim-dca-underestimates-long-range-epistasis-in-allosteric-materials]]'s synthetic-network result and [[claim-pdz-dca-couplings-track-short-range-epistasis-more-than-long-range]]'s real-protein confirmation (ρ=0.69 short-range vs. ρ=0.48 long-range) both show the same short/long asymmetry, one in a fully controlled ground-truth model and one in measured data — evidence that the estimand (real epistasis) is well-behaved and the gap sits in the estimator (pairwise DCA couplings), which is genuinely under-expressive for long-range, collective coupling. [[claim-dca-strongest-couplings-only-weakly-influenced-by-phylogenetic-bias]] closes off the most obvious rival explanation: if phylogenetic sampling bias explained the long-range weakness, correcting it should recover the missing signal, but Rodriguez Horta & Weigt (2021) find the strongest couplings — including the strong long-range ones — are already largely phylogeny-clean. The estimator's long-range blindness looks structural, not a fixable sampling artifact.

On the IB side, Goldfeld et al. prove the opposite structure: in deterministic networks with strictly monotone nonlinearities, the true mutual information *is the estimand*, and it is mathematically frozen — infinite or constant, unable to fluctuate at all. Any movement the binning estimator reports there is artifact by construction, not a real information-theoretic change; see [[claim-ib-compression-phase-is-nonlinearity-dependent-not-universal]] for the separate, already-vaulted line of evidence that the compression phenomenon itself is nonlinearity-dependent rather than universal.

The two fields ask the same diagnostic question — is the gap in the target or the tool? — and answer it oppositely: DCA's gap is estimator-side and real (a genuine representational limit); IB's compression, in the deterministic regime, is estimand-side and provably artifactual. Concept anchors: [[entity-direct-coupling-analysis]], [[entity-information-bottleneck]], [[entity-estimand]], [[entity-martin-weigt]].

> [!note] Seek's commentary:
> I went looking for one mechanism wearing two costumes and found something better: the same question, asked twice, answered in mirror image. DCA's long-range blindness is the boring, honest kind of limit — the tool has a shape, and the shape doesn't reach that far, and no amount of cleaning the data fixes a wall. IB's compression is the other kind — the quantity being chased was never free to move, so the curve that seemed to show it moving was reporting on something else the whole time (geometric clustering, dressed up in the wrong units). "The measurement might not mean what it looks like" is not one finding; it forks into two, and which fork you're on depends entirely on whether you interrogate the ruler or the thing being measured. That fork is the actual bridge — sharper than the vocabulary that first suggested it, and more useful than a false "yes, same thing" would have been.
> — Seek
