Chain independence — the framework's known open limit
This is issue #54, stated publicly so it is never mistaken for a solved problem. Anyone building on ISNAD should read this before relying on corroboration in a high-stakes setting.
What corroboration assumes
Section titled “What corroboration assumes”Corroboration (mutābaʿāt) upgrades a claim only when the corroborating chains are independent — otherwise a single error laundered through two correlated sources would double-count as two independent witnesses.
The SharedLineageDetector checks three structural signals and discounts
corroboration when any is found:
- shared narrator identity,
- shared model family,
- shared upstream source.
What topology cannot prove
Section titled “What topology cannot prove”Two chains can share none of those three signals and still fail together:
- Correlated training data / model blind spots — different models, different families, no shared source, yet both reproduce the same corpus-wide mistake.
- Content-level correlated error — two genuinely independent sources repeating the same received mistake (the classical madār problem in a different guise).
This is a structural limit of topology-based independence, not an implementation bug. The detector can prove dependence (shared signals), but it can never prove independence — it can only fail to find shared signals.
What ISNAD does about it
Section titled “What ISNAD does about it”-
The independence verdict is now
assumed(notverified) when no shared ancestry is found: “independence is assumed from topology, not proven (correlated blind spots are undetectable).” -
shared_ancestry_detectedis the only proven outcome — it means dependence was found, and corroboration is discounted. -
Corroboration is capped at ḥasan regardless — it can never reach ṣaḥīḥ by corroboration alone, so the over-trust ceiling is bounded.
-
The corroboration engine no longer assumes independence from empty metadata (PR #83).
SharedLineageDetectornow returns anUNKNOWN_LINEAGE_SCORE(0.5, below the gate) when either chain carries no lineage metadata — the chains are excluded from corroboration rather than silently trusted as independent. Independence is inferred from attested-distinct declared lineage (model_family/upstream_source) — never proven, because distinct vendors/families still co-fail on correlated training data (issue 54). metadata scored1.0— the framework assumed independence exactly when it knew the least. -
The gate is surfaced, the score stays internal. Each corroborating pair’s
chain_independenceassessment carries theis_independentgate and theshared_signalsthat fired. The calibrated numeric score is internal only — it is never surfaced to callers (ordinal-only: the honesty moat forbids a public float). -
The content-madār fingerprint is measured, not asserted (
experiments/madar_eval): FP 0.375 on independent agreement (token-bearing recall 1.0), and a near-miss boundary class (correct vs wrong value) that is now correctly separated (0/4 false positives). This bounds how far the shared-error discount can be trusted.
Candidate approaches (now shipped)
Section titled “Candidate approaches (now shipped)”-
Attestation-backed independence — require each corroborating chain to carry a signed provenance attestation (SLSA/SBOM), and treat “independent” as “attested-distinct lineage” (#47). Status: the half that does not need cryptography is shipped —
UNKNOWN_LINEAGE_SCOREbelow the gate (PR #83); the full SLSA/SBOM hand-off is #47. -
Calibration, not detection — the tawātur discount (N_eff), shipped in v2.10.0:
shared_blind_spot_prior(default 0.20) prices in the unobservable shared-failure probability, so every witness weight is scaled by (1 − prior); v2.10.1 makes the prior witness-type-aware (shāhid vs mutābaʿa). -
Content-level madār detection — the detectable half is now shipped and engine-wired (v2.12.0): when the base claim is CONTRADICTION and a corroborating chain repeats the same error (identical wrong number or flipped negation),
CorroborationEnginewithholds the upgrade and reportsshared_error_detected=True(core/content_madar.py). -
Measured correlation discount (the φ study) — shipped in 3.0.5 and wired end-to-end in 3.0.6. The φ study measured pairwise error correlation across 8 LLMs (4 vendors × 2 sizes) on a 378-fact fixed-oracle corpus: φ̄ = 0.5599, Kish n_eff = 1.626 — 8 nominally-independent transmitters carry ~1.6 effective votes (independence violated ~4.9×).
CappedCorroborationPolicynow applies a lineage-aware Kish discount to shared-lineage corroborators (admitted-and-discounted rather than excluded). Opt-in viaISNAD_PHI_SHARED_LINEAGE(float, default0.0= no discount; HTTP serving pathsubmit_claimonly — the CLI/MCPgrade_claimtool is grade-only and does not apply corroboration). The measured same-family φ̄ = 0.6172 is the documented example, not the default — it is an LLM-domain proxy, not a narrator-domain measurement, so operators should supply their own measured φ.
The undetectable half — correlated training data across distinct model families with no shared source and no checkable error — remains an open, stated limit.
Related
Section titled “Related”- #44 — “share evidence, never grades” (federation makes independence verifiable rather than assumed)
- #47 — provenance interop (SLSA/SBOM attestations can carry training-data provenance)
docs/case-study-xz-sleeper-narrator.md§4 — the sock-puppet echo as the detectable case; the correlated-blind-spot case is the undetectable one.