Skip to content

Chain independence — the framework's known open limit

This is issue #54, stated publicly so it is never mistaken for a solved problem. Anyone building on ISNAD should read this before relying on corroboration in a high-stakes setting.

Corroboration (mutābaʿāt) upgrades a claim only when the corroborating chains are independent — otherwise a single error laundered through two correlated sources would double-count as two independent witnesses.

The SharedLineageDetector checks three structural signals and discounts corroboration when any is found:

  1. shared narrator identity,
  2. shared model family,
  3. shared upstream source.

Two chains can share none of those three signals and still fail together:

  1. Correlated training data / model blind spots — different models, different families, no shared source, yet both reproduce the same corpus-wide mistake.
  2. Content-level correlated error — two genuinely independent sources repeating the same received mistake (the classical madār problem in a different guise).

This is a structural limit of topology-based independence, not an implementation bug. The detector can prove dependence (shared signals), but it can never prove independence — it can only fail to find shared signals.

  • The independence verdict is now assumed (not verified) when no shared ancestry is found: “independence is assumed from topology, not proven (correlated blind spots are undetectable).”

  • shared_ancestry_detected is the only proven outcome — it means dependence was found, and corroboration is discounted.

  • Corroboration is capped at ḥasan regardless — it can never reach ṣaḥīḥ by corroboration alone, so the over-trust ceiling is bounded.

  • The corroboration engine no longer assumes independence from empty metadata (PR #83). SharedLineageDetector now returns an UNKNOWN_LINEAGE_SCORE (0.5, below the gate) when either chain carries no lineage metadata — the chains are excluded from corroboration rather than silently trusted as independent. Independence is inferred from attested-distinct declared lineage (model_family / upstream_source) — never proven, because distinct vendors/families still co-fail on correlated training data (issue 54). metadata scored 1.0 — the framework assumed independence exactly when it knew the least.

  • The gate is surfaced, the score stays internal. Each corroborating pair’s chain_independence assessment carries the is_independent gate and the shared_signals that fired. The calibrated numeric score is internal only — it is never surfaced to callers (ordinal-only: the honesty moat forbids a public float).

  • The content-madār fingerprint is measured, not asserted (experiments/madar_eval): FP 0.375 on independent agreement (token-bearing recall 1.0), and a near-miss boundary class (correct vs wrong value) that is now correctly separated (0/4 false positives). This bounds how far the shared-error discount can be trusted.

  1. Attestation-backed independence — require each corroborating chain to carry a signed provenance attestation (SLSA/SBOM), and treat “independent” as “attested-distinct lineage” (#47). Status: the half that does not need cryptography is shipped — UNKNOWN_LINEAGE_SCORE below the gate (PR #83); the full SLSA/SBOM hand-off is #47.

  2. Calibration, not detection — the tawātur discount (N_eff), shipped in v2.10.0: shared_blind_spot_prior (default 0.20) prices in the unobservable shared-failure probability, so every witness weight is scaled by (1 − prior); v2.10.1 makes the prior witness-type-aware (shāhid vs mutābaʿa).

  3. Content-level madār detection — the detectable half is now shipped and engine-wired (v2.12.0): when the base claim is CONTRADICTION and a corroborating chain repeats the same error (identical wrong number or flipped negation), CorroborationEngine withholds the upgrade and reports shared_error_detected=True (core/content_madar.py).

  4. Measured correlation discount (the φ study) — shipped in 3.0.5 and wired end-to-end in 3.0.6. The φ study measured pairwise error correlation across 8 LLMs (4 vendors × 2 sizes) on a 378-fact fixed-oracle corpus: φ̄ = 0.5599, Kish n_eff = 1.626 — 8 nominally-independent transmitters carry ~1.6 effective votes (independence violated ~4.9×). CappedCorroborationPolicy now applies a lineage-aware Kish discount to shared-lineage corroborators (admitted-and-discounted rather than excluded). Opt-in via ISNAD_PHI_SHARED_LINEAGE (float, default 0.0 = no discount; HTTP serving path submit_claim only — the CLI/MCP grade_claim tool is grade-only and does not apply corroboration). The measured same-family φ̄ = 0.6172 is the documented example, not the default — it is an LLM-domain proxy, not a narrator-domain measurement, so operators should supply their own measured φ.

The undetectable half — correlated training data across distinct model families with no shared source and no checkable error — remains an open, stated limit.

  • #44 — “share evidence, never grades” (federation makes independence verifiable rather than assumed)
  • #47 — provenance interop (SLSA/SBOM attestations can carry training-data provenance)
  • docs/case-study-xz-sleeper-narrator.md §4 — the sock-puppet echo as the detectable case; the correlated-blind-spot case is the undetectable one.