API Reference
API Reference
Section titled “API Reference”Auto-generated from isnad v3.0.7. Every public symbol, its signature, and its docstring.
pip install isnadclass Chain
Section titled “class Chain”Chain(links: 'list[ChainLinkSpec] | None' = None, *, retrieved_rows_known: 'bool' = True)An ordered, gap-checked transmission chain for a claim.
class ChainLinkSpec
Section titled “class ChainLinkSpec”ChainLinkSpec(narrator_id: 'str', step: 'int', *, version: 'str' = 'unknown', transform_type: 'TransformType' = <TransformType.PASS_THROUGH: 'pass_through'>, trace_id: 'str' = '', domain: 'str' = 'general', confidence: 'float | None' = None, input_snapshot: 'str | None' = None, output_snapshot: 'str | None' = None, document_hashes: 'list[str] | None' = None, timestamp: 'str | None' = None, retrieved_rows: 'list[str] | None' = None)Specification for a single chain link — used to build chains.
adalah_grades_for_chain()
Section titled “adalah_grades_for_chain()”adalah_grades_for_chain(registry: 'Registry', chain: 'Chain') -> 'list[AdalahGrade]'Look up per-link ʿadālah (integrity) grades, mirroring grades_for_chain().
Kept as a separate axis (issue #11) rather than folded into NarratorGrade — a narrator’s precision/accuracy track record and its integrity status are distinct properties (paper §4.2) and should be checked independently.
grade_chain_from_registry()
Section titled “grade_chain_from_registry()”grade_chain_from_registry(registry: 'Registry', chain: 'Chain', *, corroboration_support: 'bool' = False, link_fidelity_verdicts: 'list[ContentVerdict] | None' = None, lenient_unknown: 'bool' = False) -> 'ChainGrade'Grade a chain from a registry, always threading the integrity axis.
The single canonical grading entry point. grade_chain keys the
mawdu / daif-jiddan split on the integrity axis (COMPROMISED -> MAWDU),
so every path that grades a chain MUST supply link_adalah_grades.
This helper composes the registry lookups and the grading call so no
integration can silently skip integrity and diverge from the decision
the API serves (the 3.0.0 regression).
grades_for_chain()
Section titled “grades_for_chain()”grades_for_chain(registry: 'Registry', chain: 'Chain') -> 'list[NarratorGrade]'Look up per-link narrator grades using version-aware registry keys.
make_claim_id()
Section titled “make_claim_id()”make_claim_id(text: 'str') -> 'str'Normalize + hash claim text to produce a deterministic claim_id.
normalize_claim_text()
Section titled “normalize_claim_text()”normalize_claim_text(text: 'str') -> 'str'Normalize claim text for hashing and comparison.
Normalization: lowercase, strip, collapse whitespace. This is deliberately simple — a production system would add domain-specific normalization (unit conversion, formula canonicalization).
One-call convenience
Section titled “One-call convenience”class Verdict
Section titled “class Verdict”Verdict(chain_grade: 'ChainGrade', action: 'Action | None', weakest_link: 'str', why: 'str') -> NoneThe result of grade: a chain grade, an action, and a why.
grade()
Section titled “grade()”grade(claim: 'str', chain: 'list[str]', registry: 'Registry', *, domain: 'str' = 'general', transform_types: 'list[TransformType] | None' = None) -> 'Verdict'Grade a claim’s transmission chain and decide an action.
Args:
claim: The claim text (used for the why string; content criticism is
not run here — pass a critic’s verdict through decide if you
want the full 5×3 matrix).
chain: Ordered narrator ids (source first, answer last).
registry: The graded-narrator registry.
domain: Domain tag for grading.
transform_types: Optional per-link transform types; defaults to
PASS_THROUGH (identity).
Returns:
A Verdict. action is None when no content verdict is
supplied — chain grading alone does not decide serve/review/quarantine.
Identity & versioning
Section titled “Identity & versioning”is_unknown_version()
Section titled “is_unknown_version()”is_unknown_version(version: 'str | None') -> 'bool'Return True when version should fall back to legacy alias-only lookup.
Covers both “”/“unknown” (legacy) and “latest”/“dev”/“canary” (non-resolved tags that silently drift).
parse_narrator_id()
Section titled “parse_narrator_id()”parse_narrator_id(resolved_id: 'str') -> 'tuple[str, str | None]'Split a resolved registry id into (alias, version).
resolve_narrator_id()
Section titled “resolve_narrator_id()”resolve_narrator_id(narrator_id: 'str', version: 'str | None') -> 'str'Build the registry key for a chain link.
When version is known, returns {narrator_id}@{version} unless the id
is already versioned. Otherwise returns the alias unchanged.
Corroboration
Section titled “Corroboration”class CappedCorroborationPolicy
Section titled “class CappedCorroborationPolicy”CappedCorroborationPolicy(*, shared_blind_spot_prior: 'float' = 0.2, measured_priors: 'dict[tuple[str | None, str | None], float] | None' = None, joint_failure: 'bool' = False, phi_shared_lineage: 'float' = 0.0) -> 'None'Default corroboration policy: information-theoretic, capped, minimum-gated.
This is one instantiation of a parameter the framework leaves open (see paper §4.3). Swap freely.
Uses HadithRank-style information-theoretic corroboration: multiple independent transmission chains asserting the same claim reduce the combined error probability multiplicatively.
Rules:
- Base grade is the grade of the chain under evaluation.
- Corroboration can upgrade at most one tier.
- Corroboration can never reach SAHIH (automatically capped).
- At least one independent corroborating chain must be HASAN or above (minimum gate).
- Soft shared-lineage chains (score > 0, flagged by the detector) are admitted but discounted as a group by the Kish scale; hard identity (score 0.0) and unknown lineage (0.5) are excluded.
- Chains above the threshold earn credit in proportion to their independence score (the disjointness discount, #125 tier 1): a chain at 0.85 contributes 85% of a full chain’s log-error reduction.
- The tawātur discount (#54). Even a chain with independence score 1.0
is not a full independent witness, because two agents can share an
unobservable correlated failure (shared training data, shared blind
spot) that no topology check can see. Classical tawātur never tried to
prove per-pair independence; it required a number of reporters large
enough that collusion on a falsehood was inconceivable. The computational
translation is an explicit prior
shared_blind_spot_prior— the probability that a nominally-independent chain shares a failure mode with the base chain — which discounts every chain’s witness weight. Default is conservatively high (0.20); an operator who can attest distinct lineage (#47) or exchange evidence (#44) lowers it. - All independent chains (including DAIF) contribute to the combined error reduction; even weak corroboration adds weight.
- The combined log-error ratio must reach MIN_EFFECTIVE_WEIGHT for an upgrade to fire.
class CorroborationEngine
Section titled “class CorroborationEngine”CorroborationEngine(min_independent_chains: 'int' = 1, min_gate_grade: 'ChainGrade' = <ChainGrade.HASAN: 'hasan'>, correlation_detector: 'SharedLineageDetector | None' = None, policy: 'CappedCorroborationPolicy | None' = None)Engine for cross-claim corroboration (mutābaʿāt).
Finds independent chains for a given claim and applies the information-theoretic corroboration upgrade via CappedCorroborationPolicy.
Usage: engine = CorroborationEngine(min_independent_chains=1) result = engine.evaluate( claim_text=“F = ma”, base_chain_grade=ChainGrade.DAIF, base_narrators=[“source:A”, “scraper:v1”, “model:gpt4”], all_chains=all_claim_chains, narrator_metadata=narrator_metadata, ) if result.upgraded: print(f“Upgraded from {result.base_grade.value} “ f”to {result.upgraded_grade.value}“)
class SharedLineageDetector
Section titled “class SharedLineageDetector”SharedLineageDetector()Default correlation detector: checks shared model family and upstream source.
This is one instantiation of a parameter the framework leaves open (see paper §4.3, §7). Swap freely.
A reference stub: the heuristics here (exact match on model_family, substring match on upstream source) are deliberately simple. A production version would incorporate structured model lineage data (e.g., model cards, training-data provenance) and possibly embedding similarity of model outputs to detect correlated blind spots.
Independence must be demonstrated, not assumed (issue #54). Topology can only ever falsify independence (by finding a shared signal), never prove it — the paper’s §7 concedes that a correlated pair sharing none of the three signals, and content-level correlated error, are undetectable here. So the detector distinguishes three cases, not two:
- Known-distinct lineage — both chains carry lineage metadata and it differs → high independence (earned).
- Unknown lineage — no metadata to reason about →
UNKNOWN_LINEAGE_SCORE, deliberately below the corroboration gate. Absent evidence of independence, the framework no longer assumes it (previously this returned 1.0 — “most independent when we know least”, the exact inversion issue #54 warns against). - Shared signal detected → penalised toward 0.
This is an honesty calibration, not a detection mechanism: it narrows the independence assumption, it does not claim to detect correlated blind spots.
Sybil note (issue #196): model_family/upstream_source are self-declared
caller metadata, not authenticated lineage — an adversary can spoof them to
earn a high independence score. This detector narrows the independence
assumption; it does not authenticate lineage (verified SLSA/SBOM attestation
is the 3.0 evidence model, issue #188).
evaluate_corroboration()
Section titled “evaluate_corroboration()”evaluate_corroboration(base_grade: 'ChainGrade', corroborating_chain_grades: 'list[ChainGrade]', base_narrators: 'list[str]', corroborating_narrators: 'list[list[str]]', narrator_metadata: 'dict[str, dict[str, object]]', *, policy: 'CorroborationPolicy | None' = None, detector: 'CorrelationDetector | None' = None) -> 'ChainGrade'Evaluate corroboration for a claim.
Args: base_grade: The chain grade of the claim under evaluation. corroborating_chain_grades: Grades of corroborating chains. base_narrators: Narrator IDs in the base claim’s chain. corroborating_narrators: Narrator IDs for each corroborating chain. narrator_metadata: Metadata dict for correlation detection. policy: Optional custom CorroborationPolicy. detector: Optional custom CorrelationDetector.
Returns: The (possibly upgraded) ChainGrade.
Decision
Section titled “Decision”decide()
Section titled “decide()”decide(chain_grade: 'ChainGrade', content_verdict: 'ContentVerdict') -> 'Action'Route a (chain_grade, content_verdict) pair to the correct action.
This is the decision matrix from paper §4.4, Table. It combines transmission quality (chain grade) with content quality (matn criticism) into a concrete serve/review/quarantine action.
Args: chain_grade: The ordinal chain grade (SAHIH/HASAN/DAIF/DAIF_JIDDAN/MAWDU). content_verdict: The content criticism verdict (CONSISTENT/CONTRADICTION/UNVERIFIABLE).
Returns: The Action to take.
Raises: KeyError: If an unexpected grade/verdict combination is passed.
describe_action()
Section titled “describe_action()”describe_action(chain_grade: 'ChainGrade', content_verdict: 'ContentVerdict') -> 'str'Return a human-readable description of what the matrix decided and why.
Useful for logging, review-queue entries, and debugging.
gate_serve()
Section titled “gate_serve()”gate_serve(action: 'Action', prior_only_narrators: 'list[str] | None', *, hold: 'bool' = False) -> 'Action'Cap a serve action when the chain rests on prior-only narrators (P0-B).
A narrator graded from a population prior alone (a benchmark seed with zero observed in-pipeline instances) is an unvalidated assumption, however confident the prior looks (issue #6). Classical rijāl graded on observed instances, never priors — a reputation-only transmitter is majhūl al-ḥāl, not thiqa. So a chain through a prior-only narrator must not be served as if its transmission had been observed.
This gates the action, never the chain grade: the grade stays what the weakest-link evaluation computed (a SAHIH chain through a seeded narrator is still SAHIH — we just don’t plain-SERVE it until someone observes the transmitter in this pipeline).
Args:
action: The matrix action (from decide).
prior_only_narrators: Narrator ids in the serving chain whose grade is
prior-only. Empty/None means no gate applies.
hold: When True (hard gate — high-stakes domains), a prior-only chain
downgrades any serve to REVIEW. When False (default soft gate), a
plain SERVE is capped to SERVE_WITH_CAVEAT.
Returns: The (possibly capped) action. REVIEW/QUARANTINE/REJECT are never touched — the gate only demotes serve, never upgrades anything.
Grading
Section titled “Grading”class RefinedWeakestLink
Section titled “class RefinedWeakestLink”RefinedWeakestLink()Default grading strategy: refined weakest-link with completeness cap.
This is one instantiation of a parameter the framework leaves open (see paper §4.2/§4.3). Swap freely.
The algorithm walks the transmission chain link-by-link, maintaining a running floor that represents the best grade the chain can achieve after each link:
-
Destructive (extraction, chunking, lossy summarization): the link’s grade becomes a hard floor. Nothing downstream recovers lost info.
-
Generative (broad-pretrained model synthesis) with corroboration: the floor is lifted toward the link’s own grade, capped at HASAN (hasan li-ghayrihi) and never above a permanent destructive floor - corroboration cannot recover a destructive loss. Only fires when the generative link is ACCEPTABLE or better; WEAK generative always degrades.
-
Generative without corroboration, or pass-through: standard minimum.
-
Incomplete chain → DAIF. COMPROMISED integrity → MAWDU (fabricated); a REJECTED narrator with SUSPECT integrity → DAIF_JIDDAN (very weak).
Note: the mawḍūʿ / ḍaʿīf-jiddan split follows Ibn Ḥajar (Nuzhat al-Naẓar): a narrator’s lying makes a narration mawḍūʿ (fabricated), being accused of lying makes it matrūk (abandoned → very weak). ISNAD keys the split on the integrity axis — COMPROMISED integrity → MAWDU, SUSPECT (or no) integrity → DAIF_JIDDAN — not on NarratorGrade alone. Both still quarantine.
grade_chain()
Section titled “grade_chain()”grade_chain(link_narrator_grades: 'list[NarratorGrade]', link_transform_types: 'list[TransformType]', is_complete: 'bool', *, strategy: 'GradingStrategy | None' = None, corroboration_support: 'bool' = False, link_adalah_grades: 'list[AdalahGrade]', link_fidelity_verdicts: 'list[ContentVerdict] | None' = None, lenient_unknown: 'bool' = False) -> 'ChainGrade'Grade a claim chain.
Args: link_narrator_grades: Per-link narrator grades. link_transform_types: Per-link transform types. is_complete: Chain completeness (ittiṣāl). strategy: Optional custom GradingStrategy. corroboration_support: Whether corroboration supports the claim. link_adalah_grades: Per-link ʿadālah (integrity) grades (required) — see RefinedWeakestLink.compute_chain_grade for details. link_fidelity_verdicts: Optional per-link transformation-fidelity verdicts (core/fidelity.py) — see RefinedWeakestLink.compute_chain_grade for details. lenient_unknown: Treat UNGRADED narrators as ḥasan (lenient, opt-in) instead of ḍaʿīf (strict, the default).
Returns: ChainGrade for the claim.
class BayesianTransitionPolicy
Section titled “class BayesianTransitionPolicy”BayesianTransitionPolicy(integrity_strikes_per_tier: 'int' = 1, *, reliable_threshold: 'float' = 0.9, acceptable_threshold: 'float' = 0.75, weak_threshold: 'float' = 0.6)Bayesian transition policy using Beta distribution updates.
This is one instantiation of a parameter the framework leaves open (see paper §4.2). Swap freely.
Each narrator×domain maintains a Beta(α, β) state. Evidence updates the posterior. Grades are derived from the posterior mean with calibrated thresholds.
Key advantages over threshold counting:
- Continuous confidence (posterior mean + credible interval)
- Graceful with small samples (prior provides regularization)
- Natural uncertainty quantification
- No arbitrary “3 adverse events” cutoff
The axis split (issues #9/#20) applies here too: integrity (ʿadālah) jarḥ
accumulates permanently and imposes a ceiling (_integrity_cap), while
precision (ḍabṭ) jarḥ feeds the recoverable Beta posterior. The posterior
grade is clamped to the integrity ceiling, so a narrator whose integrity
is impugned can never be lifted back up by precision evidence alone.
integrity_strikes_per_tier (issue #21, configurable) sets how many
permanent integrity strikes lower the ceiling one tier. Default 1 — the
classical matrūk bias: one proven integrity impugnment caps immediately.
class CalibratedThresholdPolicy
Section titled “class CalibratedThresholdPolicy”CalibratedThresholdPolicy(downgrade_threshold: 'int' = 5, upgrade_sustained_count: 'int' = 10, upgrade_min_corroborated: 'int' = 5, window: 'int | None' = None, integrity_strikes_per_tier: 'int | None' = None)Threshold-based policy with thresholds LEARNED from calibration data.
This is one instantiation of a parameter the framework leaves open (see paper §4.2). Swap freely.
Rather than hardcoding (3 adverse, 5 positive), these thresholds are calibrated from historical performance data via the §8 experiment methodology. The thresholds can be set per-domain and per-narrator-type.
Shares the sliding-window + edge-trigger ratchet fix of
ThresholdTransitionPolicy (issue #9 finding #1): counts are taken over
the last window evidence entries, and a transition fires only on the
kind of evidence that just arrived. The window defaults to
max(downgrade_threshold, upgrade_sustained_count) so both branches stay
reachable whatever the calibrated thresholds are.
class ThresholdTransitionPolicy
Section titled “class ThresholdTransitionPolicy”ThresholdTransitionPolicy(window: 'int | None' = None, integrity_strikes_per_tier: 'int | None' = None)Threshold jarḥ–taʿdīl transition policy over a sliding evidence window.
This is one instantiation of a parameter the framework leaves open
(see paper §4.2). Swap freely. The framework default is
BayesianTransitionPolicy; this policy is retained as a simple,
interpretable alternative.
Rules:
- Downgrade fires when adverse evidence within the recent window crosses a threshold.
- Upgrade requires sustained corroborated accuracy within the recent window (N positive evals).
- Version bump resets to UNGRADED.
- ʿAdālah COMPROMISED → REJECTED (active containment).
- REJECTED is sticky — requires explicit human review to restore.
Ratchet fix (issue #9 finding #1) — sliding window + edge trigger. The original policy counted adverse (jarḥ) evidence over the narrator’s entire history and checked the downgrade branch on every call. Two independent defects combined into a ratchet: (a) the adverse count never decayed, so it was monotonic; and (b) the downgrade was level-triggered, firing on every subsequent call — including pure taʿdīl — while the count stayed above threshold. Together they marched every narrator RELIABLE → ACCEPTABLE → WEAK → REJECTED with no path to recovery, and left the upgrade branch unreachable.
Both defects are fixed here:
- Sliding window. Adverse and favorable counts are taken over only the
last
windowevidence entries, so stale jarḥ ages out and sustained good behaviour can recover a narrator. - Edge trigger. A downgrade fires only when the arriving evidence is a jarḥ, and an upgrade only when the arriving evidence is a taʿdīl. A transition is driven by the evidence that just arrived, not by counts left standing in the window from earlier calls.
A narrator that keeps producing adverse evidence still reaches REJECTED, so active containment is preserved.
Axis split (issue #9 conceptual follow-up) — ʿadālah does not forget.
A pure sliding window forgets all jarḥ, including integrity (ʿadālah)
strikes — which silently reintroduces the fabricator-rehabilitation path
the framework exists to prevent. The tradition permits recovery only for
precision (ḍabṭ), never for integrity. So evidence carries an
EvidenceAxis:
- Integrity jarḥ (INTEGRITY, or UNSPECIFIED by conservative default) accumulates over the narrator’s entire history and never ages out. Each threshold’s worth lowers a permanent ceiling one tier; enough force REJECTED. This axis is intentionally a ratchet.
- Precision jarḥ (explicitly PRECISION-tagged) is windowed and recoverable, exactly as the base ratchet fix intends.
Integrity dominates: the permanent integrity ceiling caps the grade, and the windowed precision recovery only operates below that cap. Precision taʿdīl can never lift a grade held down by an integrity strike. The two axes are never averaged.
A production deployment would calibrate these thresholds and the window via the §8 gated-vs-ungated served-error experiment. The constants here are reference defaults, not validated values.
Registry (rijāl)
Section titled “Registry (rijāl)”class Narrator
Section titled “class Narrator”Narrator(narrator_id: 'str', domain_tag: 'str', narrator_type: 'NarratorType' = <NarratorType.MODEL: 'model'>, grade: 'NarratorGrade' = <NarratorGrade.UNGRADED: 'ungraded'>, adalah_grade: 'AdalahGrade' = <AdalahGrade.UNASSESSED: 'unassessed'>, dabt_grade: 'DabtGrade' = <DabtGrade.UNASSESSED: 'unassessed'>, known_error_rate: 'float | None' = None, model_version: 'str | None' = None, model_family: 'str | None' = None, upstream_source: 'str | None' = None, is_active: 'bool' = True, graded_at: 'datetime | None' = None, valid_until: 'datetime | None' = None, role: 'Role | None' = None, max_evidence_entries: 'int | None' = None)A narrator with its domain-conditioned grade and evidence log.
role is None for the default/integrity record (one per
(narrator, domain)) and a Role for role-scoped precision records
(one per (narrator, role, domain)). Integrity (ʿadālah) is only ever
stored on the default record and is shared across roles.
class Registry
Section titled “class Registry”Registry(transition_policy: 'TransitionPolicy | None' = None, volatility_policy: 'VolatilityPolicy | None' = None, max_evidence_entries: 'int | None' = None)The Rijāl Registry: stores and manages narrator grades per domain.
This is the pure-logic registry — usable without a database. For persistence, use RegistryDB backed by SQLAlchemy.
class RegistryDB
Section titled “class RegistryDB”RegistryDB(session: 'Session', transition_policy: 'TransitionPolicy | None' = None)Database-backed narrator registry.
Wraps the Registry in-memory store with SQLAlchemy persistence.
class Dispute
Section titled “class Dispute”Dispute(narrator_id: 'str', domain_tag: 'str', reason: 'str', disputed_grade: 'str', disputed_adalah: 'str', created_at: 'str') -> NoneA narrator’s contestation of their current grade (issue #38).
A dispute is logged (NEUTRAL — no grade change) and surfaced for operator adjudication. It is the audit trail that makes contestability possible: the original strike, the dispute, and the eventual adjudication all sit in the append-only evidence log.
accuracy_to_grade()
Section titled “accuracy_to_grade()”accuracy_to_grade(accuracy: 'float') -> 'NarratorGrade'Map a benchmark accuracy (0.0–1.0) to an ordinal narrator grade.
Cold-start bootstrapping helper (issue #33): a model’s published benchmark
accuracy seeds its precision (ḍabṭ) grade. Reference thresholds (the same
bands as BetaState.to_grade):
- ≥ 0.90 → RELIABLE (≤ ~10% error)
- ≥ 0.75 → ACCEPTABLE (≤ ~25% error)
- ≥ 0.60 → WEAK (≤ ~40% error)
- < 0.60 → REJECTED (> ~40% error)
seed_from_benchmark()
Section titled “seed_from_benchmark()”seed_from_benchmark(reg: 'Registry', narrator_id: 'str', domain: 'str', accuracy: 'float', *, role: 'Role | None' = None, benchmark: 'str' = 'published') -> 'NarratorGrade'Seed a narrator’s precision grade from a published benchmark accuracy.
Cold-start bootstrapper (issue #33): wraps Registry.seed so the seed is
recorded as BOOTSTRAP_SEED evidence carrying the benchmark name and the
accuracy (provenance: prior, not observation — see issue #6).
class SeedEntry
Section titled “class SeedEntry”SeedEntry(narrator_id: 'str', domain: 'str', grade: 'str', narrator_type: 'str', source: 'str', vertical: 'str' = 'self-maintaining-kb', metadata: 'dict[str, object]' = <factory>, model_family: 'str | None' = None, upstream_source: 'str | None' = None) -> NoneOne shipped default seed — an Estimated prior, never an observation.
The warm-registry schema (issue #203). Every entry carries an evidence
source (benchmark / publisher reputation / extraction suite) and is seeded
as BOOTSTRAP_SEED, so evidence_provenance() reports it prior_only and
gate_serve() caps it to SERVE_WITH_CAVEAT / REVIEW. “Supported”
(observation-backed) is never shipped — it is only earned by real pipeline
evidence.
default_registry()
Section titled “default_registry()”default_registry(*, vertical: 'str' = 'self-maintaining-kb') -> 'Registry'Return a Registry pre-seeded with the shipped, evidence-backed defaults.
Every seed is a BOOTSTRAP_SEED prior (“Estimated”), never an observation
(“Supported”). gate_serve() caps prior-only chains to
SERVE_WITH_CAVEAT or REVIEW, so a shipped seed can never plain-SERVE.
Grades stay operator-local: operators may override or ignore any shipped
entry.
default_seed_entries()
Section titled “default_seed_entries()”default_seed_entries(*, vertical: 'str' = 'self-maintaining-kb') -> 'list[SeedEntry]'The shipped warm-registry seed entries (Estimated priors, never observations).
Public so the serving path (isnad.api.dependencies) can warm its registry
from the same evidence-sourced defaults as default_registry(), instead of
a divergent hardcoded list.
Critics (matn)
Section titled “Critics (matn)”class ContentCritic
Section titled “class ContentCritic”ContentCritic(*args, **kwargs)Protocol for matn (content) criticism — fully decoupled from chain grading.
This is one instantiation of a parameter the framework leaves open (see paper §4.4). Swap freely.
Chain grading and content criticism are combined only at the end, by the decision matrix. They never read each other’s internals.
class EmbeddingCritic
Section titled “class EmbeddingCritic”EmbeddingCritic(contradiction_threshold: 'float' = 0.5)Default content critic — TF-IDF weighted, zero dependencies.
Works immediately with no installs, no downloads, no API keys. For better accuracy: pip install isnad[nli] and use HybridCritic.
This is one instantiation of a parameter the framework leaves open (see paper §4.4). Swap freely.
This critic is contradiction-ONLY: it returns CONTRADICTION when it finds a negation/antonym/numeric-divergence signal and UNVERIFIABLE otherwise. It never returns CONSISTENT — symmetric lexical overlap cannot affirm that a claim is correct (a <=3x dose error or an unlisted antonym pair would be blessed). Affirming consistency is the semantic critic’s job (NLI/LLM).
Args: contradiction_threshold: Cosine sim above which we check for contradiction signals.
class HybridCritic
Section titled “class HybridCritic”HybridCritic(embed_model: 'str' = 'all-MiniLM-L6-v2', nli_model: 'str' = 'cross-encoder/nli-deberta-v3-small', top_k: 'int' = 10, entailment_threshold: 'float' = 0.7, contradiction_threshold: 'float' = 0.5)Two-stage critic: fast embedding retrieval → NLI judgment.
Uses a fast embedding model to retrieve top-k relevant corpus claims, then applies the LocalNLICritic for precise entailment/contradiction.
Requires: pip install sentence-transformers
class LLMCritic
Section titled “class LLMCritic”LLMCritic(provider: 'str | None' = None, base_url: 'str | None' = None, api_key: 'str | None' = None, model: 'str | None' = None, top_k: 'int' = 5, cache_dir: 'str | None' = None, gate_affirmation: 'bool' = True)LLM-backed content critic with retrieval-augmented context.
Provider-agnostic: name a provider (openrouter, openai,
deepseek, anthropic, gemini, groq, together,
ollama) or pass a raw OpenAI-compatible base_url. When no
provider is named, the constructor auto-detects one from the
environment (see module docstring).
Args:
provider: Named provider (see list_providers()). Default: auto-detect.
base_url: OpenAI-compatible API base URL (overrides the provider default).
api_key: API key for the provider (overrides env vars).
model: Model name to use (required for providers with no default).
top_k: Number of similar corpus claims to retrieve as context.
cache_dir: Directory for on-disk cache (None = no caching).
class LocalNLICritic
Section titled “class LocalNLICritic”LocalNLICritic(model_name: 'str' = 'cross-encoder/nli-deberta-v3-small', entailment_threshold: 'float' = 0.7, contradiction_threshold: 'float' = 0.5, entailment_margin: 'float' = 0.2, contradiction_margin: 'float' = 0.2, retrieve_top_k: 'int' = 10, gate_affirmation: 'bool' = True)Local NLI-based content critic — semantic entailment/contradiction.
Uses a cross-encoder fine-tuned for NLI. For each retrieved corpus claim it computes entailment / contradiction / neutral probabilities, then decides:
- CONTRADICTION if a same-subject corpus claim clearly contradicts the claim.
- CONSISTENT if a same-subject corpus claim clearly entails it.
- UNVERIFIABLE otherwise.
Issue #110 — this critic previously had three defects, all fixed here:
- Label order — the cross-encoder outputs
[contradiction, entailment, neutral], but the code read[contradiction, neutral, entailment], so “neutral” was treated as “entailment”. - Raw logits vs probability thresholds — the thresholds were compared against raw logits (unbounded), not softmax probabilities.
- Max over the whole corpus — a claim was flagged against every corpus fact, so a “different fact” was spuriously read as a contradiction. The critic now retrieves the top-k similar claims first (TF-IDF), and only judges against those.
Args: model_name: HuggingFace cross-encoder model for NLI. entailment_threshold: entailment probability above which CONSISTENT. contradiction_threshold: contradiction probability above which CONTRADICTION. entailment_margin: entailment must exceed contradiction by this much. contradiction_margin: contradiction must exceed entailment by this much. retrieve_top_k: how many similar corpus claims to retrieve before NLI.
Example: critic = LocalNLICritic() result = critic.evaluate( “F = ma”, “f = m a”, [“force equals mass times acceleration”], “physics” )
class DeterministicRuleCritic
Section titled “class DeterministicRuleCritic”DeterministicRuleCritic()A deterministic, rule-based content critic.
This is one instantiation of a parameter the framework leaves open (see paper §4.4). Swap freely.
REFERENCE STUB: This critic detects contradictions via exact string matching against a hand-curated list of known contradictory phrase pairs. It CANNOT detect semantically equivalent contradictions phrased differently. The paper’s worked example passes because the required patterns (p=mv vs p=h/λ) are included. For production use, replace with an LLM-backed or embedding-based critic.
Patterns can be extended via add_pattern() or by subclassing
and overriding _CONTRADICTION_PATTERNS.
Types & enums
Section titled “Types & enums”Action
Section titled “Action”Enum — SERVE = 'serve', SERVE_WITH_CAVEAT = 'serve_with_caveat', REVIEW = 'review', QUARANTINE = 'quarantine', REJECT_AND_QUARANTINE_NARRATOR = 'reject_and_quarantine_narrator'
Actions from the decision matrix (paper §4.4, Table).
The 5×3 matrix: chain_grade ∈ {SAHIH, HASAN, DAIF, DAIF_JIDDAN, MAWDU} × content_verdict ∈ {CONSISTENT, CONTRADICTION, UNVERIFIABLE}
AdalahGrade
Section titled “AdalahGrade”Enum — HIGH = 'high', ACCEPTABLE = 'acceptable', SUSPECT = 'suspect', COMPROMISED = 'compromised', UNASSESSED = 'unassessed'
ʿAdālah: integrity / manipulation-resistance axis.
HIGH: trusted source, well-fenced, injection-resistant. ACCEPTABLE: no known integrity failures. SUSPECT: potential manipulation vector. COMPROMISED: known injection/poisoning source → active quarantine. UNASSESSED: never evaluated for integrity.
ChainGrade
Section titled “ChainGrade”Enum — SAHIH = 'sahih', HASAN = 'hasan', DAIF = 'daif', DAIF_JIDDAN = 'daif_jiddan', MAWDU = 'mawdu'
Ordinal chain grade for a claim, in descending trust order.
SAHIH > HASAN > DAIF > DAIF_JIDDAN > MAWDU. These are the hadith-authenticity tiers adapted to AI chains. A chain’s grade is capped by its weakest link, refined by transform type (see grading.py), and subject to completeness (ittiṣāl) enforcement.
ChainStatus
Section titled “ChainStatus”Enum — COMPLETE = 'complete', MUNQATI = 'munqati', ACTIVE = 'active', SUPERSEDED = 'superseded'
Create a collection of name/value pairs.
Example enumeration:
class Color(Enum): … RED = 1 … BLUE = 2 … GREEN = 3
Access them by:
-
attribute access:
Color.RED <Color.RED: 1>
-
value lookup:
Color(1) <Color.RED: 1>
-
name lookup:
Color[‘RED’] <Color.RED: 1>
Enumerations can be iterated over, and know how many members they have:
len(Color) 3
list(Color) [<Color.RED: 1>, <Color.BLUE: 2>, <Color.GREEN: 3>]
Methods can be added to enumerations, and members can have their own attributes – see the documentation for details.
ContentVerdict
Section titled “ContentVerdict”Enum — CONSISTENT = 'consistent', CONTRADICTION = 'contradiction', UNVERIFIABLE = 'unverifiable'
Result of matn criticism — independent of chain grade.
class CorroborationPolicy
Section titled “class CorroborationPolicy”CorroborationPolicy(*args, **kwargs)Protocol for how independent chains upgrade a claim’s grade.
This is one instantiation of a parameter the framework leaves open (see paper §4.3). Swap freely.
The default implementation applies:
- capped upgrade (never reaches SAHIH via corroboration alone)
- minimum-grade gate (at least one chain must clear threshold)
- correlation discount (correlated chains don’t count independently)
class CorrelationDetector
Section titled “class CorrelationDetector”CorrelationDetector(*args, **kwargs)Protocol for deciding whether two transmission chains are truly independent.
This is one instantiation of a parameter the framework leaves open (see paper §4.3, §7 — the madār problem). Swap freely.
The default implementation checks:
- shared narrator IDs (hard correlation)
- shared retrieved-document hashes (hard correlation — the madār case)
- shared model family (same base model / provider lineage)
- shared upstream source (both trace to the same origin)
Naive set-disjointness of narrator IDs is wrong; this detector captures correlated chains that share no explicit narrator but still fail together.
The document-hash and lineage signals arrive as optional keyword-only
arguments on the default implementation; custom detectors may ignore
them (the protocol signatures below stay minimal for backward
compatibility — callers pass document hashes through the concrete
SharedLineageDetector, not through this protocol).
DabtGrade
Section titled “DabtGrade”Enum — HIGH = 'high', ACCEPTABLE = 'acceptable', LOW = 'low', UNASSESSED = 'unassessed'
Ḍabṭ: precision / error-rate axis.
HIGH: calibrated error rate below threshold. ACCEPTABLE: adequate precision for domain. LOW: elevated error rate. UNASSESSED: never calibrated.
EvidenceAction
Section titled “EvidenceAction”Enum — JARH = 'jarh', TADIL = 'tadil', NEUTRAL = 'neutral'
Direction of evidence impact on narrator grade.
EvidenceProvenance
Section titled “EvidenceProvenance”Enum — PRIOR = 'prior', OBSERVED = 'observed', HUMAN = 'human', META = 'meta'
Where a piece of grading evidence came from — prior or observed instance.
Issue #6 (“ground rijāl grading in observed in-pipeline survival, not benchmark priors”) points at a real epistemic distinction the framework previously buried: a grade can be built on a population prior (a benchmark says this model class is ~85% accurate) or on observed instances (we watched THIS transmitter’s claims survive or fail inside the pipeline). Classical rijāl graded individuals on observed instances, never on priors.
- PRIOR: a population estimate — benchmark seeds, eval harnesses. Says nothing about whether THIS transmission is in the 85 or the 15.
- OBSERVED: an observed instance inside the operator’s own pipeline — a post-hoc audit of a served claim, or an independent-chain corroboration/contradiction. Slow to accumulate, but it is a record about this transmitter rather than a population estimate.
- HUMAN: a human reviewer verdict. Observed, but by a named human critic rather than an automated instance.
- META: not grade evidence at all — a lifecycle event (version bump) that resets the record.
This classification is a signal about the grade, not a new grading axis. It lets a caller answer “is this narrator’s grade an assumption or an observation?” without changing how grades are computed.
EvidenceType
Section titled “EvidenceType”Enum — EVAL_HARNESS = 'eval_harness', POST_HOC_AUDIT = 'post_hoc_audit', CORROBORATION_OUTCOME = 'corroboration_outcome', SURVIVAL = 'survival', HUMAN_REVIEW = 'human_review', DISPUTE = 'dispute', ADJUDICATION = 'adjudication', VERSION_BUMP = 'version_bump', BOOTSTRAP_SEED = 'bootstrap_seed', FRESHNESS_RENEWAL = 'freshness_renewal'
Named evidence types that drive narrator grade transitions.
The jarḥ–taʿdīl loop is a state machine, not a formula (paper §4.2). Transitions are driven by these evidence types, each logged immutably.
provenance_of()
Section titled “provenance_of()”provenance_of(evidence_type: 'EvidenceType') -> 'EvidenceProvenance'Classify an evidence type as prior, observed, human, or meta.
- BOOTSTRAP_SEED / EVAL_HARNESS → PRIOR (a population estimate).
- POST_HOC_AUDIT / CORROBORATION_OUTCOME / SURVIVAL → OBSERVED (an in-pipeline instance).
- HUMAN_REVIEW → HUMAN.
- VERSION_BUMP / FRESHNESS_RENEWAL → META (a reset or clock restart, not grade evidence).
class GradingStrategy
Section titled “class GradingStrategy”GradingStrategy(*args, **kwargs)Protocol for combining link grades into a chain grade.
This is one instantiation of a parameter the framework leaves open (see paper §4.2/§4.3). Swap freely.
The default implementation (RefinedWeakestLink) applies:
- strict minimum for destructive links
- bounded, corroboration-gated adjustment for generative links
- completeness cap (ittiṣāl)
NarratorGrade
Section titled “NarratorGrade”Enum — RELIABLE = 'reliable', ACCEPTABLE = 'acceptable', WEAK = 'weak', REJECTED = 'rejected', UNGRADED = 'ungraded'
Ordinal narrator grade in descending trust order.
These are the rijāl tiers from classical hadith science. The ordering is defined: RELIABLE > ACCEPTABLE > WEAK > REJECTED. Numeric error rates are optional metadata attached only where calibration data exists and MUST NOT be the primary grade or be surfaced to callers as if precise.
This is one instantiation of a parameter the framework leaves open (see paper §4.2). The ordinal categories are fixed; the transition arithmetic that moves a narrator between them is pluggable via TransitionPolicy.
NarratorType
Section titled “NarratorType”Enum — SOURCE = 'source', SCRAPER = 'scraper', MODEL = 'model', HUMAN = 'human', TOOL = 'tool'
Create a collection of name/value pairs.
Example enumeration:
class Color(Enum): … RED = 1 … BLUE = 2 … GREEN = 3
Access them by:
-
attribute access:
Color.RED <Color.RED: 1>
-
value lookup:
Color(1) <Color.RED: 1>
-
name lookup:
Color[‘RED’] <Color.RED: 1>
Enumerations can be iterated over, and know how many members they have:
len(Color) 3
list(Color) [<Color.RED: 1>, <Color.BLUE: 2>, <Color.GREEN: 3>]
Methods can be added to enumerations, and members can have their own attributes – see the documentation for details.
Enum — RETRIEVAL = 'retrieval', EXTRACTION = 'extraction', SYNTHESIS = 'synthesis', TOOL = 'tool', HUMAN = 'human', SOURCE = 'source'
What a transmitter did at a chain step (the task, not the agent).
Distinct from NarratorType (who/what the agent is). Precision (ḍabṭ)
is graded per (narrator, role, domain) because competence is
task-specific: a model can extract faithfully and synthesize carelessly
(issue #3). Integrity (ʿadālah) is NOT per-role — it is a judgment of the
person and is shared across roles.
TransformType
Section titled “TransformType”Enum — DESTRUCTIVE = 'destructive', GENERATIVE = 'generative', PASS_THROUGH = 'pass_through'
The transformation type of a chain link.
DESTRUCTIVE: extraction, chunking, lossy summarization — information is lost; downstream steps cannot recover it. The strict weakest-link minimum applies.
GENERATIVE: synthesis by a model with broad pre-training — may repair upstream noise OR introduce fresh corruption. Can raise the floor only up to its own grade and only when corroboration supports it; can always lower it.
PASS_THROUGH: identity-like transformation; does not affect grading.
class TransitionPolicy
Section titled “class TransitionPolicy”TransitionPolicy(*args, **kwargs)Protocol for how logged evidence moves a narrator between ordinal states.
This is one instantiation of a parameter the framework leaves open (see paper §4.2). Swap freely.
The framework default (BayesianTransitionPolicy) derives grades from a
Beta-distribution posterior mean, so it has no fixed “N adverse events”
cutoff. The simpler ThresholdTransitionPolicy is also available and uses:
- windowed threshold counts for adverse evidence → downgrade
- sustained corroborated accuracy → upgrade (requires N recent positive evals)
- version bump → reset to UNGRADED
(Both count over a sliding window of recent evidence and are edge-triggered on the arriving evidence — see issue #9.)