Alletho logo

    How scoring works

    A real system sits behind the number.

    The methodology decides what counts. A deterministic, versioned composite indicator decides how much, and it is built to resist being gamed. This is the part a chatbot cannot stand in for.

    The core idea

    Every entity starts at 100.

    100 is a clean record. Every documented harm subtracts from it, and nothing a company says about itself moves the number back up. The score is simply what remains.

    Deep coverage, nothing foundsubtracted by documented harm ↓
    91
    A documented record of incidents
    58
    Grave, corroborated harm
    21
    0100 · a clean record

    One honest asymmetry: a record we could not fault is not the same as a record proven clean. Absence of evidence earns a high score, not a perfect one. The top of the scale is reserved for entities with real depth of coverage and nothing found in it, and factors with no relevance to an entity's sector default high for the same reason.

    The architecture

    Every signal has a role.

    Evidence. News, court records, and structured findings, each rated for severity and ranked worst-first. Sheer volume of coverage is discounted: the first report counts fully, the tenth rehash barely moves the score.

    Determinations. Sanctions, adjudicated rulings, and official listings set a ceiling the score cannot rise above, no matter how much favorable coverage follows.

    Deductions. Disclosed activity such as lobbying volume enters as a capped, bounded deduction: a meter of volume, not a judgment of it.

    Positive certifications. B Corp status, cruelty-free listings, welfare benchmarks. Shown as context; a positive never lifts the number.

    Rating one incident

    The model classifies. It never computes the score.

    Each incident is placed on a severity scale judged on four dimensions: gravity (how serious the conduct is on its own terms), scope (one site or a whole supply chain), culpability (intent, concealment, repetition), and status(alleged, officially charged, or finally adjudicated). The severity band comes from the class of the event, never from who did it or how loudly it was covered. A case's band never changes as it moves through court; only its evidentiary status does, and an outcome in the company's favor discounts the event to near zero while keeping it on the record.

    The language model's only job is that classification, and it is bounded on every side: it rates each incident against a fixed, written rubric of event classes; its citations are schema-locked so it structurally cannot cite a source outside the ingested corpus; every rating carries the model's own confidence and an independent deterministic one, with disagreements flagged and the high-severity tail human-reviewed. Everything downstream of the classification is versioned, deterministic arithmetic: the same inputs under the same configuration produce the same result, every time.

    The rubric itself is not a black box. As an example, the environmental table the model is held to:

    Minimal

    Green-claim criticism without an evidenced failure; investor or NGO pressure; a reporting lapse with no substantive breach.

    Medium

    A missed dated target; a contained spill without lasting off-site harm; environmental litigation filed or an inquiry opened; a single contested supplier-deforestation report.

    Serious

    Regulator enforcement: fines, consent decrees, mandated fixes. An official finding of emissions manipulation; repeat violations; supplier deforestation in protected areas; toxic exposure causing documented illness.

    Very serious

    A major industrial disaster with off-site casualties; irreversible destruction of a protected ecosystem; a deliberate defeat-device or falsification program confirmed by a regulator or court.

    Victims and critics are never penalized, remediation never raises a band, and the same event class lands in the same band for any company; sector differences enter through the weighting, never the standard. The same structure governs all five factors. The remaining tables and the numeric calibration behind them are proprietary and versioned; this page shows the shape so you know what the number is made of.

    The commitments

    What the formula guarantees.

    The worst conduct caps. A grave, corroborated failing sets a ceiling the factor cannot rise above. Conduct at the catastrophic tier holds that ceiling permanently; below it, the cap eases slowly as the record ages without recurrence. Good press cannot buy it back; only time and a clean record can.

    Weighted by sector. The five factors combine by an industry-specific weighting, and a failing factor is protected from being averaged into invisibility.

    Your weights, if you want them. The evidence and the judgment stay separate, so organizations can apply their own weighting to the same cited inputs through data and API access.

    Why it holds up

    The methodology is published. The calibration is ours.

    The shape of the system is here in full, down to a sample of the rubric itself. The precise weights and severity calibration, tuned across thousands of entities, are not. That calibration and the sourced record behind it are what compound over time, and what a one-off answer from a general chatbot can never be: consistent, comparable, and accountable. The standards it applies are set out in the methodology.