Methodology
Every story we publish carries three scores. This page explains exactly what they measure, how they are produced, and where the process can go wrong.
Last updated 29 July 2026
Barnet Press is an aggregating newsroom. We read what other outlets report about The London Borough of Barnet in north London — High Barnet, Chipping Barnet, New Barnet, East Barnet, Finchley, Hendon, Edgware, Whetstone, Totteridge, Golders Green, Mill Hill, Colindale and Burnt Oak., assess the quality of that reporting, and write our own account from the facts it establishes — always linking back to the originals.
The three scores are the reason to read us rather than any single outlet. They are published on every story, with a one-line explanation of why each landed where it did.
The three scores
Trust and Craft are not opinions of a story rounded to a number. Each starts at a floor and earns points only for things that can be pointed at in the text — a source named on the record, a figure whose origin is stated, a right of reply offered. Where the coverage falls short in specific ways, it loses points the same way. The lists below are the actual checks.
Trust — 0 to 100, higher is better
How well corroborated and evidenced the original reporting is. Starts at 30 and earns:
- a source named on the record, with their role
- a document, official statement or dataset quoted directly
- two or more genuinely independent outlets — syndicated copies of one article do not count
- the central claim attributed to a named body rather than to “sources”
- figures given with their origin stated, and specific dates and places
It loses points where the central claim rests on anonymous sourcing alone, or where key facts are hedged throughout. A short council report with one named officer typically lands in the forties or fifties; the seventies require corroboration that most daily local news simply does not have. A low Trust Score is not an accusation of dishonesty — it means there is less here that you or we can check.
Craft — 0 to 100, higher is better
The quality of the journalism itself, separately from how well sourced it is. Starts at 45 and earns:
- telling the reader why this matters, or what led to it
- quoting anyone affected, opposed or responsible, or offering a right of reply
- keeping fact clearly separate from comment
- specific figures, dates, streets and names rather than vague quantities
- saying what happens next, or naming what is still unknown
It loses points for a press release reproduced with no reporting added, for speculation presented as established fact, and for a headline that promises more than the body delivers.
Why the numbers changed in July 2026
An earlier revision described each band in words and asked the model to pick one. We checked what that produced across 393 stories and it barely discriminated at all: the Trust Score used six distinct values in total, and 262 of those stories sat in the same 70–79 band. Two very different pieces of reporting both scoring 75 tells you nothing, which hollows out the only thing this site offers.
So the scores are now built by adding up the checks above, and every story published before the change has been re-scored on the new basis. Numbers across the site are lower than they were and mean something different: they are a count of evidence present, not an impression. Eight older articles whose source text was never fully captured could not be re-scored, and rather than show you a number from a scale we have retired, we removed their scores entirely.
Hype — 0 to 100, lower is better
How far the presentation runs ahead of the substance. This is the one score where a high number is bad.
- 0–20 — measured, proportionate language.
- 41–60 — notable sensationalism, loaded adjectives, a thin story dressed up.
- 81–100 — fabricated urgency, or a headline contradicted by the article.
How a story is produced
Three separate passes run in a fixed order, and the order matters.
Syndicated copy counts once
Local titles frequently republish each other's articles word for word. Three mastheads carrying the same piece is one piece of reporting, not three confirmations, and treating it otherwise would push Trust Scores up for no good reason — an audit found half our multi-source stories contained syndicated copy. Identical source text is now detected before scoring and the assessment is told how many genuinely independent reports it is looking at.
1. Scoring runs first
The scoring pass sees only the original source articles — never our own write-up, because at that point it does not exist. This is deliberate. If we scored our own rewrite, the score would measure our writing rather than the reporting it came from, and would drift upwards over time.
2. Rewriting
A second pass writes an original article from the factual claims the scoring pass isolated. It is instructed to write from the facts rather than the sentences, to attribute every claim to who said it and which outlet reported it, and to invent nothing.
Every draft is then checked mechanically against its sources. We compare ten-word runs, ignoring anything inside quotation marks, since reproducing what someone actually said with attribution is journalism rather than copying. If more than 12 per cent of our unquoted prose matches a source, the draft is discarded and no article is published. For calibration: a hand-checked, genuinely original rewrite measures around 4 per cent, almost entirely on unavoidable phrasing such as statutory wording and restated statistics.
3. Packaging
A third pass writes the headline and social copy. It is constrained by the Hype Score we just published: our own headline may not read as more hyped than the coverage we are assessing. Publishing a hype measure and then writing clickbait would make the measure worthless.
We check that we keep to it. An audit of every published headline found none carrying a figure the article did not support, and one that escalated its source's register — the BBC reported that a translator had written a scathing review and we published “slams”. That headline was changed and logged in the corrections list, and the packaging instructions now name the tabloid verbs that are not to be used and rank alternative headlines by how well they inform a reader who will never click, rather than by how tempting they are.
What runs the passes
Current prompt revision: 2.1.0. Each story records the revision that judged it, so a score can always be traced to the exact instructions that produced it.
- Scoring — google/gemini-2.5-flash-lite, temperature 0
- Rewriting — google/gemini-2.5-flash-lite
- Packaging — google/gemini-2.5-flash-lite
Scoring runs at temperature zero, which is the setting that asks a model to be as consistent as it can be. It does not make it deterministic, and we would rather give you the measurement than the assurance. We took 15 stories and scored each one twice from identical input: the Trust Score came back unchanged 11 times out of 15 and the Craft Score 10 times out of 15, with an average difference of two to three points and a worst case of 16. Identical inputs can be routed to different machines, and the answer moves a little when they are.
So a score is repeatable to within about five points, not to the point. That is the honest resolution of the number, and it is why the bands below matter more than the digits.
Models are chosen for cost as well as quality: this is a local newspaper, not a research lab, and running it sustainably matters.
Where this can go wrong
We would rather state the limitations than have readers discover them.
- The scores are judgements, not measurements. Two careful editors would disagree at the margins, and so would two runs of a model. Treat a score of 72 and a score of 78 as the same verdict.
- We can only assess what we can read. If an outlet has strong off-the-record sourcing it cannot describe, the Trust Score will understate the story.
- Corroboration is not truth. Several outlets running the same press release is not the same as several outlets checking it independently, and the scoring pass cannot always tell the difference.
- Language models make mistakes. They can misread a date or conflate two people. This is why every story links its sources, and why corrections are logged publicly.
Getting it wrong
When we are wrong, we correct it and record it in the public corrections log, including any case where a published score was changed and why. If you think a story or a score is wrong, tell us at editor@barnetpress.co.uk.
Common questions
How is the Trust Score calculated?
The Trust Score assesses how well corroborated and evidenced the original reporting is. A story scores 90 or above when several independent outlets carry it with named primary sources, documents or official statements quoted directly. It scores 50 to 69 when a single outlet carries it, or when official claims are repeated without independent checks. Below 30 means no verifiable sourcing at all. It is not a judgement of whether the underlying events are good or bad.
What does the Hype Score mean, and why is lower better?
The Hype Score measures how far the presentation of a story runs ahead of its substance. A low score means measured, proportionate language. A score above 60 means heavy clickbait, manufactured outrage, or a headline the article body does not support. Unlike Trust and Craft, a low Hype Score is the good outcome.
Do the scores judge the events or the reporting?
The reporting. A well-sourced, carefully written article about a terrible event will score highly on Trust and Craft. The scores describe how the story was told, not what happened.
Is this AI-generated content?
It is AI-assisted journalism under human editorial supervision, and we say so on every article. Language models score the source coverage, write the article from the established facts, and draft the headline. Editors set the rules those models follow, review the output, and are accountable for what is published.
Do you copy other outlets’ articles?
No. Articles are written from the facts, not from the sentences. Every draft is checked automatically against its sources: if more than 12 per cent of our unquoted prose matches a source across ten-word runs, the draft is rejected and never published. Direct quotes are reproduced only with clear attribution, and every original report is linked.