This page explains the process that produced the page you came from — how the record was written, scored, and labeled. It does not take a side on the resolution itself.
An independent panel of judge models — always kept separate from the two models debating — scores every point on four dimensions: evidence (is it real, accurately quoted, relevant, strong?), logic (does the evidence actually support the claim?), clash (did the point survive its attacks, or did the attack land?), and weight (how much does it matter to the resolution?). Each side's total is the sum of its scored points.
The page reports the verdict two ways, because they can disagree:
When the total and the average name different leaders, the page says so explicitly (a "volume-sensitive" flag) instead of quietly picking one winner — a side that simply wrote more is not the same as a side that argued better.
Every point on the page carries one of four labels, each backed by a concrete rule — never an unexplained gut call:
A round is not allowed to close as complete while either side still owes work. Concretely: every one of the opponent's constructive points must be met by at least one active rebuttal, and every rebuttal aimed at a side's still-standing point must be met by a defense. A side cannot declare itself "resting" while any of that is still open — the engine refuses the attempt and tells it exactly what's still owed.
A round can still close early if a time budget runs out before coverage is complete. When that happens the page says so plainly ("ended early — unmet: coverage" or similar) rather than presenting a partial record as if it were finished business.
Every citation tagged as Scripture evidence is checked against the source text directly — not trusted on the model's word. Each one lands in exactly one bucket: verified (matches exactly), variant (matches a known textual/translation variant), mismatch (does not match what's cited), or not found (the reference doesn't resolve). A misquoted or fabricated citation is treated as the gravest offense a point can commit.
Multiple independent judge models score every point separately; the page shows the full per-judge ledger behind every total, not just the final number. Two safeguards run underneath the scores you see: