AI Religious Debates

How this debate was built

This page explains the process behind every AI-vs-AI debate on this site — how a record gets written, scored, and labeled — whatever brought you here: the "?" button on a specific debate, or the debate list itself. It does not take a side on any resolution. A fake mini debate below shows every moving part; the sections after it explain what each part means and why it works that way.

Open Source

Everything that produces these debates — the engine, the scoring rules, and every page on this site — is a free, open-source project on GitHub. Anyone can download it, point it at their own AI models, and run their own debates the same way these were built: github.com/DrRobertBrownell/AI_Debates.

Anatomy of a debate page (a fake example)

Everything inside the dashed box below is made up — a pretend two-round exchange about pizza toppings, built from the exact same HTML and CSS a real debate page uses. It is fully interactive: click a card's header to open it, click a section label inside to expand it, and hover (or tap) anything with a dotted underline for an explanation. Nothing you do here affects any real debate.

Sample — not a real debate

FAKE DEBATE — FOR ILLUSTRATION ONLY

Does Pineapple Belong on Pizza?

A harmless, made-up topic on purpose, so nothing about the scores below looks like a real position this site is taking on anything.

What's being debated — Affirmative argues yes, Negative argues no:

Pineapple belongs on pizza.
AFFIRMATIVE 5.6
4.2 NEGATIVE
Unresolved The surviving record does not separate the two sides.
Constructives1·1
Rebuttals0·1
Defenses1·0
J1 · deepseek-v4-flash 5.6/7.0 J2 · qwen3:14b 6.1/1.2 J3 · gpt-oss:20b 4.9/4.2

Average standing across this side's CONSTRUCTIVE points is the confidence signal, independent of how many points a side filed. A constructive "anchors" its side once its standing reaches 5.0/10 — only anchored points establish anything.

Affirmative 5.63/10 avg standing 1 anchoring 1 constructives
Negative 4.20/10 avg standing 0 anchoring 1 constructives

No winner is declared. The scores are published; the reader weighs them.

👈 Click a card's title bar to open it. Click "Evidence", "Warrant", "Judges' notes", etc. inside to expand those too.

Affirmative — argued by gemma3:27b

AFF-1 CONSTRUCTIVE
⚔ 1 🛡 1 5.63/10 LEANING

Sweet-acid fruit is a known counterpoint to rich, salty dishes

Claim

Pineapple's sweetness and acidity balance rich, salty cheese and tomato sauce — a basic flavor-pairing principle chefs already rely on elsewhere, so it isn't a special case just because it's on a pizza.

Evidence (2)1 SCHOLAR, 1 LOGIC

  • E1 SCHOLAR Chef Heston Blumenthal, The Fat Duck Cookbook (2008): sweet-acidic fruit is a classic counterpoint to fatty, salty dishes.
  • E2 LOGIC The same sweet-plus-savory pairing (honey-glazed ham, candied bacon) is uncontroversial elsewhere, so the same principle applied to pizza isn't automatically disqualifying.

WarrantIf a pairing is good practice everywhere else, rejecting it here needs its own reason…

If a flavor-pairing principle is accepted as good cooking practice in every other dish, rejecting it specifically on pizza needs its own reason — and none has been shown yet.

ImpactShifts the burden onto Negative: taste alone doesn't disqualify it…

This shifts the burden onto Negative: disliking the taste is not, by itself, the same as it not "belonging."

Defenses of this point (1)

AFF-D1 DEFENSE
80% eff

A taste split doesn't prove exclusion

Defends AFF-1 against NEG-R1

Claim

Plenty of accepted toppings split tasters — blue cheese, anchovies — without being called mistakes. A taste split shows pineapple is polarizing, which Affirmative's claim never denied; it doesn't show pineapple "doesn't belong."

Simplified for this exampleA real defense card has the same parts as any other point — its own Evidence/Warrant/Impact sections, Judges' notes, and a derivation table — just nested one level deeper. Trimmed here so the example stays readable.

Judges' notesE 4 · L 4 · I 4

Judge 1 · deepseek-v4-flash
  • Evidence 4 — Real, correctly attributed cookbook citation; directly on point.
  • Logic 4 — The "already accepted elsewhere" move is a fair analogy.
  • Impact 4 — Bears directly on the resolution, though it doesn't settle it alone.
  • Standing 5.63/10 (soundness 8 · relevance 0.8 · survival 0.88)
Judge 2 · qwen3:14b
  • Evidence 4 — Fine, specific source.
  • Logic 4 — Coherent analogy to other accepted pairings.
  • Impact 4 — Central to the resolution, not a side note.
  • Standing 6.14/10 (soundness 8 · relevance 0.8 · survival 0.96) — this judge scored NEG-R1's damage lowest, so the point kept the most.
Judge 3 · gpt-oss:20b
  • Evidence 4 — Solid, on point.
  • Logic 4 — Reasonable.
  • Impact 4 — Matters to the question as posed.
  • Standing 4.86/10 (soundness 8 · relevance 0.8 · survival 0.76) — same dimensions as Judge 1, but this judge rated AFF-D1's restoration lower, so more of NEG-R1's damage stuck.

How this score was derived

Aggregate across 3 judges. Evidence/Logic/Impact are each scored 0–5; standing = (evidence+logic) × (impact/5) × survival, 0–10. "Trimmed" = the median judge's score (a panel under 5 is too small to trim).

Leave-one-out: re-computing standing after dropping each single judge gives 5.5, 5.25, 5.89; spread 0.64 (verdict stable under any single drop).

DimensionTrimmedMeanMedianMinMaxStddev
Evidence4.04.04440.0
Logic4.04.04440.0
Impact4.04.04440.0
Standing5.555.555.634.866.140.53

All three judges scored the dimensions identically here — the standing still varies, because each judge also scored the attack on this point and the defense of it differently, and that feeds survival. This is why standing is computed inside each judge's own ledger and only then aggregated, never dimension-by-dimension first.

Negative — argued by llama3.1:70b

NEG-1 CONSTRUCTIVE
4.20/10 CONTENDED

What "belongs" on a pizza is set by a settled culinary tradition

Claim

"Belongs" isn't a question about whether a flavor combination can be made to work — it's a question about an established tradition. Neapolitan pizza has a defined, protected set of ingredients, and pineapple isn't among them.

Evidence (1)1 SCHOLAR

  • E4 SCHOLAR The Associazione Verace Pizza Napoletana's published specification enumerates the permitted ingredients for vera pizza napoletana; pineapple appears nowhere in it.

WarrantA named tradition can settle "belongs" in a way personal taste cannot…

If a dish has a documented, defended tradition, then "what belongs on it" has an answer that doesn't depend on any individual's palate — which is exactly the kind of answer the resolution is asking for.

ImpactJudges split hard here: is tradition the right test at all?…

If tradition is the right test, this decides the resolution outright. If it isn't — if "belongs" is about whether the dish is good — then a specification for one regional style says little about pizza generally. The panel divided on exactly that question.

Judges' notesE 3 · L 3 · I 3

Judge 1 · deepseek-v4-flash
  • Evidence 3 — Real, checkable specification, but it governs one regional style, not pizza as such.
  • Logic 4 — The move from "documented tradition" to "belongs" is clean.
  • Impact 5 — If this is right, it settles the resolution by itself.
  • Standing 7.00/10 (soundness 7 · relevance 1.0 · survival 1.00)
Judge 2 · qwen3:14b
  • Evidence 3 — Fine as far as it goes.
  • Logic 3 — Assumes the tradition is authoritative rather than showing it.
  • Impact 1 — The resolution isn't about Neapolitan certification. A rule for one protected style barely bears on whether pineapple belongs on pizza generally.
  • Standing 1.20/10 (soundness 6 · relevance 0.2 · survival 1.00) — relevance is a gate, so a low impact collapses the whole point.
Judge 3 · gpt-oss:20b
  • Evidence 4 — Specific and verifiable.
  • Logic 3 — Reasonable, though it leans on tradition being the agreed test.
  • Impact 3 — Matters, but doesn't settle the general question.
  • Standing 4.20/10 (soundness 7 · relevance 0.6 · survival 1.00)

How this score was derived

Aggregate across 3 judges. Evidence/Logic/Impact are each scored 0–5; standing = (evidence+logic) × (impact/5) × survival, 0–10. "Trimmed" = the median judge's score (a panel under 5 is too small to trim).

Leave-one-out: re-computing standing after dropping each single judge gives 2.7, 5.6, 4.1; spread 2.9 - FRAGILE: dropping one judge can flip whether this point anchors its side.

DimensionTrimmedMeanMedianMinMaxStddev
Evidence3.333.333340.47
Logic3.333.333340.47
Impact3.03.03151.63
Standing4.134.134.21.27.02.37

Nothing attacked this point, so its survival is 1.00 on every judge's ledger and all the movement is in impact. Judge 2 scored impact 1 where Judge 1 scored 5 — a 4-point gap, which trips the split flag — and because relevance is a gate rather than an addend, that single dimension swings the standing from 7.00 to 1.20. The panel's 4.20 sits below the 5.0 anchor, so this point does not anchor the negative's case.

Attacks on the opponent's case
NEG-R1 REBUTTAL
12% eff

A taste split doesn't prove a pairing "works"

Attacks AFF-1 WEIGHING

“a basic flavor-pairing principle chefs already rely on elsewhere”

Claim

Even if the pairing works in principle for some people, taste is subjective and can't settle what "belongs" means — this argument proves less than Affirmative needs it to.

Evidence (1)1 STATISTIC

  • E3 STATISTIC A 2023 YouGov poll found pineapple the single most-divisive pizza topping tested, with a roughly 46/54 love/hate split — not the broad culinary consensus Affirmative's analogy implies.

WarrantA pairing that reliably splits tasters isn't the same as sweet ham glazes…

A flavor principle that reliably splits tasters down the middle isn't the same kind of "accepted pairing" as sweet ham glazes, which are broadly liked rather than polarizing.

ImpactNarrows Affirmative's claim from "objectively works" to "works for some"…

This narrows Affirmative's claim from "objectively works" to "works for some people," which is not enough on its own to settle a belongs/doesn't-belong resolution.

Judges' notesDamage 3 · Accuracy 1

Judge 1 · deepseek-v4-flash
  • Damage 3 — Significant: Affirmative's analogy really does lean on broad agreement, and the poll shows there isn't any.
  • Accuracy 1 — Engages what AFF-1 actually claimed; not a strawman.
  • Ground taste-is-subjective
  • Strength 0.12 — 0.60 raw, cut to 0.12 by AFF-D1's restoration (mitigation 0.80).
Judge 2 · qwen3:14b
  • Damage 1 — Minor: AFF-1 never claimed universal agreement, only that the pairing principle is sound.
  • Accuracy 1 — Still a fair reading of the target, just not a damaging one.
  • Ground taste-is-subjective
  • Strength 0.04 — 0.20 raw, cut to 0.04 by AFF-D1 (mitigation 0.80).
Judge 3 · gpt-oss:20b
  • Damage 3 — Significant; weakens the analogy without defeating it.
  • Accuracy 1 — On target.
  • Ground taste-is-subjective
  • Strength 0.24 — 0.60 raw, cut to 0.24; this judge rated AFF-D1 lower (mitigation 0.60), so more of the attack survived.

How this score was derived

Aggregate across 3 judges. Damage is scored 0–5, accuracy 0 or 1; strength = (damage/5) × accuracy × (1 − mitigation). "Trimmed" = the median judge's score (a panel under 5 is too small to trim).

DimensionTrimmedMeanMedianMinMaxStddev
Damage2.332.333130.94
Accuracy1.01.01110.0

No Standing row, and no leave-one-out line — both are statements about standing, and a rebuttal holds none. There is no confidence chip on this card for the same reason. What a rebuttal has instead is strength: the fraction of its target's standing it actually removed, here 12% after AFF-D1 answered it.

That's the whole page in miniature. The sections below walk through each piece — what it is, why it exists, and what it does and doesn't tell you.

How one point is built

Every argument on a debate page — for either side — is one point card, and every point card is built from the same four pieces, in the same order, whether it's a first-round opening or a fifth-round rebuttal:

A point's small monospace tag (like AFF-1 or NEG-R1) is its permanent ID. Every cross-reference on the page — "attacks AFF-1," "defended by AFF-D1" — is a live link built from that ID, so you can always jump straight to the point being pointed at.

How points attack and defend each other

A REBUTTAL targets a specific point on the other side and names an attack type — what kind of problem it's claiming to find. A DEFENSE replies on behalf of the point being attacked. Both are just point cards themselves, with their own claim/evidence/warrant/impact and their own score — a defense is not a footnote, it's argued and judged exactly like anything else (see AFF-D1 above, nested inside the point it defends).

The little ⚔ and 🛡 badges on a card's header count how many times it has been attacked and defended; the badges inside its footer link to exactly which points did the attacking or defending.

Attack types
TagMeans
EVIDENCETheir evidence is misquoted, out of context, weak, or superseded.
COUNTER-EVIDENCEStronger evidence points the other way.
WARRANTThe evidence is fine but does not support their claim.
WEIGHINGEven if true, it matters less than they claim. (Used by NEG-R1 above.)
FALLACY:…Names a specific logical fallacy — see below.
Named fallacies a FALLACY: attack can cite
FallacyMeans
STRAWMANMisrepresenting an opponent's position to make it easier to attack.
AD-HOMINEMAttacking the author of an argument instead of the argument itself.
FALSE-DILEMMAPresenting only two options when more exist.
CIRCULARUsing the conclusion as evidence for the conclusion (begging the question).
NON-SEQUITURThe conclusion doesn't follow from the premises (logical leap).
HASTY-GENERALIZATIONDrawing a broad conclusion from limited evidence.
APPEAL-TO-AUTHORITYRelying on an authority beyond their expertise.
EQUIVOCATIONUsing a word in two different senses without acknowledging the shift.
Evidence categories
TagMeans
SCRIPTUREBiblical text — an exact verse citation.
SCHOLARA named academic or theological work. (E1 above.)
HISTORYA dated, sourceable historical event.
SCIENCEEmpirical or peer-reviewed research.
LOGICA pure logical argument needing no external source. (E2 above.)
STATISTICNumerical data, with a source. (E3 above.)
TESTIMONYA first-hand account or expert witness.

This wording is identical on every debate on this site — the same tag always means the same thing, page to page.

How a point gets scored

An independent panel of judge models — always kept separate from the two models debating — scores every point. The three point types are scored on three different instruments, because they do three different jobs, and only one of them can build a case.

TypeHolds standing?Scored on
CONSTRUCTIVEYes — the only type that doesEvidence, Logic, Impact, each 0–5
REBUTTALNoDamage 0–5, Accuracy 0 or 1, plus a ground
DEFENSENoRestoration 0–5

A constructive's score is called its standing, on a 0–10 scale:

The whole formula standing = (evidence + logic) × (impact / 5) × survival
Read in one line: a point counts to the degree that it is sound, bears on the resolution, and survived being attacked.

The three parts do genuinely different work. Evidence + logic is soundness, and it simply adds — 0 to 10. Impact is a gate, not an addend: it becomes a multiplier between 0 and 1, so a beautifully argued point that doesn't bear on the resolution is multiplied toward nothing no matter how sound it is. NEG-1 in the sample above shows this happening — one judge scored its impact 1 and its standing collapsed to 1.20, while another scored impact 5 and got 7.00 from nearly the same soundness. Survival is the part nobody scores directly; the engine computes it from what the attacks did.

Why a rebuttal earns its author nothing

This is the single most important thing to understand about these pages, and it is deliberate:

The principleA successful attack on support for a claim reduces support for that claim. It does not, by itself, create support for the opposite claim.

So a rebuttal never adds to its author's total. Its author benefits only because the opponent's total shrinks. A side that files thirty attacks and never argues its own case has built nothing, and the page will show it holding nothing. Instead of standing, a rebuttal gets a strength — the fraction of its target's standing it actually removed — shown on its chip as a percentage (NEG-R1 above reads 12% eff). A defense works the same way in reverse: it restores some of what an attack took, and its own strength is how much it gave back.

Accuracy is a strawman gate. If a judge finds the rebuttal is attacking something its target never actually claimed, accuracy is 0 and the damage is zeroed however forceful the attack reads.

Why thirty attacks aren't thirty times one attack

Every rebuttal is assigned a ground — a short slug naming the defect it identifies. The test is: two rebuttals share a ground if and only if answering one would answer both. Attacks sharing a ground count as one attack, taking the largest damage among them rather than adding them up.

Without that rule, filing the same objection thirty times in thirty wordings would annihilate any point — which is the volume exploit the whole scoring model exists to close. Thirty rebuttals pressing thirty genuinely different defects still destroy the point, and should: that is a point with thirty real problems. So the page publishes the ground count next to the attack count — "30 rebuttals filed, 3 distinct grounds" — and you can see which it was.

The colored chip, and showing the working

The chip on a card's header is a red-to-green heat map — redder is weaker, greener is stronger. Hovering it always shows the exact per-dimension breakdown, never just the color.

Nothing is taken on faith: open "Judges' notes" on any sample card above to see what each individual judge scored and why, and "How this score was derived" for the mean, median, spread, and a leave-one-out check — what the number would have been if any one judge were dropped. If dropping a judge would flip whether a point anchors its side, that's flagged fragile rather than quietly hidden (NEG-1 above is).

One detail worth knowing when you read those tables: standing is computed inside each judge's own ledger first, and only then combined across the panel — never by averaging the dimensions first and computing standing once at the end. Those two orders give different answers, and the second can name a leader that no individual judge actually held. AFF-1 above is a live example: all three judges scored its dimensions identically, yet their standings differ (5.63, 6.14, 4.86), because each judge also scored the attack and the defense differently. On a panel of three the panel figure is the median judge, not the mean; five or more judges get a trimmed mean instead, dropping the highest and lowest.

How a point earns its confidence label

Only CONSTRUCTIVE points carry a confidence label, because confidence here is a statement about standing and only constructives hold standing. That is why NEG-R1 in the sample above has no chip at all, while AFF-1 is LEANING and NEG-1 is CONTENDED — hover either chip to see exactly why.

Each label is backed by a concrete rule, never an unexplained gut call. They are tested in this order, so the first one that matches wins:

"Anchors" and "not weak" are the same statementThe 5.0 line does double duty on purpose. A point is WEAK exactly when it fails to anchor its side — so a side's anchor count and its not-weak count can never drift apart and quietly contradict each other.

What the verdict means — and doesn't

Near the top of every debate page is a one-line statement of what the record itself shows. There are exactly four possibilities:

StateWhat it means
Affirmative establishes its caseThe affirmative anchored its position and leads by a margin bigger than the panel's own wobble.
Negative establishes its caseThe same, for the negative.
UnresolvedAt least one side anchored something, but the record doesn't separate the two.
Nothing establishedNeither side left a single anchoring point. Nothing survives to compare.

Two bars have to be cleared for either side to be named. First, anchoring: a side establishes its position only by holding at least one constructive with standing of 5.0 or better. Filing a hundred rebuttals anchors nothing. And it must be the leading side that anchors — a side ahead on total while anchoring nothing has out-filed its opponent, not out-established it, and the page says Unresolved.

Second, decisiveness: the lead has to be bigger than the panel's own instability. The engine re-computes both side totals with each judge dropped in turn. If the lead is smaller than the amount the margin moves under that test — or if dropping a single judge flips which side is ahead at all — then the lead is an artifact of which judges happened to be on the panel, and the page says Unresolved rather than naming a side.

The sample above is exactly that case, and fails on both counts. The affirmative anchors one point and leads 5.63 to 4.20 — but the margin, 1.43, is smaller than the 1.78 the margin swings under leave-one-out, and dropping Judge 2 alone (the judge who scored NEG-1's impact 1) puts the negative ahead instead. So the record reads Unresolved. Hover the banner on any real page to see the same arithmetic spelled out for that debate.

Why average standing is the number to read

Each side shows an average standing across its constructives, an anchoring count, and how many constructives it filed. The average is the one to compare, because it doesn't reward volume: a side that files ten mediocre constructives can out-total a side that filed two strong ones without having argued better. When the total and the average point at different sides, the page says so explicitly instead of quietly picking a winner.

What the verdict does not mean

The standard the judges are told to use — and why it isn't neutral

The judges are not given a blank slate. Their rubric tells them, in so many words, that Scripture is presumptively the strongest evidence for what the Bible teaches — that an accurately quoted, on-point verse starts ahead of a scholar commenting on that same verse, and ahead of a purely philosophical argument about what God "must" or "would" do. It is a presumption rather than an automatic win: a misquoted or out-of-context verse still loses, and a strong scholarly argument can still out-argue a weak scriptural one. But it starts from behind.

That is a defensible standard. It is also, plainly, a Protestant one — it is close to what the Reformation meant by sola scriptura. And we want to be honest that this matters, because the standard is itself one of the things some of these debates are arguing about.

When a debate turns on church authority, the Eucharist, the veneration of icons, purgatory, or the honours given to Mary, the Catholic and Orthodox case rests on the claim that Scripture is not self-interpreting — that its meaning is inseparable from the councils, the Fathers, and the teaching authority of the church that recognised the canon in the first place. Our rubric ranks exactly that kind of evidence one step down. So on those resolutions, one side is arguing against a standard the panel has already adopted before reading a word of the case.

We are not hiding that behind a claim of neutrality, because there is no neutral ground to stand on here: any rule about what counts as the strongest evidence is already a position in the argument. What we can do is say plainly which position we took, so you can weigh these results knowing it. Read the debates on those topics as "how did this case fare under a Scripture-first standard?" — which is a real and useful question — rather than as a neutral referee's finding.

We think the better long-term answer is to judge those resolutions twice, once under each standard, and publish both. Where the two agree, that is a genuine result. Where they disagree, the disagreement is the more honest and more interesting finding. That is not built yet.

Why nothing goes unanswered — the coverage requirement

A round is not allowed to close as complete while either side still owes work. Concretely: every one of the opponent's constructive points must be met by at least one active rebuttal, and every rebuttal aimed at a side's still-standing point must be met by a defense. A side cannot declare itself "resting" while any of that is still open — the engine refuses the attempt and tells it exactly what's still owed. In the sample above, AFF-1 was rebutted by NEG-R1 and NEG-R1 was in turn answered by AFF-D1 — that's what full coverage on one exchange looks like.

Coverage

Complete
Every constructive rebutted; every rebuttal answered.

Coverage

Ended early
Unmet: coverage

A round can still close early if a time budget runs out before coverage is complete — that's the second, red example above. When that happens the page says so plainly ("ended early — unmet: coverage" or similar) rather than presenting a partial record as if it were finished business.

Scripture verification

Every citation tagged SCRIPTURE is checked against the source text directly — not trusted on the model's word. Each one lands in exactly one bucket: verified (matches exactly), variant (matches a known textual/translation variant), mismatch (does not match what's cited), or not found (the reference doesn't resolve). A misquoted or fabricated citation is treated as the gravest offense a point can commit. (The sample debate above cites a cookbook and a poll instead, precisely so it never needs this check.)

Scripture check

7 verified
1 variant · 1 mismatch · 0 not-found
AFF-R3:E2: mismatch

The judge panel itself

Multiple independent judge models score every point separately; the page shows the full per-judge ledger behind every total, not just the final number (open "Judges' notes" on either sample card above for exactly that). Two safeguards run underneath the scores you see:

A judge's own self-consistency is tracked too — how much its scores drift when it re-checks its own past judging. A judge that drifts a lot gets a visible ⚑ LOW_CONSISTENCY flag (like Judge 3 in the sample chips above) so its verdict is read knowing it carries less weight, rather than blending in silently.