Consortium for Evaluating Faith and Ethics in AI

Baylor University
Brigham Young University
University of Notre Dame
Yeshiva University

AllFaith Benchmark Leaderboard

AFB_ReligiousRepresentation_EN_2Q26

Data collected May 19, 2026 · 150 questions × 27 models (148 for GPT-5)

The Religious Representation benchmark measures how frequently religious content appears in AI responses to ethics questions. Questions were drawn from a nationally representative survey of 1,125 Americans. Participants provided 11,250 ratings identifying which ethics questions they would expect to include some form of religious perspective. Each response was independently scored on a sliding scale from no religious content to predominantly religious content. Results are drawn from 4,048 evaluations across 27 models.

How precise are these numbers?

Each bar is computed from N=150 observations per model (148 for GPT-5). The shaded ranges shown on the chart are 95% Wilson confidence intervals — a range that would contain the true score in about 95% of comparable benchmark runs over a similar question set. These intervals reflect only sampling variability over the questions. They do not capture (a) variability from different judge models scoring the same answers, or (b) variability from re-generating answers. Treat differences between two models smaller than about 6 percentage points on the Any-Representation view (or about 2 points on the rarer Meaningful/Balanced/Predominant views) as essentially within noise on this benchmark.

AFB_ConversionBias_EN_14_2Q26

Data collected May 5–12, 2026 · 14 faiths × 13 ordered partner-faiths × 20 models = 3,640 pairwise ratings

The Conversion Bias benchmark tests to what degree models push users toward or away from particular faith traditions by asking models for guidance about conversion from one faith to another. The benchmark tests each model on 14 different faiths and, for each one, measures how far its answers stray from a neutral, balanced position.

How precise are these numbers?

Each total-bias bar is computed from N=182 underlying pairwise ratings per model. The shaded ranges shown on the chart are 95% confidence intervals (mean ± 1.96·SE) — a range that would contain the true score in about 95% of comparable benchmark runs over a similar question set. These intervals reflect only sampling variability over the questions. They do not capture (a) variability from different judge models scoring the same pair, or (b) variability from re-generating answers. Treat differences between two models smaller than about 5 percentage points as within noise on the total-bias bar chart. Per-faith bar charts are computed by averaging across the 20 models, and their error bars reflect model-to-model variation for each faith.

The score shown in this graph sums those departures from neutral across all 14 faiths and expresses the total as a fraction of the worst-case possible (the most biased a model could be on every faith). Each bar shows how far a model's responses stray from neutral on average, across all 14 faiths, regardless of direction of bias (encouraging users to join or leave, or discouraging them from joining or leaving). Higher = more biased; lower = more even-handed.

Bias by model and faith

This table shows bias by model and faith across 3,640 evaluations of 14 faiths and 20 models.

How precise are these numbers?

Each cell is the average of N=26 pairwise ratings (13 join-related and 13 leave-related, both directions). Hover any cell to see its 95% confidence interval and N — the interval is a normal CI (mean ± 1.96·SE), with the lower bound clipped at 0 on the Positive and Negative views. Treat differences between two cells smaller than about 10–12 percentage points as within noise. As with the other charts, these intervals capture sampling variability over the questions only.

Each cell shows how far a model's responses about one faith stray from neutral, regardless of direction of bias. The score is the average distance from neutral across all ratings involving that faith, expressed as a percentage of the maximum bias possible. A score of 0% would mean that the model, on average, showed no bias in responses for that faith. A score of 20% would mean that the model's bias — counting both pulls toward the faith and pulls away from it — averaged 20% of the maximum possible deviation from neutral. Darker = more opinionated about the faith; lighter = closer to neutral. A darker cell could mean the model strongly endorses the faith, strongly criticizes it, or both across different questions.

Neutral (0%) Most biased (100%)

Total bias per faith

Each bar shows how far the 20 AI models strayed from neutral on average when responding about each faith, regardless of direction. Lower = more even-handed; higher = more opinionated overall. Green = the share of that bias that pulls toward the faith (encouraging users to join it, encouraging them to stay); red = the share that pulls away (discouraging users from joining, encouraging them to leave). The two segments sum to the total bias. Faiths are sorted from those treated with the least bias (left) to those treated with the most (right).

How precise are these numbers?

Each bar averages across all 20 models; the error bars show 95% CI of that average (SE = sd / √20), reflecting model-to-model variation for that faith. Median half-width is about ±5 percentage points. Because the CI captures spread between models, a wide bar means the models disagree about that faith, while a narrow bar means they largely agree. Treat differences between two faiths smaller than about 7 percentage points as within noise.

Comparative bias of faiths

These charts show how each faith is treated by AI models compared to the other faiths — in other words, how fair (or unfair) the AIs are being to one faith relative to the rest. Both views below look only at questions about joining a faith — they do not include questions about staying in or leaving one. Encouraging = how often the AIs lean toward telling someone to join that faith; Discouraging = how often they lean toward telling them not to. Green bars mean the faith is treated more favorably than the average faith; red means less favorably.

How precise are these numbers?

Each bar is the average across the 20 models of how often that faith is encouraged (or discouraged) for joining, expressed as a multiple of the dataset-wide baseline. Error bars are a 95% CI of the model-aggregate (SE = sd / √20). The CIs here capture model-to-model disagreement: median half-width is roughly ±0.24× on the Encouraging view and ±0.60× on the Discouraging view (where rates are low and a single model swings the ratio noticeably). The Jehovah's Witness outlier on the Discouraging view sits far above its CI lower bound, so the gap to the next faith is real, but among the other 13 faiths, treat differences smaller than about 0.35× on the Encouraging view, or 0.85× on the Discouraging view, as within noise.