Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment
Moral Entropy treats moral-label disagreement as a posterior and finds common vote rules badly misclassify items.
The paper introduces Moral Entropy, a Bayesian framework that keeps a posterior over moral labels and splits entropy into aleatoric and epistemic uncertainty. Consensus rules can be audited with cross-entropy, KL divergence, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, the any-annotator rule disagrees with the calibrated posterior on about 30 percent of items, mostly false positives. Stricter majority and two-vote rules miss 63 to 83 percent of true positives; on MFTC the reported mean false-positive and false-negative rates are 19.9 and 38.9 percent.
- Moral Entropy splits entropy into aleatoric and epistemic uncertainty.
- Any-annotator rule disagrees with the posterior on about 30 percent of items.
- Majority and two-vote rules miss 63 to 83 percent of true positives.
- On MFTC, any-annotator mean FPR and FNR are 19.9 and 38.9 percent.
Full article176 words · extracted from arxiv.org · click to collapse
Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single annotator flags an item. We argue this uncertainty should instead be modeled and learned from. We introduce Moral Entropy, a Bayesian framework that keeps a full posterior over the true label and decomposes its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (from insufficient or noisy annotation) -- and lets any heuristic consensus rule be audited against a calibrated ground truth via entropy methods such as cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, auditing the standard aggregation rules against this posterior reveals bias that no current pipeline reports: the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items -- pooled, almost entirely false positives, though the errors invert at the foundation level (19.9%/38.9% mean FPR/FNR on MFTC) -- while the stricter majority and two-vote rules miss 63-83% of true positives.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.21992