Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales
Fine-tuned DFM-large retrieves biblical allusions in Blixen at R@10 0.508 on a 189-reference benchmark.
The paper builds a benchmark of 189 annotated biblical references in Karen Blixen's Seven Gothic Tales and retrieves them against 31,170 verses from historically plausible Danish translations. Linguistically normalized BM25 reaches overall R@10 of 0.365 and retrieves every quotation in its top ten, while the best zero-shot dense model scores 0.360 and does better on allusions. Fine-tuning DFM-large raises overall R@10 from 0.265 to 0.508 and allusion R@10 from 0.138 to 0.339. A literary scholar judged seven of 30 selected rank-one false positives to be meaningful additional references.
- Benchmark has 189 references against 31,170 Danish Bible verses.
- Normalized BM25 reaches overall R@10 of 0.365 and finds every quotation.
- Fine-tuned DFM-large lifts overall R@10 from 0.265 to 0.508.
- A scholar judged 7 of 30 false positives as meaningful references.
Full article245 words · extracted from arxiv.org · click to collapse
Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Tales. Drawing on the commentary to a critical edition, we construct a benchmark of 189 annotated references and evaluate retrieval against all 31,170 verses of historically plausible Danish Old and New Testament translations. We compare TF-IDF and BM25 with multilingual and Danish sentence encoders, examine the effect of linguistic normalization, and fine-tune a Danish encoder using hard negatives and five-fold cross-validation. We analyze performance across automatically derived lexical-overlap strata representing quotations, paraphrases, and allusions. Linguistically normalized BM25 provides a strong zero-shot baseline, attaining an overall R@10 of 0.365 and retrieving every quotation within its ten highest-ranked verses. The best zero-shot dense model achieves a comparable overall score of 0.360 while performing better on allusions. Fine-tuning DFM-large raises its overall R@10 from 0.265 to 0.508 and more than doubles its performance on allusions, from 0.138 to 0.339. However, evaluation against editorial annotations alone understates the model's scholarly usefulness: a literary scholar judged seven of 30 selected rank-one predictions counted as false positives to be meaningful additional references. These findings show both the potential and the epistemic limits of computational intertextual retrieval. Rather than treating scholarly annotations as exhaustive or model outputs as discoveries, we propose retrieval models as heuristic co-readers that recover documented references and generate candidates for expert-led close reading.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35765