Jev for Scientific Decisions: Evaluating Semantic Choices and Their Consequences
Jev matches peer models on semantic scientific choices, but wrong selections can still change reused quantities.
The paper evaluates Jev as a semantic decision component that selects scientific relations before arithmetic is executed in code. Twelve model configurations were compared on 20 source-grounded choices across 10 scientific cases, each repeated five times, with selections, downstream outputs, and final claim labels measured separately. Jev matched five configurations on complete semantic correctness and had the lowest median latency among successful responses, while some wrong selections changed downstream counts without changing the final label.
- Twelve model configurations were tested on twenty source-grounded choices.
- Ten scientific cases were each repeated five times.
- Jev tied five configurations on semantic correctness with the lowest median latency.
- Seven wrong culture-history selections changed counts but kept the correct label.
Full article153 words · extracted from arxiv.org · click to collapse
Scientific workflows often require choosing among known relations before a deterministic calculation can proceed. Whether observations share a culture, treatment or reference standard can change the scientific meaning of the resulting count or comparison. We evaluate Jev as a semantic decision component using a harness that follows its documented guidance and assigns arithmetic to code. The study compares twelve model configurations on twenty source-grounded Choices across ten scientific cases, each repeated five times. We measure semantic selections, downstream outputs and final claim labels separately. Jev matched five other configurations at complete semantic correctness and achieved the lowest observed median latency among successful responses. Across three comparison models, seven wrong selections on one culture-history question changed downstream counts while preserving the correct final label. These results identify a useful role for Jev in prepared scientific decision tasks and show why evaluating that role requires checking the relations and quantities that a workflow will reuse.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.24965