Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
Chinese-Jev, a lightweight encoder-only decision model, beats closed-source Jev on Chinese tasks with up to 20x speedup.
Chinese-Jev is a lightweight encoder-only System One model for Chinese decision tasks, trained by converting annotations into probability targets over candidate options. After pretraining on 10 million examples it exceeds closed-source Jev by 1.24% on general-domain tasks with a 20.3x speedup. Domain fine-tuning adds 4.0% accuracy in medicine and reaches 92% of Jev's specialized-domain average at 17x speed and 15 ms average latency. An INT8 mobile build decides in about 1.0 second; evaluation uses the new Chinese-Jev Bench.
- Encoder-only System One model scores candidate answers rather than generating text.
- Pretrained on 10 million examples, then fine-tuned for medicine, law, and finance.
- Beats closed-source Jev by 1.24% generally and 4.0% in medicine, with about 17-20x speedup.
- INT8 mobile deployment averages about 1.0 second per decision; server latency is 15 ms.
Full article249 words · extracted from huggingface.co · click to collapse
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized settings. In this paper, we introduce Chinese-Jev, a System One model that addresses this gap through a unified data processing and training pipeline. Our data processing protocol converts heterogeneous Chinese-language annotations into probability targets over candidate options, enabling a shared training formulation across domains and question formats. To enable efficient inference, Chinese-Jev adopts a lightweight encoder-only backbone for text encoding and learns to score candidate answers through decision-oriented training. To address the misalignment between the pre-training distribution and downstream Chinese-language scenarios, we first train the model on a general-purpose corpus of 10 million examples, then fine-tune it separately for the medical, legal, and financial domains. To evaluate decision accuracy and calibration in both general and domain-specific Chinese-language settings, we introduce Chinese-Jev Bench (CJ-Bench). After first-stage pre-training, Chinese-Jev exceeds the accuracy of the closed-source Jev model by 1.24% on general-domain tasks while achieving a 20.3x speedup. Subsequent domain-specific fine-tuning yields a 4.0% accuracy improvement over Jev in medicine and achieves 92% of Jev's average accuracy across specialized domains, with a 17x speedup and an average latency of only 15 ms per example. We further demonstrate on-device deployment of an INT8-quantized model on mobile devices, achieving an inference latency of approximately 1.0 second per decision. The project is available at https://gulucaptain.github.io/Chinese-Jev/.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.36965