Conditional Rank Allocation for Taxonomy-Aware Medical Language Model Adaptation
ARBOR's conditional rank allocation lifts Qwen3-8B medical QA accuracy to 69.69% across four benchmarks.
ARBOR is a parameter-efficient method that selects rank-one components from a shared low-rank basis for each medical question, using an additive gate over question text, specialty tags, operation tags, and their interaction. An orthogonal-subtask illustration motivates conditional selection over a fixed update of the same active rank, without claiming that bound for medical corpora. Five-seed experiments on Qwen3-8B across CMB, CMExam, MedQA, and MedMCQA reach 69.69% mean accuracy, exceeding LoRA r16 and MoELoRA by 1.26 and 1.30 points. The LoRA advantage grows from 0.08 to 1.94 points from one to seven specialties, and atom clusters match specialty labels with adjusted Rand index 0.62; clinical safety remains untested.
- ARBOR selects rank-one adapter components using specialty and operation tags.
- Qwen3-8B averages 69.69% on CMB, CMExam, MedQA, and MedMCQA.
- It beats LoRA r16 by 1.26 points and MoELoRA by 1.30 points.
- The LoRA gap grows from 0.08 to 1.94 points across seven specialties.
- Atom clusters align with specialty labels at adjusted Rand index 0.62.
Full article190 words · extracted from arxiv.org · click to collapse
Medical question answering spans specialties and clinical operations that may benefit from different adaptation directions. We propose ARBOR, a parameter-efficient method that selects rank-one components from a shared low-rank basis for each question. An additive gate combines question representations, specialty tags, operation tags, and their interaction; a learned coefficient scales the adapter residual. An illustrative separation under orthogonal, equiprobable subtasks shows how conditional selection can avoid an approximation floor faced by a fixed update with the same active rank. This result motivates the design without asserting a corresponding bound for medical corpora. On Qwen3-8B across CMB, CMExam, MedQA, and MedMCQA, five-seed experiments yield 69.69% mean accuracy across benchmarks, exceeding LoRA r16 and MoELoRA by 1.26 and 1.30 percentage points, respectively. The reported advantage over LoRA r16 increases from 0.08 to 1.94 points as training expands from one to seven specialties. Tag perturbations and atom masking support the usefulness of clinical routing, while atom clusters align with the supplied specialty labels (adjusted Rand index 0.62). Calibration, transfer, and measured costs further characterize the method. These findings support structured conditional adaptation for medical QA, while leaving clinical safety and broader deployment untested.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.06765