LoRA-generating hypernetworks for efficient on-device LLM generative personalization
A hypernetwork generates user-specific LoRA adapters on device for personalized long-form LLM generation.
The paper trains a hypernetwork that maps a user's context tokens to a low-rank adaptation suited to that user. After shared artifacts are deployed, each device synthesizes its own LoRA with forward passes only, combining the feasibility of in-context learning with weight-level customization from parameter-efficient fine-tuning. The method reuses the target LLM's weights to limit extra storage and is evaluated on long-form text generation personalization against ICL and PEFT baselines.
- A hypernetwork maps user context tokens to a personalized LoRA.
- On-device personalization needs only forward passes, not full fine-tuning.
- Generated LoRA weights avoid extra latency from longer in-context prompts.
- The design reuses target LLM weights to limit added storage.
- Experiments target harder long-form generation personalization tasks.
Full article276 words · extracted from arxiv.org · click to collapse
On-device large language models (`LLMs'), e.g. running on mobile phones, are ripe for improvement via personalization. The limited compute resources of mobile devices impose limits on model scale and thus model quality, making any realizable quality gains highly impactful. At the same time, their personal nature (i.e., the close coupling to a particular user) means that a given on-device LLM tends to be used in similar, predictable patterns over the course of time. This paper presents a novel method for personalizing on-device LLMs. It trains a hypernetwork to map a user's context tokens to a low-rank adaptation (`LoRA') well-suited to that user. Once the trained common artifacts are deployed to users' devices, each user uses the hypernetwork to synthesize (entirely on device) a personalized LoRA. This approach blends the benefits while avoiding the drawbacks of two existing approaches to LLM customization: in-context learning (`ICL') and parameter-efficient fine-tuning (`PEFT'). Like ICL (and unlike PEFT), the on-device phase of our approach is computationally feasible, requiring only forward passes through neural networks. Like PEFT (and unlike ICL), our approach modifies the `target' base LLM via weights (the LoRA), avoiding negative consequences (e.g. increased latency) associated with extending the input sequence. Our approach is particularly well-suited to the mobile device regime. Apart from the on-device compute and latency benefits mentioned, it also requires minimal additional storage, as internally its architecture partly leverages the same LLM weights as belong to the target LLM to be personalized. We demonstrate the benefits of LoRA-generating hypernetworks on several representative personalization datasets, comparing against baselines like ICL and PEFT. Of note, our personalization experiments focus on more challenging and less studied long-form text generation tasks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.24979