apple/LensVLM-9B — new model trending #30 on Hugging Face
Apple released LensVLM-9B, a vision-language model that selectively expands compressed document pages.
Apple published LensVLM-9B on Hugging Face, a 9-billion-parameter vision-language model for reading compressed images of text. It scans compressed pages and uses learned tools to expand only relevant pages, with 5x, 10x, and 15x compression options. Weights, including Apple's modifications to a Qwen model, are under the Apple Machine Learning Research Model License. A companion paper (arXiv:2605.07019) and inference code were released with the model.
- 9B VLM scans compressed document images then expands relevant pages
- Supports 5x, 10x, and 15x compression settings
- Weights modify a Qwen model under Apple's research license
- Paper arXiv:2605.07019 and inference code published with the weights
Full article207 words · extracted from huggingface.co · click to collapse
# LensVLM-9B
LensVLM is a 9B Vision Language Model (VLM) that scans compressed images of text,
then selectively expands only the relevant pages to their uncompressed form via
learned tools.
- Paper: [LensVLM: Selective Context Expansion for Compressed Visual Representation of Text](https://arxiv.org/abs/2605.07019)
- Code: https://github.com/apple-aiml-research/ml-lensvlm
## License
All ML model files in this repository, including Apple's modifications to the Qwen
model, are provided under the terms of the
[Apple Machine Learning Research Model License](https://huggingface.co/apple/LensVLM-9B/blob/main/LICENSE).
The source code that accompanies this model is distributed separately and is provided
under the terms of the Apple Sample Code License.
## Usage
Install the LensVLM code and run inference:
```bash
git clone https://github.com/apple-aiml-research/ml-lensvlm
cd ml-lensvlm
pip install -r requirements.txt
python scripts/run_demo.py --model apple/LensVLM-9B
```
For a custom document:
```bash
python demo.py \
--model apple/LensVLM-9B \
--text_file document.txt \
--question "What is the main finding?" \
--compression 10x
```
Compression options: `5x`, `10x`, `15x`. See the
[repository README](https://github.com/apple-aiml-research/ml-lensvlm) for data preparation
and evaluation.
## Citation
```bibtex
@article{xie2026lensvlm,
title={LensVLM: Selective Context Expansion for Compressed Visual Representation of Text},
author={Xie, Roy and Friedman, Dan and Yu, Donghan and Pan, Bowen and Fifty, Christopher and Kim, Jang-Hyun and Du, Xianzhi and Gan, Zhe and Rathod, Vivek and Dhingra, Bhuwan},
journal={arXiv preprint arXiv:2605.07019},
year={2026}
}
```
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/apple/LensVLM-9B