apple/LensVLM-9B

Apple releases LensVLM-9B, a vision-language model that compresses text images and selectively expands relevant pages for processing. The repository provides installation instructions and a demo script supporting 5x to 15x compression ratios. It is based on Qwen3.5-9B and includes a research paper describing the selective context expansion method for document analysis.

Apple developed LensVLM-9B to analyze documents by first processing compressed visual representations of text. The system then applies learned tools to selectively expand specific pages back to their original uncompressed state. This approach aims to reduce computational load during the initial scanning phase. A research paper describes the selective context expansion method used for this document analysis workflow. Users can install the code and run inference to test the vision-language capabilities of the model. A demo script allows inputting custom text files and specific questions while setting a desired compression level. The documented options include ratios of 5x, 10x, or 15x for controlling the compression intensity. The base architecture relies on the Qwen3.5-9B model as noted in the card description. The license terms require attention from both the model files and the accompanying source code. Apple Machine Learning Research Model License provisions apply to the ML model files including modifications. The source code operates under a separate Apple Sample Code License agreement. Specific performance benchmarks or accuracy metrics across different document types remain unspecified in the provided material, making independent verification necessary.

README

apple/LensVLM-9B View on Hugging Face

Loading the README from Hugging Face…