SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

A new research introduces SPARC, a framework designed to address this challenge by decoupling visual perception and reasoning into two distinct computational stages. By separating these processes, the approach enables more efficient use of computational resources while preserving, and in many cases improving, model performance. 

More details on the methodology and experimental evaluation are provided in the paper “SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of Vision-Language Models” (arXiv:2602.06566).

Screenshot from 2026 03 09 09 38 49

Decoupling Perception and Reasoning

Traditional multimodal models often lead to inefficient computation and large memory requirements.

The SPARC framework proposes a different design inspired by biological neural systems. Early visual processing extracts relevant information from images, while higher cognitive functions interpret this information to reach conclusions.

The framework introduces a two-stage pipeline:

  1. Perception Stage: This stage identifies key visual elements by producing coordinates for image regions that likely contain important information.
  2. Reasoning Stage: Once the relevant regions are identified, the model performs reasoning only on those selected visual crops.

This separation significantly reduces the amount of visual data processed by the reasoning component, enabling more efficient inference without sacrificing accuracy.

Efficiency Gains Through Structured Visual Context

One of the central insights of the SPARC framework is that accurate localization of relevant visual content can compensate for lower image resolution. Experimental results show that when the model focuses on the correct regions of an image, performance approaches that of full-resolution processing—even when the overall image resolution is significantly reduced. This means that instead of analyzing a full high-resolution image, the model can process smaller crops that capture the essential details needed for reasoning. As a result, the system dramatically reduces computational overhead while maintaining strong task performance.

Improved Performance with Lower Computational Cost

Extensive experiments demonstrate that the SPARC approach consistently improves visual reasoning accuracy compared with traditional multimodal inference pipelines. In particular, the framework shows strong results in scenarios that require detailed visual understanding.

In some experiments, models operating on reduced image resolution combined with targeted visual crops outperform baseline models processing full-resolution images, while using only a fraction of the computational resources. Rather than relying solely on larger models or higher resolution inputs, intelligent context engineering and modular inference pipelines can achieve better efficiency-performance trade-offs.

Towards Scalable and Sustainable AI Systems

Beyond its technical contributions, the SPARC framework reflects a broader trend in artificial intelligence research: designing systems that are both powerful and computationally efficient. By enabling models to process visual information more selectively and efficiently, approaches like SPARC help reduce unnecessary computation while preserving strong reasoning capabilities. The modular nature of the framework also enables independent optimization of the perception and reasoning components.

Looking Ahead

As multimodal AI continues to expand into fields such as autonomous systems, environmental monitoring, and scientific analysis, innovations like SPARC will play a key role in ensuring that these technologies remain both powerful and computationally sustainable.

Logo ALMA

SustainML is among these nine innovative projects dedicated to creating a sustainable ML framework for Green AI.

Coordinator Office Address

Plaza de la Encina 10-11, Núcleo 4, 2ª Pl.
28760 Tres cantos - Madrid (España)

  • X
  • LinkedIn

EN-Funded_by_the_EU-POSThis project has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No 101070408.