arXiv cs.AI→ original

ImagingBench revealed weaknesses of Gemini, GPT, and Qwen in physical image processing tasks

Researchers created ImagingBench—a benchmark of 20 computational photography and image processing tasks to test modern AI model capabilities. Testing showed that Gemini, GPT-5, and Qwen, despite strong semantic abilities, significantly lag behind specialized methods on computational sensing tasks (holography, lensless imaging, and time-of-flight).

AI-processed from arXiv cs.AI; edited by Hamidun News
ImagingBench revealed weaknesses of Gemini, GPT, and Qwen in physical image processing tasks
Source: arXiv cs.AI. Collage: Hamidun News.
◐ Listen to article

Researchers introduced ImagingBench — a systematic benchmark for evaluating the capabilities of vision-language models (VLMs) and agentic AI in solving computational photography and image processing tasks. The benchmark includes 20 tasks covering five categories: ray optics and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration.

What the benchmark revealed

Testing leading proprietary and open multimodal systems, including Gemini 3.1, GPT-5, and Qwen, showed that agentic models consistently underperform specialized methods. Particularly striking are the differences in computational sensing tasks, such as lensless imaging, event-based reconstruction, time-of-flight imaging, and holography.

  • 20 computational photography tasks in the benchmark
  • Tested Gemini, GPT-5, Qwen, and other VLMs
  • Agentive guidance (Planner) provided only modest improvements
  • Visually plausible but physically inaccurate results

The gap between semantics and physics

Models often generate visually acceptable results that appear correct to the human eye, but when checked against reference-based metrics show poor quality. This reveals a massive gap between models' ability to work with semantic visual tasks (what they see) and physically grounded visualization (how physical systems work).

Development roadmap

ImagingBench provides a unified testbed for measuring the progress of agentic AI in computational photography. This allows researchers to systematically track improvements in the ability of modern LLM models to understand and solve physical tasks as opposed to purely semantic ones.

What this means

The discovery shows that current LLM models work well with semantic visual tasks but remain insufficiently adapted to physically grounded problems. This opens new directions for model improvement and points to the need for specialized approaches in computational photography, even in the era of general-purpose AI agents.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…