Scaling-up Perceptual Video Quality Assessment

Jia, Ziheng; Zhang, Zicheng; Zhang, Zeyu; Liang, Yingji; Zhu, Xiaorong; Li, Chunyi; Han, Jinliang; Wu, Haoning; Wang, Bin; Zhang, Haoran; Zhu, Guanyu; Zhao, Qiyong; Liu, Xiaohong; Zhai, Guangtao; Min, Xiongkuo

Computer Science > Computer Vision and Pattern Recognition

arXiv:2505.22543 (cs)

[Submitted on 28 May 2025]

Title:Scaling-up Perceptual Video Quality Assessment

Authors:Ziheng Jia, Zicheng Zhang, Zeyu Zhang, Yingji Liang, Xiaorong Zhu, Chunyi Li, Jinliang Han, Haoning Wu, Bin Wang, Haoran Zhang, Guanyu Zhu, Qiyong Zhao, Xiaohong Liu, Guangtao Zhai, Xiongkuo Min

View PDF HTML (experimental)

Abstract:The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptual video quality assessment (VQA), the potential of scaling law remains unprecedented due to the scarcity of labeled resources and the insufficient scale of datasets. To address this, we propose \textbf{OmniVQA}, an efficient framework designed to efficiently build high-quality, human-in-the-loop VQA multi-modal instruction databases (MIDBs). We then scale up to create \textbf{OmniVQA-Chat-400K}, the largest MIDB in the VQA field concurrently. Our focus is on the technical and aesthetic quality dimensions, with abundant in-context instruction data to provide fine-grained VQA knowledge. Additionally, we have built the \textbf{OmniVQA-MOS-20K} dataset to enhance the model's quantitative quality rating capabilities. We then introduce a \textbf{complementary} training strategy that effectively leverages the knowledge from datasets for quality understanding and quality rating tasks. Furthermore, we propose the \textbf{OmniVQA-FG (fine-grain)-Benchmark} to evaluate the fine-grained performance of the models. Our results demonstrate that our models achieve state-of-the-art performance in both quality understanding and rating tasks.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2505.22543 [cs.CV]
	(or arXiv:2505.22543v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2505.22543

Submission history

From: Ziheng Jia [view email]
[v1] Wed, 28 May 2025 16:24:52 UTC (9,778 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Scaling-up Perceptual Video Quality Assessment

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Scaling-up Perceptual Video Quality Assessment

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators