Skip to main content
Cornell University
We gratefully acknowledge support from the Simons Foundation, member institutions, and all contributors. Donate
arxiv logo > cs.MM

Help | Advanced Search

arXiv logo
Cornell University Logo

quick links

  • Login
  • Help Pages
  • About

Multimedia

Authors and titles for May 2024

Total of 100 entries
Showing up to 2000 entries per page: fewer | more | all
[1] arXiv:2405.00344 [pdf, html, other]
Title: Expert Insight-Enhanced Follow-up Chest X-Ray Summary Generation
Zhichuan Wang, Kinhei Lee, Qiao Deng, Tiffany Y. So, Wan Hang Chiu, Yeung Yu Hui, Bingjing Zhou, Edward S. Hui
Comments: accepted by 22nd International Conference on Artificial Intelligence in medicine (AIME2024)
Subjects: Multimedia (cs.MM)
[2] arXiv:2405.03500 [pdf, html, other]
Title: A Rate-Distortion-Classification Approach for Lossy Image Compression
Yuefeng Zhang
Comments: 15 pages
Journal-ref: Digital Signal Processing Volume 141, September 2023, 104163
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Information Theory (cs.IT)
[3] arXiv:2405.04279 [pdf, html, other]
Title: Task Presentation and Human Perception in Interactive Video Retrieval
Nina Willis, Abraham Bernstein, Luca Rossetto
Subjects: Multimedia (cs.MM)
[4] arXiv:2405.04963 [pdf, html, other]
Title: Audio Matters Too! Enhancing Markerless Motion Capture with Audio Signals for String Performance Capture
Yitong Jin, Zhiping Qiu, Yi Shi, Shuangpeng Sun, Chongwu Wang, Donghao Pan, Jiachen Zhao, Zhenghao Liang, Yuan Wang, Xiaobing Li, Feng Yu, Tao Yu, Qionghai Dai
Comments: SIGGRAPH2024
Subjects: Multimedia (cs.MM)
[5] arXiv:2405.05170 [pdf, html, other]
Title: Picking watermarks from noise (PWFN): an improved robust watermarking model against intensive distortions
Sijing Xie, Chengxin Zhao, Nan Sun, Wei Li, Hefei Ling
Subjects: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
[6] arXiv:2405.07229 [pdf, html, other]
Title: MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
Xiaocui Yang, Wenfang Wu, Shi Feng, Ming Wang, Daling Wang, Yang Li, Qi Sun, Yifei Zhang, Xiaoming Fu, Soujanya Poria
Comments: Accepted by the Information Fusion Journal
Subjects: Multimedia (cs.MM)
[7] arXiv:2405.07689 [pdf, html, other]
Title: Quality of Experience Optimization for Real-time XR Video Transmission with Energy Constraints
Guangjin Pan, Shugong Xu, Shunqing Zhang, Xiaojing Chen, Yanzan Sun
Comments: 6 pages, 5 figures
Subjects: Multimedia (cs.MM); Networking and Internet Architecture (cs.NI); Systems and Control (eess.SY)
[8] arXiv:2405.07759 [pdf, html, other]
Title: MADRL-Based Rate Adaptation for 360° Video Streaming with Multi-Viewpoint Prediction
Haopeng Wang, Zijian Long, Haiwei Dong, Abdulmotaleb El Saddik
Comments: Accepted by IEEE Internet of Things Journal
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Networking and Internet Architecture (cs.NI); Image and Video Processing (eess.IV)
[9] arXiv:2405.07827 [pdf, html, other]
Title: Automatic Recognition of Food Ingestion Environment from the AIM-2 Wearable Sensor
Yuning Huang, Mohamed Abul Hassan, Jiangpeng He, Janine Higgins, Megan McCrory, Heather Eicher-Miller, Graham Thomas, Edward O Sazonov, Fengqing Maggie Zhu
Comments: Accepted at CVPRw 2024
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[10] arXiv:2405.07930 [pdf, html, other]
Title: Improving Multimodal Learning with Multi-Loss Gradient Modulation
Konstantinos Kontras, Christos Chatzichristos, Matthew Blaschko, Maarten De Vos
Subjects: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[11] arXiv:2405.09286 [pdf, html, other]
Title: MVBIND: Self-Supervised Music Recommendation For Videos Via Embedding Space Binding
Jiajie Teng, Huiyu Duan, Yucheng Zhu, Sijing Wu, Guangtao Zhai
Subjects: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV)
[12] arXiv:2405.10029 [pdf, html, other]
Title: AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
Ziyu Gong, Chengcheng Mai, Yihua Huang
Comments: This work has been strong-accepted as the oral conference paper by IEEE International Conference on Multimedia & Expo (ICME) 2024
Subjects: Multimedia (cs.MM)
[13] arXiv:2405.10497 [pdf, html, other]
Title: SMP Challenge: An Overview and Analysis of Social Media Prediction Challenge
Bo Wu, Peiye Liu, Wen-Huang Cheng, Bei Liu, Zhaoyang Zeng, Jia Wang, Qiushi Huang, Jiebo Luo
Comments: ACM Multimedia. arXiv admin note: text overlap with arXiv:1910.01795
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Social and Information Networks (cs.SI)
[14] arXiv:2405.11742 [pdf, html, other]
Title: Universal Organizer of SAM for Unsupervised Semantic Segmentation
Tingting Li, Gensheng Pei, Xinhao Cai, Huafeng Liu, Qiong Wang, Yazhou Yao
Comments: accepted by IEEE International Conference on Multimedia & Expo
Subjects: Multimedia (cs.MM)
[15] arXiv:2405.12775 [pdf, html, other]
Title: Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
Hanlei Zhang, Hua Xu, Fei Long, Xin Wang, Kai Gao
Comments: Accepted by ACL 2024, Main Conference, Long Paper
Subjects: Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
[16] arXiv:2405.14040 [pdf, html, other]
Title: Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
Dingyi Yang, Chunru Zhan, Ziheng Wang, Biao Wang, Tiezheng Ge, Bo Zheng, Qin Jin
Comments: 15 pages, 13 figures
Journal-ref: https://aclanthology.org/2024.acl-long.513/
Subjects: Multimedia (cs.MM)
[17] arXiv:2405.15026 [pdf, html, other]
Title: Enhancing Student Feedback Using Predictive Models in Visual Literacy Courses
Alon Friedman, Kevin Hawley, Paul Rosen, Md Dilshadur Rahman
Comments: 8 pages, 6 figures, IEEE EDUCON 2024 conference
Subjects: Multimedia (cs.MM); Computers and Society (cs.CY)
[18] arXiv:2405.17147 [pdf, html, other]
Title: Large Language Models (LLMs): Deployment, Tokenomics and Sustainability
Haiwei Dong, Shuang Xie
Comments: Accepted by IEEE CTSoc-NCT
Subjects: Multimedia (cs.MM)
[19] arXiv:2405.19802 [pdf, html, other]
Title: Exploring the Robustness of Decision-Level Through Adversarial Attacks on LLM-Based Embodied Models
Shuyuan Liu, Jiawei Chen, Shouwei Ruan, Hang Su, Zhaoxia Yin
Subjects: Multimedia (cs.MM)
[20] arXiv:2405.20078 [pdf, other]
Title: NeRF View Synthesis: Subjective Quality Assessment and Objective Metrics Evaluation
Pedro Martin, Antonio Rodrigues, Joao Ascenso, Maria Paula Queluz
Subjects: Multimedia (cs.MM)
[21] arXiv:2405.00233 (cross-list from cs.SD) [pdf, html, other]
Title: SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
Haohe Liu, Xuenan Xu, Yi Yuan, Mengyue Wu, Wenwu Wang, Mark D. Plumbley
Comments: Accepted by Journal of Selected Topics in Signal Processing (JSTSP). Demo and code: this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Audio and Speech Processing (eess.AS); Signal Processing (eess.SP)
[22] arXiv:2405.00248 (cross-list from cs.SD) [pdf, html, other]
Title: Who is Authentic Speaker
Qiang Huang
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[23] arXiv:2405.00351 (cross-list from cs.HC) [pdf, html, other]
Title: Learning High-Quality Navigation and Zooming on Omnidirectional Images in Virtual Reality
Zidong Cao, Zhan Wang, Yexin Liu, Yan-Pei Cao, Ying Shan, Wei Zeng, Lin Wang
Comments: 11 pages
Subjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[24] arXiv:2405.00384 (cross-list from cs.CV) [pdf, html, other]
Title: Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol
Konstantinos Apostolidis, Jakob Abesser, Luca Cuccovillo, Vasileios Mezaris
Comments: Accepted for publication, 3rd ACM Int. Workshop on Multimedia AI against Disinformation (MAD'24) at ACM ICMR'24, June 10, 2024, Phuket, Thailand. This is the "accepted version"
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[25] arXiv:2405.00483 (cross-list from cs.CV) [pdf, other]
Title: In Anticipation of Perfect Deepfake: Identity-anchored Artifact-agnostic Detection under Rebalanced Deepfake Detection Protocol
Wei-Han Wang, Chin-Yuan Yeh, Hsi-Wen Chen, De-Nian Yang, Ming-Syan Chen
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[26] arXiv:2405.00574 (cross-list from cs.CV) [pdf, html, other]
Title: EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
Deng Li, Xin Liu, Bohao Xing, Baiqiang Xia, Yuan Zong, Bihan Wen, Heikki Kälviäinen
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[27] arXiv:2405.03318 (cross-list from cs.CV) [pdf, html, other]
Title: Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation
Yingying Zhang, Chuangji Shi, Xin Guo, Jiangwei Lao, Jian Wang, Jiaotuan Wang, Jingdong Chen
Comments: 11 pages, 7 figures
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[28] arXiv:2405.03436 (cross-list from cs.CV) [pdf, html, other]
Title: DBDH: A Dual-Branch Dual-Head Neural Network for Invisible Embedded Regions Localization
Chengxin Zhao, Hefei Ling, Sijing Xie, Nan Sun, Zongyi Li, Yuxuan Shi, Jiazhong Chen
Comments: 7 pages, 6 figures (Have been accepted by IJCNN 2024)
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[29] arXiv:2405.03920 (cross-list from cs.CL) [pdf, html, other]
Title: A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
Dainis Boumber, Rakesh M. Verma, Fatima Zahra Qachfar
Comments: 6 pages, 1 figure, shorter version in SIAM International Conference on Data Mining (SDM) 2024
Journal-ref: Proc. SDM 2024, 396-399
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
[30] arXiv:2405.04097 (cross-list from cs.CV) [pdf, html, other]
Title: Unmasking Illusions: Understanding Human Perception of Audiovisual Deepfakes
Ammarah Hashmi, Sahibzada Adil Shahzad, Chia-Wen Lin, Yu Tsao, Hsin-Min Wang
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG); Multimedia (cs.MM)
[31] arXiv:2405.05039 (cross-list from cs.CV) [pdf, html, other]
Title: Reviewing Intelligent Cinematography: AI research for camera-based video production
Adrian Azzarelli, Nantheera Anantrasirichai, David R Bull
Comments: This paper has been accepted for publication with "Artificial Intelligence Review" Journal (this https URL) and we are in the procress of publishing it
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[32] arXiv:2405.05130 (cross-list from cs.CV) [pdf, html, other]
Title: Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection
Shengyang Sun, Xiaojin Gong
Comments: Accepted by ICME 2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[33] arXiv:2405.05244 (cross-list from eess.AS) [pdf, html, other]
Title: SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan
You Zhang, Yongyi Zang, Jiatong Shi, Ryuichi Yamamoto, Jionghao Han, Yuxun Tang, Tomoki Toda, Zhiyao Duan
Comments: Evaluation plan of the SVDD Challenge @ SLT 2024
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Sound (cs.SD)
[34] arXiv:2405.05524 (cross-list from cs.CV) [pdf, html, other]
Title: Universal Adversarial Perturbations for Vision-Language Pre-trained Models
Peng-Fei Zhang, Zi Huang, Guangdong Bai
Comments: 9 pages, 5 figures
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[35] arXiv:2405.05691 (cross-list from cs.CV) [pdf, html, other]
Title: StableMoFusion: Towards Robust and Efficient Diffusion-based Motion Generation Framework
Yiheng Huang, Hui Yang, Chuanchen Luo, Yuxi Wang, Shibiao Xu, Zhaoxiang Zhang, Man Zhang, Junran Peng
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[36] arXiv:2405.06143 (cross-list from cs.CV) [pdf, html, other]
Title: Perceptual Crack Detection for Rendered 3D Textured Meshes
Armin Shafiee Sarvestani, Wei Zhou, Zhou Wang
Comments: Accepted by IEEE QoMEX 2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computational Geometry (cs.CG); Multimedia (cs.MM)
[37] arXiv:2405.06217 (cross-list from cs.CV) [pdf, html, other]
Title: DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding
Ting Liu, Xuyang Liu, Siteng Huang, Honggang Chen, Quanjun Yin, Long Qin, Donglin Wang, Yue Hu
Comments: Accepted by ICME 2024 (Oral)
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[38] arXiv:2405.06995 (cross-list from cs.SD) [pdf, html, other]
Title: Benchmarking Cross-Domain Audio-Visual Deception Detection
Xiaobao Guo, Zitong Yu, Nithish Muthuchamy Selvaraj, Bingquan Shen, Adams Wai-Kin Kong, Alex C. Kot
Comments: 12 pages
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[39] arXiv:2405.07202 (cross-list from cs.CV) [pdf, html, other]
Title: Unified Video-Language Pre-training with Synchronized Audio
Shentong Mo, Haofan Wang, Huaxia Li, Xu Tang
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[40] arXiv:2405.07354 (cross-list from cs.SD) [pdf, html, other]
Title: SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
Sushant Gautam, Mehdi Houshmand Sarkhoosh, Jan Held, Cise Midoglu, Anthony Cioppa, Silvio Giancola, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen, Mubarak Shah
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[41] arXiv:2405.07682 (cross-list from cs.SD) [pdf, html, other]
Title: FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
Jianyi Chen, Wei Xue, Xu Tan, Zhen Ye, Qifeng Liu, Yike Guo
Comments: IJCAI 2024
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[42] arXiv:2405.08465 (cross-list from cs.IR) [pdf, html, other]
Title: How to Surprisingly Consider Recommendations? A Knowledge-Graph-based Approach Relying on Complex Network Metrics
Oliver Baumann, Durgesh Nandini, Anderson Rossanez, Mirco Schoenfeld, Julio Cesar dos Reis
Subjects: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Social and Information Networks (cs.SI)
[43] arXiv:2405.08555 (cross-list from cs.CV) [pdf, html, other]
Title: Dual-Branch Network for Portrait Image Quality Assessment
Wei Sun, Weixia Zhang, Yanwei Jiang, Haoning Wu, Zicheng Zhang, Jun Jia, Yingjie Zhou, Zhongpeng Ji, Xiongkuo Min, Weisi Lin, Guangtao Zhai
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[44] arXiv:2405.08619 (cross-list from cs.CL) [pdf, html, other]
Title: ALMol: Aligned Language-Molecule Translation LLMs through Offline Preference Contrastive Optimisation
Dimitris Gkoumas
Subjects: Computation and Language (cs.CL); Multimedia (cs.MM)
[45] arXiv:2405.08745 (cross-list from eess.IV) [pdf, html, other]
Title: Enhancing Blind Video Quality Assessment with Rich Quality-aware Features
Wei Sun, Haoning Wu, Zicheng Zhang, Jun Jia, Zhichao Zhang, Linhan Cao, Qiubo Chen, Xiongkuo Min, Weisi Lin, Guangtao Zhai
Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[46] arXiv:2405.08813 (cross-list from cs.CV) [pdf, html, other]
Title: CinePile: A Long Video Question Answering Dataset and Benchmark
Ruchit Rawal, Khalid Saifullah, Miquel Farré, Ronen Basri, David Jacobs, Gowthami Somepalli, Tom Goldstein
Comments: Project page with all the artifacts - this https URL. Updated version with adversarial refinement pipeline and more model evaluations
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM)
[47] arXiv:2405.09152 (cross-list from cs.CV) [pdf, html, other]
Title: Scalable Image Coding for Humans and Machines Using Feature Fusion Network
Takahiro Shindo, Taiju Watanabe, Yui Tatsumi, Hiroshi Watanabe
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[48] arXiv:2405.09191 (cross-list from cs.CR) [pdf, html, other]
Title: QMedShield: A Novel Quantum Chaos-based Image Encryption Scheme for Secure Medical Image Storage in the Cloud
Arun Amaithi Rajan, Vetriselvi V
Comments: 20 pages, 17 Figures, 9 Tables
Journal-ref: 2024
Subjects: Cryptography and Security (cs.CR); Multimedia (cs.MM)
[49] arXiv:2405.09266 (cross-list from cs.CV) [pdf, html, other]
Title: Dance Any Beat: Blending Beats with Visuals in Dance Video Generation
Xuanchen Wang, Heng Wang, Dongnan Liu, Weidong Cai
Comments: WACV2025, 11 pages, 7 figures, demo page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[50] arXiv:2405.09321 (cross-list from cs.CV) [pdf, html, other]
Title: ReconBoost: Boosting Can Achieve Modality Reconcilement
Cong Hua, Qianqian Xu, Shilong Bao, Zhiyong Yang, Qingming Huang
Comments: This paper has been accepted by ICML2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)
[51] arXiv:2405.09539 (cross-list from eess.IV) [pdf, html, other]
Title: MMFusion: Multi-modality Diffusion Model for Lymph Node Metastasis Diagnosis in Esophageal Cancer
Chengyu Wu, Chengkai Wang, Yaqi Wang, Huiyu Zhou, Yatao Zhang, Qifeng Wang, Shuai Wang
Comments: Early accepted to MICCAI 2024 (6/6/5)
Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[52] arXiv:2405.10121 (cross-list from cs.CL) [pdf, html, other]
Title: Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
Bo Zhang, Hui Ma, Jian Ding, Jian Wang, Bo Xu, Hongfei Lin
Comments: Accepted by Information Fusion. The code is available at this https URL
Subjects: Computation and Language (cs.CL); Multimedia (cs.MM)
[53] arXiv:2405.11093 (cross-list from eess.AS) [pdf, html, other]
Title: AudioSetMix: Enhancing Audio-Language Datasets with LLM-Assisted Augmentations
David Xu
Comments: typos corrected
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)
[54] arXiv:2405.11145 (cross-list from cs.CV) [pdf, html, other]
Title: Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
Junzhang Liu, Zhecan Wang, Hammad Ayyubi, Haoxuan You, Chris Thomas, Rui Sun, Shih-Fu Chang, Kai-Wei Chang
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
[55] arXiv:2405.11273 (cross-list from cs.AI) [pdf, html, other]
Title: Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma, Min Zhang
Comments: 22 pages, 13 figures. Project Website: this https URL. Working in progress
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[56] arXiv:2405.11295 (cross-list from eess.IV) [pdf, other]
Title: Medical Image Analysis for Detection, Treatment and Planning of Disease using Artificial Intelligence Approaches
Nand Lal Yadav, Satyendra Singh, Rajesh Kumar, Sudhakar Singh
Comments: 10 pages, 3 figures
Journal-ref: International Journal of Microsystems and IoT, Vol. 1, Issue 5, pp.278- 287, 2023
Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM)
[57] arXiv:2405.12126 (cross-list from cs.CV) [pdf, html, other]
Title: Alzheimer's Magnetic Resonance Imaging Classification Using Deep and Meta-Learning Models
Nida Nasir, Muneeb Ahmed, Neda Afreen, Mustafa Sameer
Subjects: Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET); Machine Learning (cs.LG); Multimedia (cs.MM)
[58] arXiv:2405.12221 (cross-list from cs.CV) [pdf, html, other]
Title: Images that Sound: Composing Images and Sounds on a Single Canvas
Ziyang Chen, Daniel Geng, Andrew Owens
Comments: Accepted to NeurIPS 2024. Project site: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[59] arXiv:2405.12336 (cross-list from cs.CR) [pdf, other]
Title: Interoperable Provenance Authentication of Broadcast Media using Open Standards-based Metadata, Watermarking and Cryptography
John C. Simmons, Joseph M. Winograd
Comments: 17 pages, 9 figures. Submitted to IBC2024 Technical Papers Programme
Journal-ref: IBC2024 Technical Papers Programme. https://www.ibc.org/technical-papers/ibc2024-tech-papers-interoperable-provenance-authentication-of-broadcast-media-using-open-standards-based-metadata-watermarking-and-cryptography/12063.article
Subjects: Cryptography and Security (cs.CR); Multimedia (cs.MM)
[60] arXiv:2405.12512 (cross-list from cs.CV) [pdf, html, other]
Title: Rethink Predicting the Optical Flow with the Kinetics Perspective
Yuhao Cheng, Siru Zhang, Yiqiang Yan
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
[61] arXiv:2405.12540 (cross-list from cs.CV) [pdf, html, other]
Title: Context-Enhanced Video Moment Retrieval with Large Language Models
Weijia Liu, Bo Miao, Jiuxin Cao, Xuelin Zhu, Bo Liu, Mehwish Nasim, Ajmal Mian
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[62] arXiv:2405.12564 (cross-list from q-bio.QM) [pdf, html, other]
Title: ProtT3: Protein-to-Text Generation for Text-based Protein Understanding
Zhiyuan Liu, An Zhang, Hao Fei, Enzhi Zhang, Xiang Wang, Kenji Kawaguchi, Tat-Seng Chua
Comments: ACL 2024, 9 pages
Subjects: Quantitative Methods (q-bio.QM); Computation and Language (cs.CL); Multimedia (cs.MM)
[63] arXiv:2405.12847 (cross-list from cs.IR) [pdf, html, other]
Title: A Dataset and Baselines for Measuring and Predicting the Music Piece Memorability
Li-Yang Tseng, Tzu-Ling Lin, Hong-Han Shuai, Jen-Wei Huang, Wen-Whei Chang
Journal-ref: Proceedings of the 24th International Society for Music Information Retrieval Conference, 174-181. Milan, Italy, November 5-9, 2023
Subjects: Information Retrieval (cs.IR); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[64] arXiv:2405.12983 (cross-list from eess.AS) [pdf, html, other]
Title: Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
Maxime Burchi, Krishna C. Puvvada, Jagadeesh Balam, Boris Ginsburg, Radu Timofte
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)
[65] arXiv:2405.13049 (cross-list from cs.CL) [pdf, html, other]
Title: SemEval-2024 Task 3: Multimodal Emotion Cause Analysis in Conversations
Fanfan Wang, Heqing Ma, Jianfei Yu, Rui Xia, Erik Cambria
Comments: Accepted to the 18th International Workshop on Semantic Evaluation (SemEval-2024). 12 pages, 3 figures, 4 Tables
Journal-ref: https://aclanthology.org/2024.semeval-1.277/
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
[66] arXiv:2405.13127 (cross-list from cs.CV) [pdf, html, other]
Title: Towards Retrieval-Augmented Architectures for Image Captioning
Sara Sarto, Marcella Cornia, Lorenzo Baraldi, Alessandro Nicolosi, Rita Cucchiara
Comments: ACM Transactions on Multimedia Computing, Communications and Applications (2024)
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
[67] arXiv:2405.13389 (cross-list from cs.CV) [pdf, html, other]
Title: HR-INR: Continuous Space-Time Video Super-Resolution via Event Camera
Yunfan Lu, Zipeng Wang, Yusheng Wang, Hui Xiong
Comments: 30 pages, 20 figures, 8 tables. This work was submitted for review in the second half of 2023. Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Robotics (cs.RO)
[68] arXiv:2405.13403 (cross-list from eess.IV) [pdf, html, other]
Title: Adaptive Wireless Image Semantic Transmission and Over-The-Air Testing
Jiarun Ding, Peiwen Jiang, Chao-Kai Wen, Shi Jin
Subjects: Image and Video Processing (eess.IV); Multimedia (cs.MM)
[69] arXiv:2405.13762 (cross-list from cs.CV) [pdf, html, other]
Title: A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual Generation
Gwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou, José Lezama, Agrim Gupta, Lijun Yu, Lu Jiang, Aren Jansen, Jacob Walker, Krishna Somandepalli
Journal-ref: In Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024)
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[70] arXiv:2405.13984 (cross-list from cs.CL) [pdf, html, other]
Title: Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI Evaluation
Dimitris Gkoumas, Maria Liakata
Comments: ACL25 Main
Subjects: Computation and Language (cs.CL); Multimedia (cs.MM)
[71] arXiv:2405.14225 (cross-list from q-bio.QM) [pdf, html, other]
Title: ReactXT: Understanding Molecular "Reaction-ship" via Reaction-Contextualized Molecule-Text Pretraining
Zhiyuan Liu, Yaorui Shi, An Zhang, Sihang Li, Enzhi Zhang, Xiang Wang, Kenji Kawaguchi, Tat-Seng Chua
Comments: ACL 2024 Findings, 9 pages
Subjects: Quantitative Methods (q-bio.QM); Computation and Language (cs.CL); Multimedia (cs.MM)
[72] arXiv:2405.14312 (cross-list from cs.CV) [pdf, html, other]
Title: Improving Gloss-free Sign Language Translation by Reducing Representation Density
Jinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang, Hui Xiong
Comments: Accepted at NeurIPS'24; Representation Density Problem and Performance Drop in Gloss-free SLT
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Multimedia (cs.MM)
[73] arXiv:2405.14556 (cross-list from cs.CV) [pdf, other]
Title: Deep Learning Classification of Photoplethysmogram Signal for Hypertension Levels
Nida Nasir, Mustafa Sameer, Feras Barneih, Omar Alshaltone, Muneeb Ahmed
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[74] arXiv:2405.14598 (cross-list from cs.CV) [pdf, html, other]
Title: Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
Shiqi Yang, Zhi Zhong, Mengjie Zhao, Shusuke Takahashi, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji
Comments: 10 pages
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[75] arXiv:2405.14709 (cross-list from cs.CV) [pdf, html, other]
Title: OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance
Shuheng Ge, Haoyu Xing, Li Zhang, Xiangqian Wu
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[76] arXiv:2405.15451 (cross-list from cs.CV) [pdf, html, other]
Title: Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
Yiming Wu, Hangfei Li, Fangfang Wang, Yilong Zhang, Ronghua Liang
Comments: ICASSP 2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR); Multimedia (cs.MM)
[77] arXiv:2405.15757 (cross-list from cs.CV) [pdf, html, other]
Title: Looking Backward: Streaming Video-to-Video Translation with Feature Banks
Feng Liang, Akio Kodaira, Chenfeng Xu, Masayoshi Tomizuka, Kurt Keutzer, Diana Marculescu
Comments: ICLR 2025. Project page: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[78] arXiv:2405.16000 (cross-list from cs.SD) [pdf, html, other]
Title: Carnatic Raga Identification System using Rigorous Time-Delay Neural Network
Sanjay Natesan, Homayoon Beigi
Comments: 7 pages, 2 tables, 3 figures
Journal-ref: Recognition Technologies, Inc. Technical Report (2024), RTI-20240524-01
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[79] arXiv:2405.16296 (cross-list from cs.CV) [pdf, html, other]
Title: Neural Network-Based Tracking and 3D Reconstruction of Baseball Pitch Trajectories from Single-View 2D Video
Jhen Hsieh
Subjects: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Multimedia (cs.MM)
[80] arXiv:2405.16640 (cross-list from cs.AI) [pdf, html, other]
Title: A Survey of Multimodal Large Language Model from A Data-centric Perspective
Tianyi Bai, Hao Liang, Binwang Wan, Yanran Xu, Xi Li, Shiyu Li, Ling Yang, Bozhou Li, Yifan Wang, Bin Cui, Ping Huang, Jiulong Shan, Conghui He, Binhang Yuan, Wentao Zhang
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[81] arXiv:2405.16728 (cross-list from cs.CV) [pdf, other]
Title: Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
Lijun Yu
Comments: PhD thesis
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)
[82] arXiv:2405.16807 (cross-list from cs.CV) [pdf, other]
Title: Extreme Compression of Adaptive Neural Images
Leo Hoshikawa, Marcos V. Conde, Takeshi Ohashi, Atsushi Irie
Comments: Technical Report. Work in progress
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR); Multimedia (cs.MM)
[83] arXiv:2405.16961 (cross-list from eess.IV) [pdf, other]
Title: Blind Data Adaptation to tackle Covariate Shift in Operational Steganalysis
Rony Abecidan (CRIStAL), Vincent Itier (IMT Nord Europe, CRIStAL), Jérémie Boulanger (CRIStAL), Patrick Bas (CRIStAL), Tomáš Pevný (CTU)
Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Multimedia (cs.MM)
[84] arXiv:2405.17729 (cross-list from cs.CV) [pdf, html, other]
Title: Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
Rui Zhang, Shuailong Li, Junxiao Xue, Feng Lin, Qing Zhang, Xiao Ma, Xiaoran Yan
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[85] arXiv:2405.17730 (cross-list from cs.CV) [pdf, html, other]
Title: MMPareto: Boosting Multimodal Learning with Innocent Unimodal Assistance
Yake Wei, Di Hu
Comments: Accepted by ICML2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)
[86] arXiv:2405.17842 (cross-list from cs.CV) [pdf, html, other]
Title: MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
Akio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji
Comments: ICLR 2025
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[87] arXiv:2405.18386 (cross-list from cs.SD) [pdf, html, other]
Title: Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
Yixiao Zhang, Yukara Ikemiya, Woosung Choi, Naoki Murata, Marco A. Martínez-Ramírez, Liwei Lin, Gus Xia, Wei-Hsiang Liao, Yuki Mitsufuji, Simon Dixon
Comments: Code and demo are available at: this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[88] arXiv:2405.18726 (cross-list from cs.SD) [pdf, html, other]
Title: Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
Che Liu, Changde Du, Xiaoyu Chen, Huiguang He
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[89] arXiv:2405.18790 (cross-list from cs.CV) [pdf, html, other]
Title: Opinion-Unaware Blind Image Quality Assessment using Multi-Scale Deep Feature Statistics
Zhangkai Ni, Yue Liu, Keyan Ding, Wenhan Yang, Hanli Wang, Shiqi Wang
Comments: Accepted to IEEE Transactions on Multimedia 2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Image and Video Processing (eess.IV)
[90] arXiv:2405.18887 (cross-list from cs.HC) [pdf, html, other]
Title: 4Doodle: Two-handed Gestures for Immersive Sketching of Architectural Models
Fernando Fonseca, Maurício Sousa, Daniel Mendes, Alfredo Ferreira, Joaquim Jorge
Comments: 9 pages; 15 Figures
Subjects: Human-Computer Interaction (cs.HC); Computational Engineering, Finance, and Science (cs.CE); Graphics (cs.GR); Multimedia (cs.MM)
[91] arXiv:2405.18959 (cross-list from cs.CV) [pdf, html, other]
Title: Transcending Fusion: A Multi-Scale Alignment Method for Remote Sensing Image-Text Retrieval
Rui Yang, Shuang Wang, Yingping Han, Yuanheng Li, Dong Zhao, Dou Quan, Yanhe Guo, Licheng Jiao
Comments: 16 pages, 9 figures
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[92] arXiv:2405.18991 (cross-list from cs.CV) [pdf, html, other]
Title: EasyAnimate: A High-Performance Long Video Generation Method based on Transformer Architecture
Jiaqi Xu, Xinyi Zou, Kunzhe Huang, Yunkuo Chen, Bo Liu, MengLi Cheng, Xing Shi, Jun Huang
Comments: 8 pages, 6 figures
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Multimedia (cs.MM)
[93] arXiv:2405.19226 (cross-list from cs.CV) [pdf, html, other]
Title: ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions
Honglin Lin, Siyu Li, Guoshun Nan, Chaoyue Tang, Xueting Wang, Jingxin Xu, Rong Yankai, Zhili Zhou, Yutong Gao, Qimei Cui, Xiaofeng Tao
Comments: Accepted in ACL 2024 Findings
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[94] arXiv:2405.19334 (cross-list from cs.AI) [pdf, html, other]
Title: LLMs Meet Multimodal Generation and Editing: A Survey
Yingqing He, Zhaoyang Liu, Jingye Chen, Zeyue Tian, Hongyu Liu, Xiaowei Chi, Runtao Liu, Ruibin Yuan, Yazhou Xing, Wenhai Wang, Jifeng Dai, Yong Zhang, Wei Xue, Qifeng Liu, Yike Guo, Qifeng Chen
Comments: 52 Pages with 16 Figures, 12 Tables, and 545 References. GitHub Repository at: this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)
[95] arXiv:2405.19889 (cross-list from eess.SP) [pdf, html, other]
Title: Deep Joint Semantic Coding and Beamforming for Near-Space Airship-Borne Massive MIMO Network
Minghui Wu, Zhen Gao, Zhaocheng Wang, Dusit Niyato, George K. Karagiannidis, Sheng Chen
Comments: Major Revision by IEEE JSAC
Subjects: Signal Processing (eess.SP); Information Theory (cs.IT); Machine Learning (cs.LG); Multimedia (cs.MM)
[96] arXiv:2405.20032 (cross-list from cs.NI) [pdf, html, other]
Title: Promptus: Can Prompts Streaming Replace Video Streaming with Stable Diffusion
Jiangkai Wu, Liming Liu, Yunpeng Tan, Junlin Hao, Xinggong Zhang
Subjects: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
[97] arXiv:2405.20606 (cross-list from cs.CV) [pdf, html, other]
Title: Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
Yang Chen, Tian He, Junfeng Fu, Ling Wang, Jingcai Guo, Ting Hu, Hong Cheng
Comments: Accepted by IEEE Transactions on Multimedia
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)
[98] arXiv:2405.20675 (cross-list from cs.CV) [pdf, html, other]
Title: Adv-KD: Adversarial Knowledge Distillation for Faster Diffusion Sampling
Kidist Amde Mekonnen, Nicola Dall'Asen, Paolo Rota
Comments: 7 pages, 11 figures, ELLIS Doctoral Symposium 2023 in Helsinki, Finland
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)
[99] arXiv:2405.20687 (cross-list from cs.CV) [pdf, html, other]
Title: Conditioning GAN Without Training Dataset
Kidist Amde Mekonnen
Comments: 5 pages, 2 figures, Part of my MSc project course, School Project Course 2022
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM)
[100] arXiv:2405.20775 (cross-list from cs.CR) [pdf, html, other]
Title: Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models
Xijie Huang, Xinyuan Wang, Hantao Zhang, Yinghao Zhu, Jiawen Xi, Jingkun An, Hao Wang, Hao Liang, Chengwei Pan
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)
Total of 100 entries
Showing up to 2000 entries per page: fewer | more | all
  • About
  • Help
  • contact arXivClick here to contact arXiv Contact
  • subscribe to arXiv mailingsClick here to subscribe Subscribe
  • Copyright
  • Privacy Policy
  • Web Accessibility Assistance
  • arXiv Operational Status
    Get status notifications via email or slack