Computer Vision and Pattern Recognition

Authors and titles for June 2025

Total of 3130 entries : 1-2000 2001-3130 2701-3130

Showing up to 2000 entries per page: fewer | more | all

[2701] arXiv:2506.10028 (cross-list from cs.CR) [pdf, other]: Title: Secure Data Access in Cloud Environments Using Quantum Cryptography

S. Vasavi Venkata Lakshmi, Ziaul Haque Choudhury

Subjects: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2702] arXiv:2506.10054 (cross-list from cs.LG) [pdf, html, other]: Title: Omni-DPO: A Dual-Perspective Paradigm for Dynamic Preference Learning of LLMs

Shangpin Peng, Weinong Wang, Zhuotao Tian, Senqiao Yang, Xing Wu, Haotian Xu, Chengquan Zhang, Takashi Isobe, Baotian Hu, Min Zhang

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2703] arXiv:2506.10142 (cross-list from eess.IV) [pdf, html, other]: Title: Rethinking Brain Tumor Segmentation from the Frequency Domain Perspective

Minye Shao, Zeyu Wang, Haoran Duan, Yawen Huang, Bing Zhai, Shizheng Wang, Yang Long, Yefeng Zheng

Comments: Accepted by IEEE Transactions on Medical Imaging

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2704] arXiv:2506.10146 (cross-list from cs.LG) [pdf, html, other]: Title: Balanced Hyperbolic Embeddings Are Natural Out-of-Distribution Detectors

Tejaswi Kasarla, Max van Spengler, Pascal Mettes

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2705] arXiv:2506.10172 (cross-list from cs.RO) [pdf, html, other]: Title: A Navigation Framework Utilizing Vision-Language Models

Yicheng Duan, Kaiyu tang

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2706] arXiv:2506.10177 (cross-list from cs.LG) [pdf, html, other]: Title: Geometric Regularity in Deterministic Sampling of Diffusion-based Generative Models

Defang Chen, Zhenyu Zhou, Can Wang, Siwei Lyu

Comments: 50 pages. The short version appeared in ICML 2024. arXiv admin note: substantial text overlap with arXiv:2405.11326

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
[2707] arXiv:2506.10230 (cross-list from eess.IV) [pdf, html, other]: Title: Prompt-Guided Latent Diffusion with Predictive Class Conditioning for 3D Prostate MRI Generation

Emerson P. Grabke, Masoom A. Haider, Babak Taati

Comments: MAH and BT are co-senior authors on the work. This work has been submitted to the IEEE for possible publication

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2708] arXiv:2506.10233 (cross-list from eess.IV) [pdf, html, other]: Title: Conditional diffusion models for guided anomaly detection in brain images using fluid-driven anomaly randomization

Ana Lawry Aguila, Peirong Liu, Oula Puonti, Juan Eugenio Iglesias

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2709] arXiv:2506.10251 (cross-list from eess.SY) [pdf, other]: Title: Energy Aware Camera Location Search Algorithm for Increasing Precision of Observation in Automated Manufacturing

Rongfei Li, Francis Assadian

Comments: 35 pages, 24 figures, Journal, Published in: Applied Sciences, 2024, vol. 14, article 9140. For published version, see this http URL: this https URL

Journal-ref: Appl. Sci. 2024, 14, 9140

Subjects: Systems and Control (eess.SY); Computer Vision and Pattern Recognition (cs.CV)
[2710] arXiv:2506.10265 (cross-list from eess.SP) [pdf, html, other]: Title: Ground Reaction Force Estimation via Time-aware Knowledge Distillation

Eun Som Jeon, Sinjini Mitra, Jisoo Lee, Omik M. Save, Ankita Shukla, Hyunglae Lee, Pavan Turaga

Journal-ref: IEEE Internet of Things Journal, 2025

Subjects: Signal Processing (eess.SP); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
[2711] arXiv:2506.10309 (cross-list from eess.IV) [pdf, html, other]: Title: DUN-SRE: Deep Unrolling Network with Spatiotemporal Rotation Equivariance for Dynamic MRI Reconstruction

Yuliang Zhu, Jing Cheng, Qi Xie, Zhuo-Xu Cui, Qingyong Zhu, Yuanyuan Liu, Xin Liu, Jianfeng Ren, Chengbo Wang, Dong Liang

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2712] arXiv:2506.10325 (cross-list from eess.IV) [pdf, html, other]: Title: SWDL: Stratum-Wise Difference Learning with Deep Laplacian Pyramid for Semi-Supervised 3D Intracranial Hemorrhage Segmentation

Cheng Wang, Siqi Chen, Donghua Mi, Yang Chen, Yudong Zhang, Yinsheng Li

Comments: 11 pages, 4 figures, 6 Tables

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2713] arXiv:2506.10407 (cross-list from eess.SY) [pdf, html, other]: Title: Semi-Tensor-Product Based Convolutional Neural Networks

Daizhan Cheng

Subjects: Systems and Control (eess.SY); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2714] arXiv:2506.10415 (cross-list from cs.CL) [pdf, html, other]: Title: Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?

Yingjin Song, Yupei Du, Denis Paperno, Albert Gatt

Comments: 27 pages, 14 figures. Accepted to ACL 2025

Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2715] arXiv:2506.10468 (cross-list from cs.GR) [pdf, html, other]: Title: Low-Barrier Dataset Collection with Real Human Body for Interactive Per-Garment Virtual Try-On

Zaiqiang Wu, Yechen Li, Jingyuan Liu, Yuki Shibata, Takayuki Hori, I-Chao Shen, Takeo Igarashi

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2716] arXiv:2506.10507 (cross-list from cs.GR) [pdf, html, other]: Title: Edit360: 2D Image Edits to 3D Assets from Any Angle

Junchao Huang, Xinting Hu, Shaoshuai Shi, Zhuotao Tian, Li Jiang

Comments: 11 pages, 9 figures

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2717] arXiv:2506.10540 (cross-list from cs.MA) [pdf, html, other]: Title: AniMaker: Automated Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation

Haoyuan Shi, Yunxin Li, Xinyu Chen, Longyue Wang, Baotian Hu, Min Zhang

Subjects: Multiagent Systems (cs.MA); Computer Vision and Pattern Recognition (cs.CV)
[2718] arXiv:2506.10580 (cross-list from cs.GR) [pdf, html, other]: Title: Transformer IMU Calibrator: Dynamic On-body IMU Calibration for Inertial Motion Capture

Chengxu Zuo, Jiawei Huang, Xiao Jiang, Yuan Yao, Xiangren Shi, Rui Cao, Xinyu Yi, Feng Xu, Shihui Guo, Yipeng Qin

Comments: Accepted by SIGGRAPH 2025 (TOG)

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2719] arXiv:2506.10600 (cross-list from cs.RO) [pdf, html, other]: Title: EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence

Xinjie Wang, Liu Liu, Yu Cao, Ruiqi Wu, Wenkang Qin, Dehui Wang, Wei Sui, Zhizhong Su

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2720] arXiv:2506.10617 (cross-list from cs.LG) [pdf, html, other]: Title: Deep Learning-Based Digitization of Overlapping ECG Images with Open-Source Python Code

Reza Karbasi, Masoud Rahimi, Abdol-Hossein Vahabie, Hadi Moradi

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2721] arXiv:2506.10632 (cross-list from cs.LG) [pdf, html, other]: Title: Hessian Geometry of Latent Space in Generative Models

Alexander Lobashev, Dmitry Guskov, Maria Larchenko, Mikhail Tamm

Comments: ICML 2025

Subjects: Machine Learning (cs.LG); Statistical Mechanics (cond-mat.stat-mech); Computer Vision and Pattern Recognition (cs.CV); Differential Geometry (math.DG); Statistics Theory (math.ST)
[2722] arXiv:2506.10675 (cross-list from eess.IV) [pdf, html, other]: Title: ConStyX: Content Style Augmentation for Generalizable Medical Image Segmentation

Xi Chen, Zhiqiang Shen, Peng Cao, Jinzhu Yang, Osmar R. Zaiane

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2723] arXiv:2506.10797 (cross-list from physics.med-ph) [pdf, other]: Title: Modality-AGnostic Image Cascade (MAGIC) for Multi-Modality Cardiac Substructure Segmentation

Nicholas Summerfield, Qisheng He, Alex Kuo, Ahmed I. Ghanem, Simeng Zhu, Chase Ruff, Joshua Pan, Anudeep Kumar, Prashant Nagpal, Jiwei Zhao, Ming Dong, Carri K. Glide-Hurst

Subjects: Medical Physics (physics.med-ph); Computer Vision and Pattern Recognition (cs.CV)
[2724] arXiv:2506.10825 (cross-list from eess.IV) [pdf, other]: Title: Generalist Models in Medical Image Segmentation: A Survey and Performance Comparison with Task-Specific Approaches

Andrea Moglia (1), Matteo Leccardi (1), Matteo Cavicchioli (1), Alice Maccarini (2), Marco Marcon (1), Luca Mainardi (1), Pietro Cerveri (1 and 2) ((1) Politecnico di Milano, (2) Università di Pavia)

Comments: 132 pages, 26 figures, 23 tables. Andrea Moglia and Matteo Leccardi are equally contributing authors

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2725] arXiv:2506.10858 (cross-list from eess.IV) [pdf, html, other]: Title: Med-URWKV: Pure RWKV With ImageNet Pre-training For Medical Image Segmentation

Zhenhuan Zhou

Comments: Preprint Draft, 5 pages. This paper will be updated with a formal version in the future, Copyright: College of Computer Science, Nankai University. All rights reserved

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2726] arXiv:2506.10916 (cross-list from eess.IV) [pdf, other]: Title: Semi-Automated Quality Assurance in Digital Pathology: Tile Classification Approach

Meredith VandeHaar, M. Clinch, I. Yilmaz, M.A. Rahman, Y. Xiao, F. Dogany, H.M. Alazab, A. Nassar, Z. Akkus, B. Dangott

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2727] arXiv:2506.10955 (cross-list from cs.LG) [pdf, html, other]: Title: ReGuidance: A Simple Diffusion Wrapper for Boosting Sample Quality on Hard Inverse Problems

Aayush Karan, Kulin Shah, Sitan Chen

Comments: 38 pages, 14 figures

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2728] arXiv:2506.10968 (cross-list from cs.RO) [pdf, other]: Title: Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop

Justin Kerr, Kush Hari, Ethan Weber, Chung Min Kim, Brent Yi, Tyler Bonnen, Ken Goldberg, Angjoo Kanazawa

Comments: Project page: this https URL

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2729] arXiv:2506.11004 (cross-list from cs.LG) [pdf, html, other]: Title: Developing a Dyslexia Indicator Using Eye Tracking

Kevin Cogan, Vuong M. Ngo, Mark Roantree

Comments: The 23rd International Conference on Artificial Intelligence in Medicine (AIME 2025), LNAI, Springer, 11 pages

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
[2730] arXiv:2506.11025 (cross-list from cs.LG) [pdf, html, other]: Title: When Algorithms Play Favorites: Lookism in the Generation and Perception of Faces

Miriam Doh, Aditya Gulati, Matei Mancas, Nuria Oliver

Comments: Accepted as an extended abstract at the Fourth European Workshop on Algorithmic Fairness (EWAF) (URL: this https URL)

Journal-ref: Proceedings of Machine Learning Research 294 (2025) 474-480

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2731] arXiv:2506.11035 (cross-list from cs.LG) [pdf, html, other]: Title: Tversky Neural Networks: Psychologically Plausible Deep Learning with Differentiable Tversky Similarity

Moussa Koulako Bala Doumbouya, Dan Jurafsky, Christopher D. Manning

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2732] arXiv:2506.11073 (cross-list from cs.CL) [pdf, html, other]: Title: CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention

Zekai Ye, Qiming Li, Xiaocheng Feng, Libo Qin, Yichong Huang, Baohang Li, Kui Jiang, Yang Xiang, Zhirui Zhang, Yunfei Lu, Duyu Tang, Dandan Tu, Bing Qin

Comments: ACL2025 Main

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2733] arXiv:2506.11123 (cross-list from q-bio.NC) [pdf, html, other]: Title: Sparse Autoencoders Bridge The Deep Learning Model and The Brain

Ziming Mao, Jia Xu, Zeqi Zheng, Haofang Zheng, Dabing Sheng, Yaochu Jin, Guoyuan Yang

Comments: 54 pages, 41 figures

Subjects: Neurons and Cognition (q-bio.NC); Computer Vision and Pattern Recognition (cs.CV)
[2734] arXiv:2506.11139 (cross-list from eess.IV) [pdf, html, other]: Title: Grids Often Outperform Implicit Neural Representations

Namhoon Kim, Sara Fridovich-Keil

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2735] arXiv:2506.11146 (cross-list from quant-ph) [pdf, html, other]: Title: HQFNN: A Compact Quantum-Fuzzy Neural Network for Accurate Image Classification

Jianhong Yao, Yangming Guo

Subjects: Quantum Physics (quant-ph); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2736] arXiv:2506.11150 (cross-list from eess.IV) [pdf, html, other]: Title: ADAgent: LLM Agent for Alzheimer's Disease Analysis with Collaborative Coordinator

Wenlong Hou, Guangqian Yang, Ye Du, Yeung Lau, Lihao Liu, Junjun He, Ling Long, Shujun Wang

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2737] arXiv:2506.11163 (cross-list from eess.IV) [pdf, html, other]: Title: Vector Representations of Vessel Trees

James Batten, Michiel Schaap, Matthew Sinclair, Ying Bai, Ben Glocker

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
[2738] arXiv:2506.11183 (cross-list from eess.IV) [pdf, html, other]: Title: DiffPR: Diffusion-Based Phase Reconstruction via Frequency-Decoupled Learning

Yi Zhang

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2739] arXiv:2506.11234 (cross-list from cs.RO) [pdf, other]: Title: Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving

Luke Rowe, Rodrigue de Schaetzen, Roger Girgis, Christopher Pal, Liam Paull

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2740] arXiv:2506.11252 (cross-list from cs.GR) [pdf, other]: Title: Anti-Aliased 2D Gaussian Splatting

Mae Younes, Adnane Boukhayma

Comments: Code will be available at this https URL

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2741] arXiv:2506.11261 (cross-list from cs.RO) [pdf, other]: Title: Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation

Shizhe Chen, Ricardo Garcia, Paul Pacaud, Cordelia Schmid

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2742] arXiv:2506.11283 (cross-list from eess.IV) [pdf, html, other]: Title: Joint Denoising of Cryo-EM Projection Images using Polar Transformers

Joakim Andén, Justus Sagemüller

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2743] arXiv:2506.11387 (cross-list from cs.RO) [pdf, other]: Title: Control Architecture and Design for a Multi-robotic Visual Servoing System in Automated Manufacturing Environment

Rongfei Li

Comments: 272 pages, 171 figures, PhD dissertation, University of California, Davis, 2025. To be published in ProQuest ETD

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Systems and Control (eess.SY)
[2744] arXiv:2506.11444 (cross-list from cs.CR) [pdf, html, other]: Title: GaussMarker: Robust Dual-Domain Watermark for Diffusion Models

Kecen Li, Zhicong Huang, Xinwen Hou, Cheng Hong

Comments: Accepted at ICML 2025

Subjects: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2745] arXiv:2506.11454 (cross-list from eess.IV) [pdf, other]: Title: FAD-Net: Frequency-Domain Attention-Guided Diffusion Network for Coronary Artery Segmentation using Invasive Coronary Angiography

Nan Mu, Ruiqi Song, Xiaoning Li, Zhihui Xu, Jingfeng Jiang, Chen Zhao

Comments: 35 pages, 12 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2746] arXiv:2506.11455 (cross-list from q-bio.NC) [pdf, other]: Title: Voxel-Level Brain States Prediction Using Swin Transformer

Yifei Sun, Daniel Chahine, Qinghao Wen, Tianming Liu, Xiang Li, Yixuan Yuan, Fernando Calamante, Jinglei Lv

Subjects: Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2747] arXiv:2506.11465 (cross-list from cs.LG) [pdf, html, other]: Title: RollingQ: Reviving the Cooperation Dynamics in Multimodal Transformer

Haotian Ni, Yake Wei, Hang Liu, Gong Chen, Chong Peng, Hao Lin, Di Hu

Comments: Accepted by ICML 2025

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2748] arXiv:2506.11475 (cross-list from cs.MA) [pdf, other]: Title: AutoGen Driven Multi Agent Framework for Iterative Crime Data Analysis and Prediction

Syeda Kisaa Fatima, Tehreem Zubair, Noman Ahmed, Asifullah Khan

Subjects: Multiagent Systems (cs.MA); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2749] arXiv:2506.11496 (cross-list from eess.IV) [pdf, html, other]: Title: Taming Stable Diffusion for Computed Tomography Blind Super-Resolution

Chunlei Li, Yilei Shi, Haoxi Hu, Jingliang Hu, Xiao Xiang Zhu, Lichao Mou

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2750] arXiv:2506.11545 (cross-list from eess.IV) [pdf, html, other]: Title: FCA2: Frame Compression-Aware Autoencoder for Modular and Fast Compressed Video Super-Resolution

Zhaoyang Wang, Jie Li, Wen Lu, Lihuo He, Maoguo Gong, Xinbo Gao

Comments: This work has been submitted to the IEEE TMM for possible publication

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2751] arXiv:2506.11546 (cross-list from cs.GR) [pdf, html, other]: Title: CGVQM+D: Computer Graphics Video Quality Metric and Dataset

Akshay Jindal, Nabil Sadaka, Manu Mathew Thomas, Anton Sochenov, Anton Kaplanyan

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2752] arXiv:2506.11604 (cross-list from cs.AI) [pdf, other]: Title: VLM@school -- Evaluation of AI image understanding on German middle school knowledge

René Peinl, Vincent Tischler

Comments: Peinl, René; Tischler, Vincent (2025): VLM@school - Evaluation of AI image understanding on German middle school knowledge. Future Technologies Conference (FTC) 2025, Munich, Germany 2025 (accepted)

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2753] arXiv:2506.11671 (cross-list from eess.IV) [pdf, html, other]: Title: Brain Network Analysis Based on Fine-tuned Self-supervised Model for Brain Disease Diagnosis

Yifei Tang, Hongjie Jiang, Changhong Jing, Hieu Pham, Shuqiang Wang

Comments: 13 pages, 3 figures, International Conference on Neural Computing for Advanced Applications

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2754] arXiv:2506.11753 (cross-list from eess.IV) [pdf, other]: Title: Exploring the Effectiveness of Deep Features from Domain-Specific Foundation Models in Retinal Image Synthesis

Zuzanna Skorniewska, Bartlomiej W. Papiez

Comments: To be published and presented at the MIUA 2025 conference

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2755] arXiv:2506.11796 (cross-list from nlin.AO) [pdf, html, other]: Title: Solving Inverse Problems in Stochastic Self-Organising Systems through Invariant Representations

Elias Najarro, Nicolas Bessone, Sebastian Risi

Comments: Preprint. Under review

Subjects: Adaptation and Self-Organizing Systems (nlin.AO); Disordered Systems and Neural Networks (cond-mat.dis-nn); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2756] arXiv:2506.11821 (cross-list from eess.IV) [pdf, html, other]: Title: Framework of a multiscale data-driven DT of the musculoskeletal system

Martina Paccini, Simone Cammarasana, Giuseppe Patanè

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2757] arXiv:2506.11823 (cross-list from eess.IV) [pdf, html, other]: Title: Structural Similarity-Inspired Unfolding for Lightweight Image Super-Resolution

Zhangkai Ni, Yang Zhang, Wenhan Yang, Hanli Wang, Shiqi Wang, Sam Kwong

Comments: Accepted to IEEE Transactions on Image Processing

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2758] arXiv:2506.11860 (cross-list from eess.IV) [pdf, other]: Title: MindGrab for BrainChop: Fast and Accurate Skull Stripping for Command Line and Browser

Armina Fani (1), Mike Doan (1), Isabelle Le (1), Alex Fedorov (2), Malte Hoffmann (3), Chris Rorden (4), Sergey Plis (1) ((1) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, Emory University, (2) Emory University, (3) Harvard University, (4) University of South Carolina)

Comments: 12 pages, 1 table, 4 figures. 2 supplementary tables, 1 supplementary figure. Brainchop-cli: this https URL . Brainchop web: this https URL

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE)
[2759] arXiv:2506.11925 (cross-list from cs.AR) [pdf, html, other]: Title: Real-World Deployment of a Lane Change Prediction Architecture Based on Knowledge Graph Embeddings and Bayesian Inference

M. Manzour, Catherine M. Elias, Omar M. Shehata, R. Izquierdo, M. A. Sotelo

Subjects: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2760] arXiv:2506.11967 (cross-list from cs.LG) [pdf, html, other]: Title: Visual Pre-Training on Unlabeled Images using Reinforcement Learning

Dibya Ghosh, Sergey Levine

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2761] arXiv:2506.12006 (cross-list from eess.IV) [pdf, html, other]: Title: crossMoDA Challenge: Evolution of Cross-Modality Domain Adaptation Techniques for Vestibular Schwannoma and Cochlea Segmentation from 2021 to 2023

Navodini Wijethilake, Reuben Dorent, Marina Ivory, Aaron Kujawa, Stefan Cornelissen, Patrick Langenhuizen, Mohamed Okasha, Anna Oviedova, Hexin Dong, Bogyeong Kang, Guillaume Sallé, Luyi Han, Ziyuan Zhao, Han Liu, Yubo Fan, Tao Yang, Shahad Hardan, Hussain Alasmawi, Santosh Sanjeev, Yuzhou Zhuang, Satoshi Kondo, Maria Baldeon Calisto, Shaikh Muhammad Uzair Noman, Cancan Chen, Ipek Oguz, Rongguo Zhang, Mina Rezaei, Susana K. Lai-Yuen, Satoshi Kasai, Yunzhi Huang, Chih-Cheng Hung, Mohammad Yaqub, Lisheng Wang, Benoit M. Dawant, Cuntai Guan, Ritse Mann, Vincent Jaouen, Tae-Eui Kam, Li Zhang, Jonathan Shapey, Tom Vercauteren

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2762] arXiv:2506.12007 (cross-list from cs.LG) [pdf, html, other]: Title: SIMSHIFT: A Benchmark for Adapting Neural Surrogates to Distribution Shifts

Paul Setinek, Gianluca Galletti, Thomas Gross, Dominik Schnürer, Johannes Brandstetter, Werner Zellinger

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Computational Physics (physics.comp-ph)
[2763] arXiv:2506.12015 (cross-list from cs.LG) [pdf, html, other]: Title: EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction

Hsi-Che Lin, Yu-Chu Yu, Kai-Po Chang, Yu-Chiang Frank Wang

Comments: Under review. Project page: this https URL

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2764] arXiv:2506.12040 (cross-list from cs.LG) [pdf, html, other]: Title: BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook

Hao Gu, Lujun Li, Zheyu Wang, Bei Liu, Qiyuan Zhu, Sirui Han, Yike Guo

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2765] arXiv:2506.12041 (cross-list from cs.LG) [pdf, html, other]: Title: Meta Pruning via Graph Metanetworks : A Meta Learning Framework for Network Pruning

Yewei Liu, Xiyuan Wang, Muhan Zhang

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2766] arXiv:2506.12106 (cross-list from eess.IV) [pdf, html, other]: Title: Enhancing Privacy: The Utility of Stand-Alone Synthetic CT and MRI for Tumor and Bone Segmentation

André Ferreira, Kunpeng Xie, Caroline Wilpert, Gustavo Correia, Felix Barajas Ordonez, Tiago Gil Oliveira, Maike Bode, Robert Siepmann, Frank Hölzle, Rainer Röhrig, Jens Kleesiek, Daniel Truhn, Jan Egger, Victor Alves, Behrus Puladi

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2767] arXiv:2506.12116 (cross-list from cs.CL) [pdf, html, other]: Title: Unsupervised Document and Template Clustering using Multimodal Embeddings

Phillipe R. Sampaio, Helene Maxcici

Comments: 22 pages, 12 figures

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2768] arXiv:2506.12156 (cross-list from cs.LG) [pdf, html, other]: Title: Explaining Recovery Trajectories of Older Adults Post Lower-Limb Fracture Using Modality-wise Multiview Clustering and Large Language Models

Shehroz S. Khan, Ali Abedi, Charlene H. Chu

Comments: 15 pages, 2 figures, 3 tables

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2769] arXiv:2506.12184 (cross-list from cs.RO) [pdf, other]: Title: SPLATART: Articulated Gaussian Splatting with Estimated Object Structure

Stanley Lewis, Vishal Chandra, Tom Gao, Odest Chadwicke Jenkins

Comments: 7 pages, Accepted to the 2025 RSS Workshop on Gaussian Representations for Robot Autonomy. Contact: Stanley Lewis, [email protected]

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2770] arXiv:2506.12186 (cross-list from eess.IV) [pdf, html, other]: Title: MRI-CORE: A Foundation Model for Magnetic Resonance Imaging

Haoyu Dong, Yuwen Chen, Hanxue Gu, Nicholas Konz, Yaqian Chen, Qihang Li, Maciej A. Mazurowski

Comments: 36 pages, under review

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2771] arXiv:2506.12239 (cross-list from cs.RO) [pdf, html, other]: Title: ViTaSCOPE: Visuo-tactile Implicit Representation for In-hand Pose and Extrinsic Contact Estimation

Jayjun Lee, Nima Fazeli

Comments: Accepted to RSS 2025 | Project page: this https URL

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2772] arXiv:2506.12269 (cross-list from eess.IV) [pdf, html, other]: Title: ICME 2025 Grand Challenge on Video Super-Resolution for Video Conferencing

Babak Naderi, Ross Cutler, Juhee Cho, Nabakumar Khongbantabam, Dejan Ivkovic

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[2773] arXiv:2506.12344 (cross-list from cs.CR) [pdf, html, other]: Title: Restoring Gaussian Blurred Face Images for Deanonymization Attacks

Haoyu Zhai, Shuo Wang, Pirouz Naghavi, Qingying Hao, Gang Wang

Comments: 18 pages, 16 figures, IEEE Transaction format

Subjects: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2774] arXiv:2506.12348 (cross-list from cs.GR) [pdf, html, other]: Title: Real-Time Per-Garment Virtual Try-On with Temporal Consistency for Loose-Fitting Garments

Zaiqiang Wu, I-Chao Shen, Takeo Igarashi

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2775] arXiv:2506.12364 (cross-list from cs.AI) [pdf, html, other]: Title: MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval

Mingjun Xu, Jinhan Dong, Jue Hou, Zehui Wang, Sihang Li, Zhifeng Gao, Renxin Zhong, Hengxing Cai

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2776] arXiv:2506.12375 (cross-list from cs.NE) [pdf, html, other]: Title: Optimized Spectral Fault Receptive Fields for Diagnosis-Informed Prognosis

Stan Muñoz Gutiérrez, Franz Wotawa

Comments: Submitted to The 36th International Conference on Principles of Diagnosis and Resilient Systems (DX'25)

Subjects: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2777] arXiv:2506.12395 (cross-list from eess.IV) [pdf, html, other]: Title: Shape-aware Sampling Matters in the Modeling of Multi-Class Tubular Structures

Minghui Zhang, Yaoyu Liu, Xin You, Hanxiao Zhang, Yun Gu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2778] arXiv:2506.12411 (cross-list from cs.CR) [pdf, html, other]: Title: InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning

Mengyuan Sun, Yu Li, Yuchen Liu, Bo Du, Yunjie Ge

Subjects: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2779] arXiv:2506.12430 (cross-list from cs.CR) [pdf, html, other]: Title: Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025

Zonghao Ying, Siyang Wu, Run Hao, Peng Ying, Shixuan Sun, Pengyu Chen, Junze Chen, Hao Du, Kaiwen Shen, Shangkun Wu, Jiwei Wei, Shiyuan He, Yang Yang, Xiaohai Xu, Ke Ma, Qianqian Xu, Qingming Huang, Shi Lin, Xun Wang, Changting Lin, Meng Han, Yilei Jiang, Siqi Lai, Yaozhi Zheng, Yifei Song, Xiangyu Yue, Zonglei Jing, Tianyuan Zhang, Zhilei Zhu, Aishan Liu, Jiakai Wang, Siyuan Liang, Xianglong Kong, Hainan Li, Junjie Mu, Haotong Qin, Yue Yu, Lei Chen, Felix Juefei-Xu, Qing Guo, Xinyun Chen, Yew Soon Ong, Xianglong Liu, Dawn Song, Alan Yuille, Philip Torr, Dacheng Tao

Comments: AdvML@CVPR Challenge Report

Subjects: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2780] arXiv:2506.12440 (cross-list from cs.SD) [pdf, html, other]: Title: Style-based Composer Identification and Attribution of Symbolic Music Scores: a Systematic Survey

Federico Simonetta

Comments: Accepted at the TISMIR

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Digital Libraries (cs.DL); Audio and Speech Processing (eess.AS)
[2781] arXiv:2506.12471 (cross-list from eess.IV) [pdf, html, other]: Title: Adaptive Multi-resolution Hash-Encoding Framework for INR-based Dental CBCT Reconstruction with Truncated FOV

Hyoung Suk Park, Kiwan Jeon

Comments: 18 pages, 4 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2782] arXiv:2506.12475 (cross-list from eess.IV) [pdf, html, other]: Title: Efficient Star Distillation Attention Network for Lightweight Image Super-Resolution

Fangwei Hao, Ji Du, Desheng Kong, Jiesheng Wu, Jing Xu, Ping Li

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2783] arXiv:2506.12479 (cross-list from cs.AI) [pdf, html, other]: Title: AI Flow: Perspectives, Scenarios, and Approaches

Hongjun An, Wenhan Hu, Sida Huang, Siqi Huang, Ruanjun Li, Yuanzhi Liang, Jiawei Shao, Yiliang Song, Zihan Wang, Cheng Yuan, Chi Zhang, Hongyuan Zhang, Wenhao Zhuang, Xuelong Li

Comments: Authors are with Institute of Artificial Intelligence (TeleAI), China Telecom, China. Author names are listed alphabetically by surname. This work was conducted at TeleAI, facilitated by Dr. Jiawei Shao (e-mail: [email protected]) under the leadership of Prof. Xuelong Li. The corresponding author is Prof. Xuelong Li (e-mail: xuelong [email protected]), the CTO and Chief Scientist of China Telecom

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC); Signal Processing (eess.SP)
[2784] arXiv:2506.12541 (cross-list from cs.LG) [pdf, html, other]: Title: BSA: Ball Sparse Attention for Large-scale Geometries

Catalin E. Brita, Hieu Nguyen, Lohithsai Yadala Chanchu, Domonkos Nagy, Maksim Zhdanov

Comments: Long Context Foundation Models Workshop @ ICML 2025

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2785] arXiv:2506.12542 (cross-list from cs.LG) [pdf, html, other]: Title: PLD: A Choice-Theoretic List-Wise Knowledge Distillation

Ejafa Bassam, Dawei Zhu, Kaigui Bian

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
[2786] arXiv:2506.12678 (cross-list from cs.RO) [pdf, html, other]: Title: Adapting by Analogy: OOD Generalization of Visuomotor Policies via Functional Correspondence

Pranay Gupta, Henny Admoni, Andrea Bajcsy

Comments: 15 pages, 11 figures

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2787] arXiv:2506.12693 (cross-list from eess.IV) [pdf, html, other]: Title: Zero-shot denoising via neural compression: Theoretical and algorithmic framework

Ali Zafari, Xi Chen, Shirin Jalali

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Information Theory (cs.IT)
[2788] arXiv:2506.12719 (cross-list from eess.IV) [pdf, html, other]: Title: GM-LDM: Latent Diffusion Model for Brain Biomarker Identification through Functional Data-Driven Gray Matter Synthesis

Hu Xu, Yang Jingling, Jia Sihan, Bi Yuda, Calhoun Vince

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2789] arXiv:2506.12798 (cross-list from eess.IV) [pdf, html, other]: Title: Predicting Genetic Mutations from Single-Cell Bone Marrow Images in Acute Myeloid Leukemia Using Noise-Robust Deep Learning Models

Garima Jain, Ravi Kant Gupta, Priyansh Jain, Abhijeet Patil, Ardhendu Sekhar, Gajendra Smeeta, Sanghamitra Pati, Amit Sethi

Comments: 2 figues

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2790] arXiv:2506.12847 (cross-list from cs.GR) [pdf, html, other]: Title: iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Zhelun Shen, Chenming Wu, Junsheng Zhou, Chen Zhao, Kaisiyuan Wang, Hang Zhou, Yingying Li, Haocheng Feng, Wei He, Jingdong Wang

Comments: Technical report, 12 pages

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2791] arXiv:2506.13045 (cross-list from cs.LG) [pdf, html, other]: Title: Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

Haiyang Guo, Fanhu Zeng, Fei Zhu, Jiayi Wang, Xukai Wang, Jingang Zhou, Hongbo Zhao, Wenzhuo Liu, Shijie Ma, Da-Han Wang, Xu-Yao Zhang, Cheng-Lin Liu

Comments: Preprint

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2792] arXiv:2506.13050 (cross-list from cs.GR) [pdf, html, other]: Title: NeuVAS: Neural Implicit Surfaces for Variational Shape Modeling

Pengfei Wang, Qiujie Dong, Fangtian Liang, Hao Pan, Lei Yang, Congyi Zhang, Guying Lin, Caiming Zhang, Yuanfeng Zhou, Changhe Tu, Shiqing Xin, Alla Sheffer, Xin Li, Wenping Wang

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2793] arXiv:2506.13056 (cross-list from cs.AI) [pdf, other]: Title: Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning

Haibo Qiu, Xiaohan Lan, Fanfan Liu, Xiaohu Sun, Delian Ruan, Peng Shi, Lin Ma

Comments: Project Page: this https URL

Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2794] arXiv:2506.13100 (cross-list from cs.RO) [pdf, html, other]: Title: A Novel ViDAR Device With Visual Inertial Encoder Odometry and Reinforcement Learning-Based Active SLAM Method

Zhanhua Xin, Zhihao Wang, Shenghao Zhang, Wanchao Chi, Yan Meng, Shihan Kong, Yan Xiong, Chong Zhang, Yuzhen Liu, Junzhi Yu

Comments: 12 pages, 13 figures

Journal-ref: IEEE Transactions on Industrial Informatics, pp. 1-12, 2025

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2795] arXiv:2506.13160 (cross-list from cs.LG) [pdf, html, other]: Title: CertDW: Towards Certified Dataset Ownership Verification via Conformal Prediction

Ting Qiao, Yiming Li, Jianbin Li, Yingjia Wang, Leyi Qi, Junfeng Guo, Ruili Feng, Dacheng Tao

Comments: The first two authors contributed equally to this work. 16 pages

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2796] arXiv:2506.13187 (cross-list from cs.LG) [pdf, html, other]: Title: Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence

Yibo Yang, Sihao Liu, Chuan Rao, Bang An, Tiancheng Shen, Philip H.S. Torr, Ming-Hsuan Yang, Bernard Ghanem

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2797] arXiv:2506.13195 (cross-list from eess.IV) [pdf, html, other]: Title: ViT-NeBLa: A Hybrid Vision Transformer and Neural Beer-Lambert Framework for Single-View 3D Reconstruction of Oral Anatomy from Panoramic Radiographs

Bikram Keshari Parida, Anusree P. Sunilkumar, Abhijit Sen, Wonsang You

Comments: 10 figures, 19 pages

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2798] arXiv:2506.13277 (cross-list from cs.LG) [pdf, html, other]: Title: SeqPE: Transformer with Sequential Position Encoding

Huayang Li, Yahui Liu, Hongyu Sun, Deng Cai, Leyang Cui, Wei Bi, Peilin Zhao, Taro Watanabe

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2799] arXiv:2506.13306 (cross-list from eess.IV) [pdf, html, other]: Title: Brain Imaging Foundation Models, Are We There Yet? A Systematic Review of Foundation Models for Brain Imaging and Biomedical Research

Salah Ghamizi, Georgia Kanli, Yu Deng, Magali Perquin, Olivier Keunen

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2800] arXiv:2506.13348 (cross-list from cs.GR) [pdf, html, other]: Title: TextureSplat: Per-Primitive Texture Mapping for Reflective Gaussian Splatting

Mae Younes, Adnane Boukhayma

Comments: Code will be available at this https URL

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2801] arXiv:2506.13415 (cross-list from eess.IV) [pdf, html, other]: Title: Simple is what you need for efficient and accurate medical image segmentation

Xiang Yu, Yayan Chen, Guannan He, Qing Zeng, Yue Qin, Meiling Liang, Dandan Luo, Yimei Liao, Zeyu Ren, Cheng Kang, Delong Yang, Bocheng Liang, Bin Pu, Ying Yuan, Shengli Li

Comments: 15 pages, 11 figures

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2802] arXiv:2506.13419 (cross-list from eess.IV) [pdf, html, other]: Title: Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos

Riku Takahashi, Ryugo Morita, Jinjia Zhou

Comments: Accepted to ICMR2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2803] arXiv:2506.13425 (cross-list from cs.RO) [pdf, html, other]: Title: JENGA: Object selection and pose estimation for robotic grasping from a stack

Sai Srinivas Jeevanandam, Sandeep Inuganti, Shreedhar Govil, Didier Stricker, Jason Rambach

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2804] arXiv:2506.13443 (cross-list from eess.IV) [pdf, other]: Title: PRO: Projection Domain Synthesis for CT Imaging

Kang Chen, Bin Huang, Xuebin Yang, Junyan Zhang, Qiegen Liu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2805] arXiv:2506.13477 (cross-list from cs.HC) [pdf, html, other]: Title: Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications

Pegah Salehi, Sajad Amouei Sheshkal, Vajira Thambawita, Michael A. Riegler, Pål Halvorsen

Comments: 20 pages, 9 figures, 9 tables

Subjects: Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)
[2806] arXiv:2506.13579 (cross-list from cs.LG) [pdf, html, other]: Title: Flexible-length Text Infilling for Discrete Diffusion Models

Andrew Zhang, Anushka Sivakumar, Chiawei Tang, Chris Thomas

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2807] arXiv:2506.13614 (cross-list from stat.ML) [pdf, html, other]: Title: Exploiting the Exact Denoising Posterior Score in Training-Free Guidance of Diffusion Models

Gregory Bellchambers

Subjects: Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2808] arXiv:2506.13642 (cross-list from cs.AI) [pdf, html, other]: Title: Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou, Yang Feng

Comments: Code: this https URL , Model: this https URL

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[2809] arXiv:2506.13667 (cross-list from eess.IV) [pdf, html, other]: Title: MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model

Bi Yuda, Jia Sihan, Gao Yutong, Abrol Anees, Fu Zening, Calhoun Vince

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2810] arXiv:2506.13679 (cross-list from cs.RO) [pdf, html, other]: Title: ROSA: Harnessing Robot States for Vision-Language and Action Alignment

Yuqing Wen, Kefan Gu, Haoxuan Liu, Yucheng Zhao, Tiancai Wang, Haoqiang Fan, Xiaoyan Sun

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2811] arXiv:2506.13754 (cross-list from cs.LG) [pdf, html, other]: Title: VideoPDE: Unified Generative PDE Solving via Video Inpainting Diffusion Models

Edward Li, Zichen Wang, Jiahe Huang, Jeong Joon Park

Comments: Project page: this https URL

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2812] arXiv:2506.13762 (cross-list from cs.RO) [pdf, html, other]: Title: Touch begins where vision ends: Generalizable policies for contact-rich manipulation

Zifan Zhao, Siddhant Haldar, Jinda Cui, Lerrel Pinto, Raunaq Bhirangi

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2813] arXiv:2506.13763 (cross-list from cs.LG) [pdf, html, other]: Title: Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value

Yixian Xu, Shengjie Luo, Liwei Wang, Di He, Chang Liu

Comments: 29 pages, 8 figures, 3 tables. Preprint. Work in Progress

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
[2814] arXiv:2506.13807 (cross-list from eess.IV) [pdf, html, other]: Title: BraTS orchestrator : Democratizing and Disseminating state-of-the-art brain tumor image analysis

Florian Kofler, Marcel Rosier, Mehdi Astaraki, Ujjwal Baid, Hendrik Möller, Josef A. Buchner, Felix Steinbauer, Eva Oswald, Ezequiel de la Rosa, Ivan Ezhov, Constantin von See, Jan Kirschke, Anton Schmick, Sarthak Pati, Akis Linardos, Carla Pitarch, Sanyukta Adap, Jeffrey Rudie, Maria Correia de Verdier, Rachit Saluja, Evan Calabrese, Dominic LaBella, Mariam Aboian, Ahmed W. Moawad, Nazanin Maleki, Udunna Anazodo, Maruf Adewole, Marius George Linguraru, Anahita Fathi Kazerooni, Zhifan Jiang, Gian Marco Conte, Hongwei Li, Juan Eugenio Iglesias, Spyridon Bakas, Benedikt Wiestler, Marie Piraud, Bjoern Menze

Comments: 27p, 2figs, 3tabs

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2815] arXiv:2506.13819 (cross-list from eess.IV) [pdf, html, other]: Title: Reliable Noninvasive Glucose Sensing via CNN-Based Spectroscopy

El Arbi Belfarsi, Henry Flores, Maria Valero

Comments: Submitted to the IEEE-EMBS International Conference on Biomedical and Health Informatics (BHI 2025)

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2816] arXiv:2506.13888 (cross-list from cs.CL) [pdf, html, other]: Title: VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training

Jipeng Zhang, Kehao Miao, Renjie Pi, Zhaowei Wang, Runtao Liu, Rui Pan, Tong Zhang

Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2817] arXiv:2506.14107 (cross-list from cs.DC) [pdf, html, other]: Title: Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse

Jinwoo Hwang, Daeun Kim, Sangyeop Lee, Yoonsung Kim, Guseul Heo, Hojoon Kim, Yunseok Jeong, Tadiwos Meaza, Eunhyeok Park, Jeongseob Ahn, Jongse Park

Comments: Accepted to 2025 VLDB

Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Computer Vision and Pattern Recognition (cs.CV)
[2818] arXiv:2506.14135 (cross-list from cs.RO) [pdf, html, other]: Title: GAF: Gaussian Action Field as a Dynamic World Model for Robotic Manipulation

Ying Chai, Litao Deng, Ruizhi Shao, Jiajun Zhang, Liangjun Xing, Hongwen Zhang, Yebin Liu

Comments: this http URL

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2819] arXiv:2506.14198 (cross-list from cs.RO) [pdf, html, other]: Title: AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

Jeremy A. Collins, Loránd Cheng, Kunal Aneja, Albert Wilcox, Benjamin Joffe, Animesh Garg

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2820] arXiv:2506.14209 (cross-list from eess.IV) [pdf, html, other]: Title: Latent Anomaly Detection: Masked VQ-GAN for Unsupervised Segmentation in Medical CBCT

Pengwei Wang

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2821] arXiv:2506.14303 (cross-list from eess.IV) [pdf, html, other]: Title: orGAN: A Synthetic Data Augmentation Pipeline for Simultaneous Generation of Surgical Images and Ground Truth Labels

Niran Nataraj, Maina Sogabe, Kenji Kawashima

Comments: 24 pages, 7figures

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2822] arXiv:2506.14315 (cross-list from cs.GR) [pdf, html, other]: Title: ImmerseGen: Agent-Guided Immersive World Generation with Alpha-Textured Proxies

Jinyan Yuan, Bangbang Yang, Keke Wang, Panwang Pan, Lin Ma, Xuehai Zhang, Xiao Liu, Zhaopeng Cui, Yuewen Ma

Comments: Project webpage: this https URL

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2823] arXiv:2506.14318 (cross-list from eess.IV) [pdf, html, other]: Title: BRISC: Annotated Dataset for Brain Tumor Segmentation and Classification with Swin-HAFNet

Amirreza Fateh, Yasin Rezvani, Sara Moayedi, Sadjad Rezvani, Fatemeh Fateh, Mansoor Fateh, Vahid Abolghasemi

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2824] arXiv:2506.14381 (cross-list from eess.IV) [pdf, html, other]: Title: Compressed Video Super-Resolution based on Hierarchical Encoding

Yuxuan Jiang, Siyue Teng, Qiang Zhu, Chen Feng, Chengxi Zeng, Fan Zhang, Shuyuan Zhu, Bing Zeng, David Bull

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2825] arXiv:2506.14390 (cross-list from cs.LG) [pdf, html, other]: Title: Enclosing Prototypical Variational Autoencoder for Explainable Out-of-Distribution Detection

Conrad Orglmeister, Erik Bochinski, Volker Eiselein, Elvira Fleig

Comments: This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in Computer Safety, Reliability and Security - SAFECOMP 2024 Workshops - DECSoS, SASSUR, TOASTS, and WAISE, and is available online at this https URL

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2826] arXiv:2506.14432 (cross-list from eess.IV) [pdf, html, other]: Title: A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning

Asbjørn Munk, Stefano Cerri, Jakob Ambsdorf, Julia Machnio, Sebastian Nørgaard Llambias, Vardan Nersesjan, Christian Hedeager Krag, Peirong Liu, Pablo Rocamora García, Mostafa Mehdipour Ghazi, Mikael Boesen, Michael Eriksen Benros, Juan Eugenio Iglesias, Mads Nielsen

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2827] arXiv:2506.14497 (cross-list from eess.IV) [pdf, other]: Title: Towards Reliable WMH Segmentation under Domain Shift: An Application Study using Maximum Entropy Regularization to Improve Uncertainty Estimation

Franco Matzkin, Agostina Larrazabal, Diego H Milone, Jose Dolz, Enzo Ferrante

Comments: 32 pages, 7 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2828] arXiv:2506.14513 (cross-list from cs.RO) [pdf, html, other]: Title: GAMORA: A Gesture Articulated Meta Operative Robotic Arm for Hazardous Material Handling in Containment-Level Environments

Farha Abdul Wasay, Mohammed Abdul Rahman, Hania Ghouse

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2829] arXiv:2506.14515 (cross-list from cs.LG) [pdf, html, other]: Title: Train Once, Forget Precisely: Anchored Optimization for Efficient Post-Hoc Unlearning

Prabhav Sanga, Jaskaran Singh, Arun K. Dubey

Comments: Accepted at ICML MUGen'25

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2830] arXiv:2506.14524 (cross-list from eess.IV) [pdf, html, other]: Title: Integrating Radiomics with Deep Learning Enhances Multiple Sclerosis Lesion Delineation

Nadezhda Alsahanova, Pavel Bartenev, Maksim Sharaev, Milos Ljubisavljevic, Taleb Al. Mansoori, Yauhen Statsenko

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2831] arXiv:2506.14542 (cross-list from physics.optics) [pdf, html, other]: Title: MobileHolo: A Lightweight Complex-Valued Deformable CNN for High-Quality Computer-Generated Hologram

Xie Shuyang, Zhou Jie, Xu Bo, Wang Jun, Xu Renjing

Comments: 8 pages, 9 figures

Subjects: Optics (physics.optics); Computer Vision and Pattern Recognition (cs.CV)
[2832] arXiv:2506.14582 (cross-list from cs.CR) [pdf, html, other]: Title: Busting the Paper Ballot: Voting Meets Adversarial Machine Learning

Kaleel Mahmood, Caleb Manicke, Ethan Rathbun, Aayushi Verma, Sohaib Ahmad, Nicholas Stamatakis, Laurent Michel, Benjamin Fuller

Comments: 18 Pages. Author version of article to appear at CCS 2025

Subjects: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2833] arXiv:2506.14698 (cross-list from cs.LG) [pdf, html, other]: Title: Towards Desiderata-Driven Design of Visual Counterfactual Explainers

Sidney Bender, Jan Herrmann, Klaus-Robert Müller, Grégoire Montavon

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2834] arXiv:2506.14719 (cross-list from eess.IV) [pdf, html, other]: Title: Plug-and-Play with 2.5D Artifact Reduction Prior for Fast and Accurate Industrial Computed Tomography Reconstruction

Haley Duba-Sullivan, Aniket Pramanik, Venkatakrishnan Singanallur, Amirkoushyar Ziabari

Comments: Submitted to Journal of Nondestructive Evaluation

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2835] arXiv:2506.14771 (cross-list from eess.IV) [pdf, html, other]: Title: Empirical Studies of Large Scale Environment Scanning by Consumer Electronics

Mengyuan Wang, Yang Liu, Haopeng Wang, Haiwei Dong, Abdulmotaleb El Saddik

Comments: Accepted by IEEE Consumer Electronics Magazine

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET); Multimedia (cs.MM)
[2836] arXiv:2506.14786 (cross-list from cs.LG) [pdf, html, other]: Title: PIPE: Physics-Informed Position Encoding for Alignment of Satellite Images and Time Series

Haobo Li, Eunseo Jung, Zixin Chen, Zhaowei Wang, Yueya Wang, Huamin Qu, Alexis Kai Hon Lau

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2837] arXiv:2506.14803 (cross-list from cs.MM) [pdf, html, other]: Title: Omnidirectional Video Super-Resolution using Deep Learning

Arbind Agrahari Baniya, Tsz-Kwan Lee, Peter W. Eklund, Sunil Aryal

Journal-ref: in IEEE Transactions on Multimedia, vol. 26, pp. 540-554, 2024

Subjects: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2838] arXiv:2506.14821 (cross-list from cs.LG) [pdf, html, other]: Title: Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints

Sunil Kumar, Bowen Zhao, Leo Dirac, Paulina Varshavskaya

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2839] arXiv:2506.14834 (cross-list from eess.IV) [pdf, other]: Title: Deploying and Evaluating Multiple Deep Learning Models on Edge Devices for Diabetic Retinopathy Detection

Akwasi Asare, Dennis Agyemanh Nana Gookyi, Derrick Boateng, Fortunatus Aabangbio Wulnye

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2840] arXiv:2506.14843 (cross-list from cs.LG) [pdf, html, other]: Title: CACTUS as a Reliable Tool for Early Classification of Age-related Macular Degeneration

Luca Gherardini, Imre Lengyel, Tunde Peto, Caroline C.W. Klaverd, Magda A. Meester-Smoord, Johanna Maria Colijnd, EYE-RISK Consortium, E3 Consortium, Jose Sousa

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Applications (stat.AP)
[2841] arXiv:2506.14844 (cross-list from eess.IV) [pdf, other]: Title: Improving Prostate Gland Segmenting Using Transformer based Architectures

Shatha Abudalou

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2842] arXiv:2506.14857 (cross-list from cs.RO) [pdf, html, other]: Title: Towards Perception-based Collision Avoidance for UAVs when Guiding the Visually Impaired

Suman Raj, Swapnil Padhi, Ruchi Bhoot, Prince Modi, Yogesh Simmhan

Comments: 16 pages, 7 figures; Accepted as Late-Breaking Results at the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2023

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2843] arXiv:2506.14864 (cross-list from cs.SD) [pdf, other]: Title: pycnet-audio: A Python package to support bioacoustics data processing

Zachary J. Ruff, Damon B. Lesmeister

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[2844] arXiv:2506.14909 (cross-list from eess.IV) [pdf, other]: Title: Foundation Artificial Intelligence Models for Health Recognition Using Face Photographs (FAHR-Face)

Fridolin Haugg, Grace Lee, John He, Leonard Nürnberg, Dennis Bontempi, Danielle S. Bitterman, Paul Catalano, Vasco Prudente, Dmitrii Glubokov, Andrew Warrington, Suraj Pai, Dirk De Ruysscher, Christian Guthier, Benjamin H. Kann, Vadim N. Gladyshev, Hugo JWL Aerts, Raymond H. Mak

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2845] arXiv:2506.14914 (cross-list from eess.IV) [pdf, html, other]: Title: Recursive Variational Autoencoders for 3D Blood Vessel Generative Modeling

Paula Feldman, Miguel Fainstein, Viviana Siless, Claudio Delrieux, Emmanuel Iarussi

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2846] arXiv:2506.14970 (cross-list from eess.IV) [pdf, html, other]: Title: NeuroMoE: A Transformer-Based Mixture-of-Experts Framework for Multi-Modal Neurological Disorder Classification

Wajih Hassan Raza, Aamir Bader Shah, Yu Wen, Yidan Shen, Juan Diego Martinez Lemus, Mya Caryn Schiess, Timothy Michael Ellmore, Renjie Hu, Xin Fu

Comments: Accepted at the 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2847] arXiv:2506.15029 (cross-list from cs.SD) [pdf, other]: Title: An accurate and revised version of optical character recognition-based speech synthesis using LabVIEW

Prateek Mehta, Anasuya Patil

Comments: 9 pages, 9 figures

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
[2848] arXiv:2506.15084 (cross-list from cs.SE) [pdf, other]: Title: An Empirical Study of Bugs in Data Visualization Libraries

Weiqi Lu, Yongqiang Tian, Xiaohan Zhong, Haoyang Ma, Zhenyang Xu, Shing-Chi Cheung, Chengnian Sun

Comments: Proc. ACM Softw. Eng. 2, FSE

Subjects: Software Engineering (cs.SE); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
[2849] arXiv:2506.15157 (cross-list from cs.RO) [pdf, html, other]: Title: Robust Instant Policy: Leveraging Student's t-Regression Model for Robust In-context Imitation Learning of Robot Manipulation

Hanbit Oh, Andrea M. Salcedo-Vázquez, Ixchel G. Ramirez-Alpizar, Yukiyasu Domae

Comments: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025 accepted

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2850] arXiv:2506.15182 (cross-list from eess.IV) [pdf, html, other]: Title: Classification of Multi-Parametric Body MRI Series Using Deep Learning

Boah Kim, Tejas Sudharshan Mathai, Kimberly Helm, Peter A. Pinto, Ronald M. Summers

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2851] arXiv:2506.15258 (cross-list from eess.IV) [pdf, html, other]: Title: Privacy-Preserving Chest X-ray Classification in Latent Space with Homomorphically Encrypted Neural Inference

Jonghun Kim, Gyeongdeok Jo, Sinyoung Ra, Hyunjin Park

Comments: 11 pages, 5 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2852] arXiv:2506.15290 (cross-list from cs.GR) [pdf, html, other]: Title: Human Motion Capture from Loose and Sparse Inertial Sensors with Garment-aware Diffusion Models

Andela Ilic, Jiaxi Jiang, Paul Streli, Xintong Liu, Christian Holz

Comments: Accepted by IJCAI 2025

Subjects: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
[2853] arXiv:2506.15312 (cross-list from cs.GR) [pdf, html, other]: Title: One-shot Face Sketch Synthesis in the Wild via Generative Diffusion Prior and Instruction Tuning

Han Wu, Junyao Li, Kangbo Zhao, Sen Zhang, Yukai Shi, Liang Lin

Comments: We propose a novel framework for face sketch synthesis, where merely a single pair of samples suffices to enable in-the-wild face sketch synthesis

Subjects: Graphics (cs.GR); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY)
[2854] arXiv:2506.15365 (cross-list from eess.IV) [pdf, html, other]: Title: FedWSIDD: Federated Whole Slide Image Classification via Dataset Distillation

Haolong Jin, Shenglin Liu, Cong Cong, Qingmin Feng, Yongzhi Liu, Lina Huang, Yingzi Hu

Comments: MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2855] arXiv:2506.15395 (cross-list from eess.IV) [pdf, html, other]: Title: A Real-time Endoscopic Image Denoising System

Yu Xing, Shishi Huang, Meng Lv, Guo Chen, Huailiang Wang, Lingzhi Sui

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2856] arXiv:2506.15402 (cross-list from cs.RO) [pdf, html, other]: Title: MCOO-SLAM: A Multi-Camera Omnidirectional Object SLAM System

Miaoxin Pan, Jinnan Li, Yaowen Zhang, Yi Yang, Yufeng Yue

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2857] arXiv:2506.15489 (cross-list from eess.IV) [pdf, other]: Title: Advanced cervical cancer classification: enhancing pap smear images with hybrid PMD Filter-CLAHE

Ach Khozaimi, Isnani Darti, Syaiful Anam, Wuryansari Muharini Kusumawinahyu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2858] arXiv:2506.15499 (cross-list from cs.LG) [pdf, other]: Title: Pixel-level Certified Explanations via Randomized Smoothing

Alaa Anani, Tobias Lorenz, Mario Fritz, Bernt Schiele

Journal-ref: International Conference on Machine Learning (ICML), 2025

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2859] arXiv:2506.15562 (cross-list from eess.IV) [pdf, html, other]: Title: Automated MRI Tumor Segmentation using hybrid U-Net with Transformer and Efficient Attention

Syed Haider Ali, Asrar Ahmad, Muhammad Ali, Asifullah Khan, Nadeem Shaukat

Comments: 16 pages, 5 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2860] arXiv:2506.15677 (cross-list from cs.AI) [pdf, other]: Title: Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence

Yining Hong, Rui Sun, Bingxuan Li, Xingcheng Yao, Maxine Wu, Alexander Chien, Da Yin, Ying Nian Wu, Zhecan James Wang, Kai-Wei Chang

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Robotics (cs.RO)
[2861] arXiv:2506.15680 (cross-list from cs.RO) [pdf, html, other]: Title: Particle-Grid Neural Dynamics for Learning Deformable Object Models from RGB-D Videos

Kaifeng Zhang, Baoyu Li, Kris Hauser, Yunzhu Li

Comments: Project page: this https URL

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2862] arXiv:2506.15684 (cross-list from cs.GR) [pdf, html, other]: Title: Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards

Qingming Liu, Zhen Liu, Dinghuai Zhang, Kui Jia

Comments: Technical Report (21 pages, 21 figures)

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2863] arXiv:2506.15698 (cross-list from cs.LG) [pdf, html, other]: Title: Global Context-aware Representation Learning for Spatially Resolved Transcriptomics

Yunhak Oh, Junseok Lee, Yeongmin Kim, Sangwoo Seo, Namkyeong Lee, Chanyoung Park

Comments: ICML 2025

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2864] arXiv:2506.15711 (cross-list from cs.LG) [pdf, html, other]: Title: Shadow defense against gradient inversion attack in federated learning

Le Jiang, Liyan Ma, Guang Yang

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2865] arXiv:2506.15720 (cross-list from cs.LG) [pdf, html, other]: Title: Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning

Juntae Lee, Munawar Hayat, Sungrack Yun

Comments: Accepted at CVPR 2025

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2866] arXiv:2506.15728 (cross-list from q-bio.QM) [pdf, other]: Title: Smartphone-integrated RPA-CRISPR-Cas12a Detection System with Microneedle Sampling for Point-of-Care Diagnosis of Potato Late Blight in Early Stage

Jiangnan Zhao (1 and 2), Hanbo Xu (1 and 2), Cifu Xu (1 and 2), Wenlong Yin (1 and 2), Laixin Luo (3), Gang Liu (1 and 2), Yan Wang (1 and 2) ((1) Key Laboratory of Smart Agriculture Systems, Ministry of Education, China Agricultural University, Beijing, PR China, (2) Key Laboratory of Agricultural Information Acquisition Technology, Ministry of Agriculture and Rural Affairs of China, China Agricultural University, Beijing, PR China, (3) Department of Plant Pathology, China Agricultural University, Beijing Key Laboratory of Seed Disease Testing and Control, Beijing, PR China)

Comments: 32 pages,7 figures,1 table

Subjects: Quantitative Methods (q-bio.QM); Computer Vision and Pattern Recognition (cs.CV); Biomolecules (q-bio.BM)
[2867] arXiv:2506.15734 (cross-list from cs.AI) [pdf, html, other]: Title: The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models

Peiyuan Tang, Haojie Xin, Xiaodong Zhang, Jun Sun, Qin Xia, Zijiang Yang

Comments: 23 pages, 10 figures

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2868] arXiv:2506.15744 (cross-list from eess.IV) [pdf, other]: Title: Pixel-wise Modulated Dice Loss for Medical Image Segmentation

Seyed Mohsen Hosseini

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2869] arXiv:2506.15748 (cross-list from eess.IV) [pdf, html, other]: Title: Diffusion-based Counterfactual Augmentation: Towards Robust and Interpretable Knee Osteoarthritis Grading

Zhe Wang, Yuhua Ru, Aladine Chetouani, Tina Shiang, Fang Chen, Fabian Bauer, Liping Zhang, Didier Hans, Rachid Jennane, William Ewing Palmer, Mohamed Jarraya, Yung Hsin Chen

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2870] arXiv:2506.15815 (cross-list from cs.GR) [pdf, html, other]: Title: GratNet: A Photorealistic Neural Shader for Diffractive Surfaces

Narayan Kandel, Daljit Singh J.S. Dhillon

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2871] arXiv:2506.15821 (cross-list from cs.GR) [pdf, html, other]: Title: VEIGAR: View-consistent Explicit Inpainting and Geometry Alignment for 3D object Removal

Pham Khai Nguyen Do, Bao Nguyen Tran, Nam Nguyen, Duc Dung Nguyen

Subjects: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
[2872] arXiv:2506.15835 (cross-list from eess.IV) [pdf, html, other]: Title: MoNetV2: Enhanced Motion Network for Freehand 3D Ultrasound Reconstruction

Mingyuan Luo, Xin Yang, Zhongnuo Yan, Yan Cao, Yuanji Zhang, Xindi Hu, Jin Wang, Haoxuan Ding, Wei Han, Litao Sun, Dong Ni

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2873] arXiv:2506.15849 (cross-list from cs.RO) [pdf, html, other]: Title: PRISM-Loc: a Lightweight Long-range LiDAR Localization in Urban Environments with Topological Maps

Kirill Muravyev, Vasily Yuryev, Oleg Bulichev, Dmitry Yudin, Konstantin Yakovlev

Comments: This version was submitted and rejected from IROS 2025 conference

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2874] arXiv:2506.15851 (cross-list from cs.RO) [pdf, html, other]: Title: Semantic and Feature Guided Uncertainty Quantification of Visual Localization for Autonomous Vehicles

Qiyuan Wu, Mark Campbell

Comments: Accepted by ICRA 2025

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2875] arXiv:2506.15853 (cross-list from eess.IV) [pdf, other]: Title: Cross-Modality Learning for Predicting IHC Biomarkers from H&E-Stained Whole-Slide Images

Amit Das, Naofumi Tomita, Kyle J. Syme, Weijie Ma, Paige O'Connor, Kristin N. Corbett, Bing Ren, Xiaoying Liu, Saeed Hassanpour

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2876] arXiv:2506.15888 (cross-list from physics.ins-det) [pdf, html, other]: Title: Bias Variation Compensation in Perimeter-Gated SPAD TRNGs

Md Sakibur Sajal, Hunter Guthrie, Marc Dandin

Comments: 5 pages, 8 figures, 1 software, accepted at MWSCAS 2025 conference

Subjects: Instrumentation and Detectors (physics.ins-det); Hardware Architecture (cs.AR); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2877] arXiv:2506.16050 (cross-list from cs.RO) [pdf, html, other]: Title: Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments

Jiawen Yu, Jieji Ren, Yang Chang, Qiaojun Yu, Xuan Tong, Boyang Wang, Yan Song, You Li, Xinji Mai, Wenqiang Zhang

Comments: IROS 2025 Oral

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2878] arXiv:2506.16102 (cross-list from eess.IV) [pdf, html, other]: Title: Fast Training-free Perceptual Image Compression

Ziran Zhu, Tongda Xu, Minye Huang, Dailan He, Xingtong Ge, Xinjie Zhang, Ling Li, Yan Wang

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2879] arXiv:2506.16116 (cross-list from eess.IV) [pdf, html, other]: Title: Enhanced Dermatology Image Quality Assessment via Cross-Domain Training

Ignacio Hernández Montilla, Alfonso Medela, Paola Pasquali, Andy Aguilar, Taig Mac Carthy, Gerardo Fernández, Antonio Martorell, Enrique Onieva

Comments: 9 pages, 4 figures. This manuscript has been accepted to the 2025 12th International Conference on Bioinformatics Research and Applications (ICBRA 2025). It will be published in International Conference Proceedings by ACM, which will be archived in ACM Digital Library, indexed by Ei Compendex and Scopus

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[2880] arXiv:2506.16201 (cross-list from cs.RO) [pdf, html, other]: Title: FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

Sen Wang, Le Wang, Sanping Zhou, Jingyi Tian, Jiayi Li, Haowen Sun, Wei Tang

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2881] arXiv:2506.16210 (cross-list from eess.IV) [pdf, other]: Title: From Coarse to Continuous: Progressive Refinement Implicit Neural Representation for Motion-Robust Anisotropic MRI Reconstruction

Zhenxuan Zhang, Lipei Zhang, Yanqi Cheng, Zi Wang, Fanwen Wang, Haosen Zhang, Yue Yang, Yinzhe Wu, Jiahao Huang, Angelica I Aviles-Rivero, Zhifan Gao, Guang Yang, Peter J. Lally

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2882] arXiv:2506.16213 (cross-list from eess.IV) [pdf, html, other]: Title: CF-Seg: Counterfactuals meet Segmentation

Raghav Mehta, Fabio De Sousa Ribeiro, Tian Xia, Melanie Roschewitz, Ainkaran Santhirasekaram, Dominic C. Marshall, Ben Glocker

Comments: Accepted at MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2883] arXiv:2506.16256 (cross-list from eess.IV) [pdf, html, other]: Title: AGE-US: automated gestational age estimation based on fetal ultrasound images

César Díaz-Parga, Marta Nuñez-Garcia, Maria J. Carreira, Gabriel Bernardino, Nicolás Vila-Blanco

Comments: Accepted in Iberian Conference on Pattern Recognition and Image Analysis (IbPRIA) 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2884] arXiv:2506.16299 (cross-list from cs.CG) [pdf, html, other]: Title: Wavelet-based Global Orientation and Surface Reconstruction for Point Clouds

Yueji Ma, Yanzun Meng, Dong Xiao, Zuoqiang Shi, Bin Wang

Comments: 22Pages

Subjects: Computational Geometry (cs.CG); Computer Vision and Pattern Recognition (cs.CV)
[2885] arXiv:2506.16349 (cross-list from cs.LG) [pdf, other]: Title: Watermarking Autoregressive Image Generation

Nikola Jovanović, Ismail Labiad, Tomáš Souček, Martin Vechev, Pierre Fernandez

Comments: Code: this https URL

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2886] arXiv:2506.16401 (cross-list from cs.CY) [pdf, other]: Title: TrajSceneLLM: A Multimodal Perspective on Semantic GPS Trajectory Analysis

Chunhou Ji, Qiumeng Li

Comments: Under review for ACM SIGSPATIAL 2025

Subjects: Computers and Society (cs.CY); Computer Vision and Pattern Recognition (cs.CV)
[2887] arXiv:2506.16402 (cross-list from cs.AI) [pdf, other]: Title: IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks

Xiaoya Lu, Zeren Chen, Xuhao Hu, Yijin Zhou, Weichen Zhang, Dongrui Liu, Lu Sheng, Jing Shao

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
[2888] arXiv:2506.16495 (cross-list from cs.MM) [pdf, html, other]: Title: DT-UFC: Universal Large Model Feature Coding via Peaky-to-Balanced Distribution Transformation

Changsheng Gao, Zijie Liu, Li Li, Dong Liu, Xiaoyan Sun, Weisi Lin

Subjects: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV)
[2889] arXiv:2506.16506 (cross-list from cs.LG) [pdf, html, other]: Title: Subspace-Boosted Model Merging

Ronald Skorobogat, Karsten Roth, Mariana-Iuliana Georgescu, Zeynep Akata

Comments: 21 pages (main + supp)

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2890] arXiv:2506.16556 (cross-list from eess.IV) [pdf, html, other]: Title: VesselSDF: Distance Field Priors for Vascular Network Reconstruction

Salvatore Esposito, Daniel Rebain, Arno Onken, Changjian Li, Oisin Mac Aodha

Journal-ref: International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI), 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2891] arXiv:2506.16565 (cross-list from cs.RO) [pdf, html, other]: Title: Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control

Yuxin Chen, Jianglan Wei, Chenfeng Xu, Boyi Li, Masayoshi Tomizuka, Andrea Bajcsy, Ran Tian

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2892] arXiv:2506.16572 (cross-list from eess.IV) [pdf, html, other]: Title: DiffO: Single-step Diffusion for Image Compression at Ultra-Low Bitrates

Chanung Park, Joo Chan Lee, Jong Hwan Ko

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2893] arXiv:2506.16592 (cross-list from eess.IV) [pdf, html, other]: Title: Hybrid Attention Network for Accurate Breast Tumor Segmentation in Ultrasound Images

Muhammad Azeem Aslam, Asim Naveed, Nisar Ahmed

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2894] arXiv:2506.16597 (cross-list from astro-ph.EP) [pdf, html, other]: Title: Exoplanet Classification through Vision Transformers with Temporal Image Analysis

Anupma Choudhary, Sohith Bandari, B.S.Kushvah, C. Swastik

Comments: Accepted for publication in the Astronomical Journal

Subjects: Earth and Planetary Astrophysics (astro-ph.EP); Instrumentation and Methods for Astrophysics (astro-ph.IM); Computer Vision and Pattern Recognition (cs.CV)
[2895] arXiv:2506.16627 (cross-list from cs.GR) [pdf, html, other]: Title: FlatCAD: Fast Curvature Regularization of Neural SDFs for CAD Models

Haotian Yin, Aleksander Plocharski, Michal Jan Wlodarczyk, Mikolaj Kida, Przemyslaw Musialski

Comments: Computer Graphics Forum, Proceedings of Pacific Graphics 2025, 12 pages, 10 figures, preprint

Journal-ref: Computer Graphics Forum, Volume 44 (2025), Number 7

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2896] arXiv:2506.16631 (cross-list from eess.IV) [pdf, html, other]: Title: Overfitting in Histopathology Model Training: The Need for Customized Architectures

Saghir Alfasly, Ghazal Alabtah, H.R. Tizhoosh

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2897] arXiv:2506.16652 (cross-list from cs.RO) [pdf, html, other]: Title: CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

Guang Yin, Yitong Li, Yixuan Wang, Dale McConachie, Paarth Shah, Kunimatsu Hashimoto, Huan Zhang, Katherine Liu, Yunzhu Li

Comments: Accepted to Robotics: Science and Systems (RSS) 2025. The first three authors contributed equally. Project Page: this https URL

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Software Engineering (cs.SE)
[2898] arXiv:2506.16733 (cross-list from eess.IV) [pdf, other]: Title: A Prior-Guided Joint Diffusion Model in Projection Domain for PET Tracer Conversion

Fang Chen, Weifeng Zhang, Xingyu Ai, BingXuan Li, An Li, Qiegen Liu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2899] arXiv:2506.16760 (cross-list from cs.CL) [pdf, html, other]: Title: Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models

Lei Jiang, Zixun Zhang, Zizhou Wang, Xiaobing Sun, Zhen Li, Liangli Zhen, Xiaohua Xu

Comments: 15 pages, 9 figures

Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2900] arXiv:2506.16803 (cross-list from eess.IV) [pdf, html, other]: Title: Temperature calibration of surface emissivities with an improved thermal image enhancement network

Ning Chu, Siya Zheng, Shanqing Zhang, Li Li, Caifang Cai, Ali Mohammad-Djafari, Feng Zhao, Yuanbo Song

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2901] arXiv:2506.16827 (cross-list from cs.GR) [pdf, html, other]: Title: Beyond Blur: A Fluid Perspective on Generative Diffusion Models

Grzegorz Gruszczynski, Jakub Meixner, Michal Jan Wlodarczyk, Przemyslaw Musialski

Comments: ICCV 2025 main conference, 8 pages paper, 20 pages appendix, 24 figures, supplementary pseudocode in appendix, this https URL

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2902] arXiv:2506.16890 (cross-list from cs.LG) [pdf, html, other]: Title: From Lab to Factory: Pitfalls and Guidelines for Self-/Unsupervised Defect Detection on Low-Quality Industrial Images

Sebastian Hönel, Jonas Nordqvist

Comments: 18 pages, 7 figures, 1 table. Camera-ready version for the 2025 conference European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD '25)

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Applications (stat.AP)
[2903] arXiv:2506.16898 (cross-list from cs.AI) [pdf, html, other]: Title: AI's Blind Spots: Geographic Knowledge and Diversity Deficit in Generated Urban Scenario

Ciro Beneduce, Massimiliano Luca, Bruno Lepri

Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY)
[2904] arXiv:2506.16934 (cross-list from eess.IV) [pdf, other]: Title: PET Tracer Separation Using Conditional Diffusion Transformer with Multi-latent Space Learning

Bin Huang, Feihong Xu, Xinchong Shi, Shan Huang, Binxuan Li, Fei Li, Qiegen Liu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2905] arXiv:2506.17110 (cross-list from cs.RO) [pdf, html, other]: Title: Monocular One-Shot Metric-Depth Alignment for RGB-Based Robot Grasping

Teng Guo, Baichuan Huang, Jingjin Yu

Comments: Accepted to IROS 2025

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2906] arXiv:2506.17133 (cross-list from eess.IV) [pdf, html, other]: Title: Robust Training with Data Augmentation for Medical Imaging Classification

Josué Martínez-Martínez, Olivia Brown, Mostafa Karami, Sheida Nabavi

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2907] arXiv:2506.17140 (cross-list from eess.IV) [pdf, html, other]: Title: MeDi: Metadata-Guided Diffusion Models for Mitigating Biases in Tumor Classification

David Jacob Drexlin, Jonas Dippel, Julius Hense, Niklas Prenißl, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2908] arXiv:2506.17165 (cross-list from eess.IV) [pdf, html, other]: Title: Proportional Sensitivity in Generative Adversarial Network (GAN)-Augmented Brain Tumor Classification Using Convolutional Neural Network

Mahin Montasir Afif, Abdullah Al Noman, K. M. Tahsin Kabir, Md. Mortuza Ahmmed, Md. Mostafizur Rahman, Mufti Mahmud, Md. Ashraful Babu

Comments: This papaer has been submitted to The 18th International Conference on Brain Informatics (BI'25), Italy

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2909] arXiv:2506.17198 (cross-list from cs.RO) [pdf, html, other]: Title: Dex1B: Learning with 1B Demonstrations for Dexterous Manipulation

Jianglong Ye, Keyi Wang, Chengjing Yuan, Ruihan Yang, Yiquan Li, Jiyue Zhu, Yuzhe Qin, Xueyan Zou, Xiaolong Wang

Comments: Accepted to RSS 2025. Project page: this https URL

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2910] arXiv:2506.17206 (cross-list from cs.GR) [pdf, html, other]: Title: DreamCube: 3D Panorama Generation via Multi-plane Synchronization

Yukun Huang, Yanning Zhou, Jianan Wang, Kaiyi Huang, Xihui Liu

Comments: Project page: this https URL

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2911] arXiv:2506.17232 (cross-list from cs.LG) [pdf, html, other]: Title: PCaM: A Progressive Focus Attention-Based Information Fusion Method for Improving Vision Transformer Domain Adaptation

Zelin Zang, Fei Wang, Liangyu Li, Jinlin Wu, Chunshui Zhao, Zhen Lei, Baigui Sun

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2912] arXiv:2506.17307 (cross-list from cs.LG) [pdf, html, other]: Title: Learning to Adapt Frozen CLIP for Few-Shot Test-Time Domain Adaptation

Zhixiang Chi, Li Gu, Huan Liu, Ziqiang Wang, Yanan Wu, Yang Wang, Konstantinos N Plataniotis

Comments: ICLR2025,this https URL

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2913] arXiv:2506.17320 (cross-list from cs.CY) [pdf, other]: Title: MAARTA:Multi-Agentic Adaptive Radiology Teaching Assistant

Akash Awasthi, Brandon V. Chang, Anh M. Vu, Ngan Le, Rishi Agrawal, Zhigang Deng, Carol Wu, Hien Van Nguyen

Comments: Accepted to MICCAI 2025 (Main Conference)

Subjects: Computers and Society (cs.CY); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2914] arXiv:2506.17324 (cross-list from cs.LG) [pdf, other]: Title: Origins of Creativity in Attention-Based Diffusion Models

Emma Finn, T. Anderson Keller, Manos Theodosis, Demba E. Ba

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2915] arXiv:2506.17337 (cross-list from eess.IV) [pdf, html, other]: Title: Can Common VLMs Rival Medical VLMs? Evaluation and Strategic Insights

Yuan Zhong, Ruinan Jin, Xiaoxiao Li, Qi Dou

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2916] arXiv:2506.17361 (cross-list from eess.IV) [pdf, html, other]: Title: Efficient Feedback Gate Network for Hyperspectral Image Super-Resolution

Xufei Wang, Mingjian Zhang, Fei Ge, Jinchen Zhu, Wen Sha, Jifen Ren, Zhimeng Hou, Shouguo Zheng, ling Zheng, Shizhuang Weng

Comments: 20 pages,17 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2917] arXiv:2506.17364 (cross-list from cs.CY) [pdf, html, other]: Title: AI-based Multimodal Biometrics for Detecting Smartphone Distractions: Application to Online Learning

Alvaro Becerra, Roberto Daza, Ruth Cobos, Aythami Morales, Mutlu Cukurova, Julian Fierrez

Comments: Accepted in EC-TEL25: 20th European Conference on Technology Enhanced Learning, Newcastle and Durham, UK, 15-19 September 2025

Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
[2918] arXiv:2506.17372 (cross-list from cs.CY) [pdf, html, other]: Title: Multimodal Political Bias Identification and Neutralization

Cedric Bernard, Xavier Pleimling, Amun Kharel, Chase Vickery

Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2919] arXiv:2506.17378 (cross-list from cs.RO) [pdf, html, other]: Title: A workflow for generating synthetic LiDAR datasets in simulation environments

Abhishek Phadke, Shakib Mahmud Dipto, Pratip Rana

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2920] arXiv:2506.17412 (cross-list from eess.IV) [pdf, html, other]: Title: VMRA-MaR: An Asymmetry-Aware Temporal Framework for Longitudinal Breast Cancer Risk Prediction

Zijun Sun, Solveig Thrun, Michael Kampffmeyer

Comments: MICCAI 2025, Provisional Accept

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2921] arXiv:2506.17425 (cross-list from eess.IV) [pdf, html, other]: Title: Trans${^2}$-CBCT: A Dual-Transformer Framework for Sparse-View CBCT Reconstruction

Minmin Yang, Huantao Ren, Senem Velipasalar

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2922] arXiv:2506.17462 (cross-list from cs.RO) [pdf, other]: Title: General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting

Bernard Lange, Anil Yildiz, Mansur Arief, Shehryar Khattak, Mykel Kochenderfer, Georgios Georgakis

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2923] arXiv:2506.17501 (cross-list from eess.IV) [pdf, html, other]: Title: DSA-NRP: No-Reflow Prediction from Angiographic Perfusion Dynamics in Stroke EVT

Shreeram Athreya, Carlos Olivares, Ameera Ismail, Kambiz Nael, William Speier, Corey Arnold

Comments: 12 pages, 4 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2924] arXiv:2506.17516 (cross-list from cs.RO) [pdf, html, other]: Title: EASE: Embodied Active Event Perception via Self-Supervised Energy Minimization

Zhou Chen, Sanjoy Kundu, Harsimran S. Baweja, Sathyanarayanan N. Aakur

Comments: Accepted to IEEE Robotics and Automation Letters, 2025

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2925] arXiv:2506.17540 (cross-list from eess.IV) [pdf, html, other]: Title: MTSIC: Multi-stage Transformer-based GAN for Spectral Infrared Image Colorization

Tingting Liu, Yuan Liu, Jinhui Tang, Liyin Yuan, Chengyu Liu, Chunlai Li, Xiubao Sui, Qian Chen

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2926] arXiv:2506.17552 (cross-list from cs.LG) [pdf, other]: Title: DRIMV_TSK: An Interpretable Surgical Evaluation Model for Incomplete Multi-View Rectal Cancer Data

Wei Zhang, Zi Wang, Hanwen Zhou, Zhaohong Deng, Weiping Ding, Yuxi Ge, Te Zhang, Yuanpeng Zhang, Kup-Sze Choi, Shitong Wang, Shudong Hu

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2927] arXiv:2506.17623 (cross-list from cs.MM) [pdf, html, other]: Title: Can Generated Images Serve as a Viable Modality for Text-Centric Multimodal Learning?

Yuesheng Huang, Peng Zhang, Riliang Liu, Jiaqi Liang

Comments: 4 figures,7 tables

Subjects: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV)
[2928] arXiv:2506.17636 (cross-list from cs.GR) [pdf, html, other]: Title: 3D Gaussian Splatting for Fine-Detailed Surface Reconstruction in Large-Scale Scene

Shihan Chen, Zhaojin Li, Zeyu Chen, Qingsong Yan, Gaoyang Shen, Ran Duan

Comments: IROS 2025

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
[2929] arXiv:2506.17747 (cross-list from physics.geo-ph) [pdf, html, other]: Title: Pix2Geomodel: A Next-Generation Reservoir Geomodeling with Property-to-Property Translation

Abdulrahman Al-Fakih, Ardiansyah Koeshidayatullah, Nabil A. Saraih, Tapan Mukerji, Rayan Kanfar, Abdulmohsen Alali, SanLinn I. Kaka

Comments: 34 pages, 13 figures

Subjects: Geophysics (physics.geo-ph); Computational Engineering, Finance, and Science (cs.CE); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
[2930] arXiv:2506.17770 (cross-list from cs.GR) [pdf, html, other]: Title: Collaborative Texture Filtering

Tomas Akenine-Möller, Pontus Ebelin, Matt Pharr, Bartlomiej Wronski

Comments: Accepted to ACM/EG Symposium on High Performance Graphics (HPG), 2025

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2931] arXiv:2506.17872 (cross-list from cs.LG) [pdf, html, other]: Title: Decoding Federated Learning: The FedNAM+ Conformal Revolution

Sree Bhargavi Balija, Amitash Nanda, Debashis Sahoo

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2932] arXiv:2506.17874 (cross-list from stat.ML) [pdf, html, other]: Title: DRO-Augment Framework: Robustness by Synergizing Wasserstein Distributionally Robust Optimization and Data Augmentation

Jiaming Hu, Debarghya Mukherjee, Ioannis Ch. Paschalidis

Comments: 26 pages,3 figures

Subjects: Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2933] arXiv:2506.17879 (cross-list from eess.IV) [pdf, html, other]: Title: StainPIDR: A Pathological Image Decouplingand Reconstruction Method for Stain Normalization Based on Color Vector Quantization and Structure Restaining

Zheng Chen

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2934] arXiv:2506.17954 (cross-list from eess.IV) [pdf, other]: Title: Mobile Image Analysis Application for Mantoux Skin Test

Liong Gele, Tan Chye Cheah

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2935] arXiv:2506.17966 (cross-list from cs.IR) [pdf, html, other]: Title: LLM-Enhanced Multimodal Fusion for Cross-Domain Sequential Recommendation

Wangyu Wu, Zhenhong Chen, Xianglin Qiu, Siqi Song, Xiaowei Huang, Fei Ma, Jimin Xiao

Comments: arXiv admin note: substantial text overlap with arXiv:2504.15085

Subjects: Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV)
[2936] arXiv:2506.17967 (cross-list from cs.LG) [pdf, html, other]: Title: Adapting Vision-Language Models for Evaluating World Models

Mariya Hendriksen, Tabish Rashid, David Bignell, Raluca Georgescu, Abdelhak Lemkhenter, Katja Hofmann, Sam Devlin, Sarah Parisot

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2937] arXiv:2506.17968 (cross-list from cs.LG) [pdf, html, other]: Title: h-calibration: Rethinking Classifier Recalibration with Probabilistic Error-Bounded Objective

Wenjian Huang, Guiping Cao, Jiahao Xia, Jingkun Chen, Hao Wang, Jianguo Zhang

Journal-ref: IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Probability (math.PR); Machine Learning (stat.ML)
[2938] arXiv:2506.17983 (cross-list from eess.IV) [pdf, html, other]: Title: LVPNet: A Latent-variable-based Prediction-driven End-to-end Framework for Lossless Compression of Medical Images

Chenyue Song, Chen Hui, Qing Lin, Wei Zhang, Siqiao Li, Haiqi Zhu, Zhixuan Li, Shengping Zhang, Shaohui Liu, Feng Jiang, Xiang Li

Comments: Accepted to MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2939] arXiv:2506.18017 (cross-list from cs.GR) [pdf, html, other]: Title: Auto-Regressive Surface Cutting

Yang Li, Victor Cheung, Xinhai Liu, Yuguang Chen, Zhongjin Luo, Biwen Lei, Haohan Weng, Zibo Zhao, Jingwei Huang, Zhuo Chen, Chunchao Guo

Comments: Tech. report. this https URL

Subjects: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2940] arXiv:2506.18069 (cross-list from cs.DL) [pdf, html, other]: Title: Unfolding the Past: A Comprehensive Deep Learning Approach to Analyzing Incunabula Pages

Klaudia Ropel, Krzysztof Kutt, Luiz do Valle Miranda, Grzegorz J. Nalepa

Comments: 10 pages, 8 figures; submitted to TPDL 2025; change in v2: updated e-mail address

Subjects: Digital Libraries (cs.DL); Computer Vision and Pattern Recognition (cs.CV)
[2941] arXiv:2506.18072 (cross-list from eess.IV) [pdf, html, other]: Title: Multimodal Medical Image Binding via Shared Text Embeddings

Yunhao Liu, Suyang Xi, Shiqi Liu, Hong Ding, Chicheng Jin, Chong Zhong, Junjun He, Catherine C. Liu, Yiqing Shen

Comments: 10 pages, 3 figures

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2942] arXiv:2506.18088 (cross-list from cs.RO) [pdf, html, other]: Title: RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, Weiliang Deng, Yubin Guo, Tian Nian, Xuanbing Xie, Qiangyu Chen, Kailun Su, Tianling Xu, Guodong Liu, Mengkang Hu, Huan-ang Gao, Kaixuan Wang, Zhixuan Liang, Yusen Qin, Xiaokang Yang, Ping Luo, Yao Mu

Comments: Project Page: this https URL, Code: this https URL, Doc: this https URL

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA)
[2943] arXiv:2506.18158 (cross-list from cs.AI) [pdf, html, other]: Title: Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation

Xinzge Gao, Chuanrui Hu, Bin Chen, Teng Li

Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2944] arXiv:2506.18162 (cross-list from cs.LG) [pdf, html, other]: Title: Pitfalls of Conformal Predictions for Medical Image Classification

Hendrik Mehrtens, Tabea Bucher, Titus J. Brinker

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2945] arXiv:2506.18172 (cross-list from eess.IV) [pdf, html, other]: Title: STACT-Time: Spatio-Temporal Cross Attention for Cine Thyroid Ultrasound Time Series Classification

Irsyad Adam, Tengyue Zhang, Shrayes Raman, Zhuyu Qiu, Brandon Taraku, Hexiang Feng, Sile Wang, Ashwath Radhachandran, Shreeram Athreya, Vedrana Ivezic, Peipei Ping, Corey Arnold, William Speier

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2946] arXiv:2506.18201 (cross-list from cs.CL) [pdf, html, other]: Title: Deciphering Emotions in Children Storybooks: A Comparative Analysis of Multimodal LLMs in Educational Applications

Bushra Asseri, Estabraq Abdelaziz, Maha Al Mogren, Tayef Alhefdhi, Areej Al-Wabil

Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
[2947] arXiv:2506.18251 (cross-list from cs.GR) [pdf, html, other]: Title: Morse: Dual-Sampling for Lossless Acceleration of Diffusion Models

Chao Li, Jiawei Fan, Anbang Yao

Comments: Fixed a prompt typo in Figure 18 of the Appendix. This work is accepted to ICML 2025. The project page: this https URL

Subjects: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2948] arXiv:2506.18270 (cross-list from eess.IV) [pdf, other]: Title: Adaptive Mask-guided K-space Diffusion for Accelerated MRI Reconstruction

Qinrong Cai, Yu Guan, Zhibo Chen, Dong Liang, Qiuyun Fan, Qiegen Liu

Comments: 10 pages, 9 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2949] arXiv:2506.18323 (cross-list from eess.IV) [pdf, html, other]: Title: A Multi-Scale Spatial Attention-Based Zero-Shot Learning Framework for Low-Light Image Enhancement

Muhammad Azeem Aslam, Hassan Khalid, Nisar Ahmed

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2950] arXiv:2506.18335 (cross-list from eess.IV) [pdf, html, other]: Title: Rethinking Decoder Design: Improving Biomarker Segmentation Using Depth-to-Space Restoration and Residual Linear Attention

Saad Wazir, Daeyoung Kim

Comments: Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 30861-30871

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2951] arXiv:2506.18371 (cross-list from eess.IV) [pdf, html, other]: Title: Transforming H&E images into IHC: A Variance-Penalized GAN for Precision Oncology

Sara Rehmat, Hafeez Ur Rehman

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2952] arXiv:2506.18378 (cross-list from eess.IV) [pdf, html, other]: Title: Taming Vision-Language Models for Medical Image Analysis: A Comprehensive Review

Haoneng Lin, Cheng Xu, Jing Qin

Comments: 34 pages

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2953] arXiv:2506.18407 (cross-list from cs.GR) [pdf, html, other]: Title: What You Think Is What You Get: Bridge User Intent and Transfer Function Design through Multimodal Large Language Models

Yiyao Wang, Bo Pan, Ke Wang, Han Liu, Jinyuan Mao, Yuxin Liu, Minfeng Zhu, Bo Zhang, Weifeng Chen, Xiuqi Huang, Wei Chen

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2954] arXiv:2506.18443 (cross-list from cs.RO) [pdf, html, other]: Title: Radar and Event Camera Fusion for Agile Robot Ego-Motion Estimation

Yang Lyu, Zhenghao Zou, Yanfeng Li, Chunhui Zhao, Quan Pan

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2955] arXiv:2506.18474 (cross-list from eess.IV) [pdf, html, other]: Title: A Deep Convolutional Neural Network-Based Novel Class Balancing for Imbalance Data Segmentation

Atifa Kalsoom, M.A. Iftikhar, Amjad Ali, Zubair Shah, Shidin Balakrishnan, Hazrat Ali

Comments: This is preprint of the paper submitted to Scientific Reports journal

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2956] arXiv:2506.18484 (cross-list from eess.IV) [pdf, html, other]: Title: GANs vs. Diffusion Models for virtual staining with the HER2match dataset

Pascal Klöckner, José Teixeira, Diana Montezuma, Jaime S. Cardoso, Hugo M. Horlings, Sara P. Oliveira

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2957] arXiv:2506.18512 (cross-list from eess.IV) [pdf, html, other]: Title: MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis

Yuting Zhang, Kaishen Yuan, Hao Lu, Yutao Yue, Jintai Chen, Kaishun Wu

Subjects: Image and Video Processing (eess.IV); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Quantitative Methods (q-bio.QM)
[2958] arXiv:2506.18598 (cross-list from cs.LG) [pdf, html, other]: Title: No Training Wheels: Steering Vectors for Bias Correction at Inference Time

Aviral Gupta, Armaan Sethi, Ameesh Sethi

Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2959] arXiv:2506.18601 (cross-list from cs.GR) [pdf, html, other]: Title: BulletGen: Improving 4D Reconstruction with Bullet-Time Generation

Denys Rozumnyi, Jonathon Luiten, Numair Khan, Johannes Schönberger, Peter Kontschieder

Subjects: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2960] arXiv:2506.18671 (cross-list from cs.SD) [pdf, html, other]: Title: TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography

Yuqin Dai, Wanlu Zhu, Ronghui Li, Xiu Li, Zhenyu Zhang, Jun Li, Jian Yang

Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Audio and Speech Processing (eess.AS)
[2961] arXiv:2506.18680 (cross-list from cs.GR) [pdf, html, other]: Title: DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling

Anindita Ghosh, Bing Zhou, Rishabh Dabral, Jian Wang, Vladislav Golyanik, Christian Theobalt, Philipp Slusallek, Chuan Guo

Comments: 11 pages, 7 figures, 2 tables, accepted in ACM Siggraph 2025 conference track

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[2962] arXiv:2506.18720 (cross-list from eess.IV) [pdf, html, other]: Title: Temporal Neural Cellular Automata: Application to modeling of contrast enhancement in breast MRI

Daniel M. Lang, Richard Osuala, Veronika Spieker, Karim Lekadir, Rickmer Braren, Julia A. Schnabel

Comments: MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2963] arXiv:2506.18725 (cross-list from cs.RO) [pdf, html, other]: Title: TopoRec: Point Cloud Recognition Using Topological Data Analysis

Anirban Ghosh, Iliya Kulbaka, Ian Dahlin, Ayan Dutta

Subjects: Robotics (cs.RO); Computational Geometry (cs.CG); Computer Vision and Pattern Recognition (cs.CV)
[2964] arXiv:2506.18810 (cross-list from cs.AI) [pdf, html, other]: Title: ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation

Siao Tang, Xinyin Ma, Gongfan Fang, Xinchao Wang

Comments: Codes are available at this https URL

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2965] arXiv:2506.18842 (cross-list from cs.DB) [pdf, html, other]: Title: LIGHTHOUSE: Fast and precise distance to shoreline calculations from anywhere on earth

Patrick Beukema, Henry Herzog, Yawen Zhang, Hunter Pitelka, Favyen Bastani

Comments: 8 pages, 7 figures, 1 table, ICML 2025 ML4RS

Subjects: Databases (cs.DB); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2966] arXiv:2506.18844 (cross-list from cs.RO) [pdf, other]: Title: Reproducible Evaluation of Camera Auto-Exposure Methods in the Field: Platform, Benchmark and Lessons Learned

Olivier Gamache, Jean-Michel Fortin, Matěj Boxan, François Pomerleau, Philippe Giguère

Comments: 19 pages, 11 figures, pre-print version of the accepted paper for IEEE Transactions on Field Robotics (T-FR)

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2967] arXiv:2506.18885 (cross-list from cs.RO) [pdf, html, other]: Title: GRAND-SLAM: Local Optimization for Globally Consistent Large-Scale Multi-Agent Gaussian SLAM

Annika Thomas, Aneesa Sonawalla, Alex Rose, Jonathan P. How

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2968] arXiv:2506.18919 (cross-list from cs.CL) [pdf, html, other]: Title: MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection

Hexiang Gu, Qifan Yu, Saihui Hou, Zhiqin Fang, Huijia Wu, Zhaofeng He

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2969] arXiv:2506.19051 (cross-list from eess.IV) [pdf, html, other]: Title: NIC-RobustBench: A Comprehensive Open-Source Toolkit for Neural Image Compression and Robustness Analysis

Georgii Bychkov, Khaled Abud, Egor Kovalev, Alexander Gushchin, Dmitriy Vatolin, Anastasia Antsiferova

Comments: arXiv admin note: text overlap with arXiv:2411.11795

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[2970] arXiv:2506.19055 (cross-list from eess.IV) [pdf, html, other]: Title: Xray2Xray: World Model from Chest X-rays with Volumetric Context

Zefan Yang, Xinrui Song, Xuanang Xu, Yongyi Shi, Ge Wang, Mannudeep K. Kalra, Pingkun Yan

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2971] arXiv:2506.19106 (cross-list from eess.IV) [pdf, html, other]: Title: Staining normalization in histopathology: Method benchmarking using multicenter dataset

Umair Khan, Jouni Härkönen, Marjukka Friman, Leena Latonen, Teijo Kuopio, Pekka Ruusuvuori

Comments: 18 pages, 9 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Tissues and Organs (q-bio.TO)
[2972] arXiv:2506.19139 (cross-list from cs.GR) [pdf, html, other]: Title: SOF: Sorted Opacity Fields for Fast Unbounded Surface Reconstruction

Lukas Radl, Felix Windisch, Thomas Deixelberger, Jozef Hladky, Michael Steiner, Dieter Schmalstieg, Markus Steinberger

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[2973] arXiv:2506.19167 (cross-list from eess.IV) [pdf, other]: Title: A Deep Learning Based Method for Fast Registration of Cardiac Magnetic Resonance Images

Benjamin Graham

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2974] arXiv:2506.19222 (cross-list from eess.IV) [pdf, html, other]: Title: Deformable Medical Image Registration with Effective Anatomical Structure Representation and Divide-and-Conquer Network

Xinke Ma, Yongsheng Pan, Qingjie Zeng, Mengkang Lu, Bolysbek Murat Yerzhanuly, Bazargul Matkerim, Yong Xia

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2975] arXiv:2506.19234 (cross-list from eess.IV) [pdf, html, other]: Title: Quantitative Benchmarking of Anomaly Detection Methods in Digital Pathology

Can Cui, Xindong Zheng, Ruining Deng, Quan Liu, Tianyuan Yao, Keith T Wilson, Lori A Coburn, Bennett A Landman, Haichun Yang, Yaohong Wang, Yuankai Huo

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2976] arXiv:2506.19266 (cross-list from q-bio.NC) [pdf, other]: Title: Convergent and divergent connectivity patterns of the arcuate fasciculus in macaques and humans

Jiahao Huang, Ruifeng Li, Wenwen Yu, Anan Li, Xiangning Li, Mingchao Yan, Lei Xie, Qingrun Zeng, Xueyan Jia, Shuxin Wang, Ronghui Ju, Feng Chen, Qingming Luo, Hui Gong, Andrew Zalesky, Xiaoquan Yang, Yuanjing Feng, Zheng Wang

Comments: 34 pages, 6 figures

Subjects: Neurons and Cognition (q-bio.NC); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
[2977] arXiv:2506.19297 (cross-list from eess.IV) [pdf, html, other]: Title: Explicit Residual-Based Scalable Image Coding for Humans and Machines

Yui Tatsumi, Ziyue Zeng, Hiroshi Watanabe

Comments: Accepted to IEEE 27th International Workshop on Multimedia Signal Processing (MMSP 2025)

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2978] arXiv:2506.19360 (cross-list from cs.CR) [pdf, html, other]: Title: SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation

Yunsung Chung, Yunbei Zhang, Nassir Marrouche, Jihun Hamm

Comments: Accepted at the 34th USENIX Security Symposium (USENIX Security '25). 21 pages, plus a 6-page appendix

Subjects: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2979] arXiv:2506.19363 (cross-list from eess.IV) [pdf, html, other]: Title: Reconsidering Explicit Longitudinal Mammography Alignment for Enhanced Breast Cancer Risk Prediction

Solveig Thrun, Stine Hansen, Zijun Sun, Nele Blum, Suaiba A. Salahuddin, Kristoffer Wickstrøm, Elisabeth Wetzer, Robert Jenssen, Maik Stille, Michael Kampffmeyer

Comments: MICCAI 2025, early accepted

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2980] arXiv:2506.19387 (cross-list from eess.IV) [pdf, other]: Title: NAADA: A Noise-Aware Attention Denoising Autoencoder for Dental Panoramic Radiographs

Khuram Naveed, Bruna Neves de Freitas, Ruben Pauwels

Comments: 10 pages, 8 figures

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2981] arXiv:2506.19415 (cross-list from cs.GR) [pdf, html, other]: Title: Virtual Memory for 3D Gaussian Splatting

Jonathan Haberl, Philipp Fleck, Clemens Arth

Comments: Based on the Master Thesis from Jonathan Haberl from 2024, Submitted to TVCG in Feb. 2025;

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
[2982] arXiv:2506.19455 (cross-list from eess.IV) [pdf, html, other]: Title: Angio-Diff: Learning a Self-Supervised Adversarial Diffusion Model for Angiographic Geometry Generation

Zhifeng Wang, Renjiao Yi, Xin Wen, Chenyang Zhu, Kai Xu, Kunlun He

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2983] arXiv:2506.19464 (cross-list from eess.IV) [pdf, html, other]: Title: Assessing Risk of Stealing Proprietary Models for Medical Imaging Tasks

Ankita Raj, Harsh Swaika, Deepankar Varma, Chetan Arora

Comments: Accepted to MICCAI 2024

Subjects: Image and Video Processing (eess.IV); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[2984] arXiv:2506.19491 (cross-list from cs.ET) [pdf, html, other]: Title: Experimental Assessment of Neural 3D Reconstruction for Small UAV-based Applications

Genís Castillo Gómez-Raya, Álmos Veres-Vitályos, Filip Lemic, Pablo Royo, Mario Montagud, Sergi Fernández, Sergi Abadal, Xavier Costa-Pérez

Comments: 6 pages, 7 figures, 2 tables, accepted at IEEE International Symposium on Personal, Indoor and Mobile Radio Communications 2025

Subjects: Emerging Technologies (cs.ET); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Networking and Internet Architecture (cs.NI); Image and Video Processing (eess.IV)
[2985] arXiv:2506.19558 (cross-list from cs.LG) [pdf, html, other]: Title: ConCM: Consistency-Driven Calibration and Matching for Few-Shot Class-Incremental Learning

QinZhe Wang, Zixuan Chen, Keke Huang, Xiu Su, Chunhua Yang, Chang Xu

Comments: 9 pages, 5 figures(Excluding the appendix)

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[2986] arXiv:2506.19579 (cross-list from cs.RO) [pdf, html, other]: Title: Fake or Real, Can Robots Tell? Evaluating Embodied Vision-Language Models on Real and 3D-Printed Objects

Federico Tavella, Kathryn Mearns, Angelo Cangelosi

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[2987] arXiv:2506.19590 (cross-list from eess.IV) [pdf, html, other]: Title: Learning from Anatomy: Supervised Anatomical Pretraining (SAP) for Improved Metastatic Bone Disease Segmentation in Whole-Body MRI

Joris Wuts, Jakub Ceranka, Nicolas Michoux, Frédéric Lecouvet, Jef Vandemeulebroucke

Comments: This preprint is currently under review at *Computers in Biology and Medicine* (Elsevier). This version has not been peer-reviewed

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2988] arXiv:2506.19600 (cross-list from eess.IV) [pdf, html, other]: Title: Filling of incomplete sinograms from sparse PET detector configurations using a residual U-Net

Klara Leffler, Luigi Tommaso Luppino, Samuel Kuttner, Karin Söderkvist, Jan Axelsson

Comments: 15 pages, 9 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Medical Physics (physics.med-ph)
[2989] arXiv:2506.19687 (cross-list from eess.IV) [pdf, html, other]: Title: ReCoGNet: Recurrent Context-Guided Network for 3D MRI Prostate Segmentation

Ahmad Mustafa, Reza Rastegar, Ghassan AlRegib

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2990] arXiv:2506.19708 (cross-list from cs.GR) [pdf, html, other]: Title: Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders

Matyas Bohacek, Thomas Fel, Maneesh Agrawala, Ekdeep Singh Lubana

Subjects: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2991] arXiv:2506.19741 (cross-list from cs.LG) [pdf, html, other]: Title: Noise Consistency Training: A Native Approach for One-Step Generator in Learning Additional Controls

Yihong Luo, Shuchen Xue, Tianyang Hu, Jing Tang

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
[2992] arXiv:2506.19742 (cross-list from eess.IV) [pdf, html, other]: Title: NeRF-based CBCT Reconstruction needs Normalization and Initialization

Zhuowei Xu, Han Li, Dai Sun, Zhicheng Li, Yujia Li, Qingpeng Kong, Zhiwei Cheng, Nassir Navab, S. Kevin Zhou

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[2993] arXiv:2506.19797 (cross-list from eess.IV) [pdf, html, other]: Title: Systematic Review of Pituitary Gland and Pituitary Adenoma Automatic Segmentation Techniques in Magnetic Resonance Imaging

Mubaraq Yakubu, Navodini Wijethilake, Jonathan Shapey, Andrew King, Alexander Hammers

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[2994] arXiv:2506.19807 (cross-list from cs.AI) [pdf, other]: Title: KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality

Baochang Ren, Shuofei Qiao, Wenhao Yu, Huajun Chen, Ningyu Zhang

Comments: Work in progress

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
[2995] arXiv:2506.19816 (cross-list from cs.RO) [pdf, html, other]: Title: CronusVLA: Transferring Latent Motion Across Time for Multi-Frame Prediction in Manipulation

Hao Li, Shuai Yang, Yilun Chen, Yang Tian, Xiaoda Yang, Xinyi Chen, Hanqing Wang, Tai Wang, Feng Zhao, Dahua Lin, Jiangmiao Pang

Comments: 36 pages, 21 figures

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2996] arXiv:2506.19827 (cross-list from cs.RO) [pdf, html, other]: Title: Look to Locate: Vision-Based Multisensory Navigation with 3-D Digital Maps for GNSS-Challenged Environments

Ola Elmaghraby, Eslam Mounier, Paulo Ricardo Marques de Araujo, Aboelmagd Noureldin

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[2997] arXiv:2506.19847 (cross-list from cs.LG) [pdf, html, other]: Title: Orthogonal Finetuning Made Scalable

Zeju Qiu, Weiyang Liu, Adrian Weller, Bernhard Schölkopf

Comments: Technical report (17 pages, 7 figures, project page: this https URL)

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[2998] arXiv:2506.19860 (cross-list from eess.SP) [pdf, html, other]: Title: A Multi-Modal Spatial Risk Framework for EV Charging Infrastructure Using Remote Sensing

Oktay Karakuş, Padraig Corcoran

Comments: 11 pages, 4 figures, 2 tables

Subjects: Signal Processing (eess.SP); Computer Vision and Pattern Recognition (cs.CV)
[2999] arXiv:2506.19935 (cross-list from cs.LG) [pdf, html, other]: Title: Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture

Shuchen Xue, Tianyu Xie, Tianyang Hu, Zijin Feng, Jiacheng Sun, Kenji Kawaguchi, Zhenguo Li, Zhi-Ming Ma

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
[3000] arXiv:2506.19975 (cross-list from eess.IV) [pdf, html, other]: Title: VoxelOpt: Voxel-Adaptive Message Passing for Discrete Optimization in Deformable Abdominal CT Registration

Hang Zhang, Yuxi Zhang, Jiazheng Wang, Xiang Chen, Renjiu Hu, Xin Tian, Gaolei Li, Min Liu

Comments: Accepted for publication at MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)
[3001] arXiv:2506.20045 (cross-list from cs.RO) [pdf, html, other]: Title: Consensus-Driven Uncertainty for Robotic Grasping based on RGB Perception

Eric C. Joyce, Qianwen Zhao, Nathaniel Burgdorfer, Long Wang, Philippos Mordohai

Comments: Accepted to IROS 2025

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[3002] arXiv:2506.20100 (cross-list from cs.LG) [pdf, html, other]: Title: MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations

Vardhan Dongre, Chi Gui, Shubham Garg, Hooshang Nayyeri, Gokhan Tur, Dilek Hakkani-Tür, Vikram S. Adve

Comments: 66 pages, 32 figures, 23 tables

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[3003] arXiv:2506.20200 (cross-list from eess.IV) [pdf, html, other]: Title: MS-IQA: A Multi-Scale Feature Fusion Network for PET/CT Image Quality Assessment

Siqiao Li, Chen Hui, Wei Zhang, Rui Liang, Chenyue Song, Feng Jiang, Haiqi Zhu, Zhixuan Li, Hong Huang, Xiang Li

Comments: Accepted to MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3004] arXiv:2506.20245 (cross-list from cs.LG) [pdf, html, other]: Title: FedBKD: Distilled Federated Learning to Embrace Gerneralization and Personalization on Non-IID Data

Yushan Zhao, Jinyuan He, Donglai Chen, Weijie Luo, Chong Xie, Ri Zhang, Yonghong Chen, Yan Xu

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3005] arXiv:2506.20267 (cross-list from cs.GR) [pdf, html, other]: Title: X-SiT: Inherently Interpretable Surface Vision Transformers for Dementia Diagnosis

Fabian Bongratz, Tom Nuno Wolf, Jaume Gual Ramon, Christian Wachinger

Comments: MICCAI 2025

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3006] arXiv:2506.20282 (cross-list from eess.IV) [pdf, html, other]: Title: Opportunistic Osteoporosis Diagnosis via Texture-Preserving Self-Supervision, Mixture of Experts and Multi-Task Integration

Jiaxing Huang, Heng Guo, Le Lu, Fan Yang, Minfeng Xu, Ge Yang, Wei Luo

Comments: Accepted by MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3007] arXiv:2506.20303 (cross-list from eess.IV) [pdf, other]: Title: FundaQ-8: A Clinically-Inspired Scoring Framework for Automated Fundus Image Quality Assessment

Lee Qi Zun, Oscar Wong Jin Hao, Nor Anita Binti Che Omar, Zalifa Zakiah Binti Asnir, Mohamad Sabri bin Sinal Zainal, Goh Man Fye

Subjects: Image and Video Processing (eess.IV); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[3008] arXiv:2506.20305 (cross-list from cs.LG) [pdf, html, other]: Title: Learning Moderately Input-Sensitive Functions: A Case Study in QR Code Decoding

Kazuki Yoda, Kazuhiko Kawamoto, Hiroshi Kera

Comments: 17 pages, 13 figures

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3009] arXiv:2506.20333 (cross-list from eess.IV) [pdf, html, other]: Title: EAGLE: An Efficient Global Attention Lesion Segmentation Model for Hepatic Echinococcosis

Jiayan Chen, Kai Li, Yulu Zhao, Jianqiang Huang, Zhan Wang

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3010] arXiv:2506.20355 (cross-list from quant-ph) [pdf, html, other]: Title: Practical insights on the effect of different encodings, ansätze and measurements in quantum and hybrid convolutional neural networks

Jesús Lozano-Cruz, Albert Nieto-Morales, Oriol Balló-Gimbernat, Adan Garriga, Antón Rodríguez-Otero, Alejandro Borrallo-Rentero

Comments: 20 pages, 22 figures

Subjects: Quantum Physics (quant-ph); Computer Vision and Pattern Recognition (cs.CV)
[3011] arXiv:2506.20367 (cross-list from cs.GR) [pdf, html, other]: Title: DreamAnywhere: Object-Centric Panoramic 3D Scene Generation

Edoardo Alberto Dominici, Jozef Hladky, Floor Verhoeven, Lukas Radl, Thomas Deixelberger, Stefan Ainetter, Philipp Drescher, Stefan Hauswiesner, Arno Coomans, Giacomo Nazzaro, Konstantinos Vardis, Markus Steinberger

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[3012] arXiv:2506.20407 (cross-list from eess.IV) [pdf, html, other]: Title: Fusing Radiomic Features with Deep Representations for Gestational Age Estimation in Fetal Ultrasound Images

Fangyijie Wang, Yuan Liang, Sourav Bhattacharjee, Abey Campbell, Kathleen M. Curran, Guénolé Silvestre

Comments: Accepted at MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3013] arXiv:2506.20430 (cross-list from cs.CL) [pdf, html, other]: Title: An Agentic System for Rare Disease Diagnosis with Traceable Reasoning

Weike Zhao, Chaoyi Wu, Yanjie Fan, Xiaoman Zhang, Pengcheng Qiu, Yuze Sun, Xiao Zhou, Yanfeng Wang, Xin Sun, Ya Zhang, Yongguo Yu, Kun Sun, Weidi Xie

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA)
[3014] arXiv:2506.20566 (cross-list from cs.RO) [pdf, html, other]: Title: HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction

Zhonghao Shi, Enyu Zhao, Nathaniel Dennler, Jingzhen Wang, Xinyang Xu, Kaleen Shrestha, Mengxue Fu, Daniel Seita, Maja Matarić

Comments: Accepted to the 19th International Symposium on Experimental Robotics (ISER 2025)

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[3015] arXiv:2506.20614 (cross-list from eess.IV) [pdf, html, other]: Title: Weighted Mean Frequencies: a handcraft Fourier feature for 4D Flow MRI segmentation

Simon Perrin, Sébastien Levilly, Huajun Sun, Harold Mouchère, Jean-Michel Serfaty

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3016] arXiv:2506.20652 (cross-list from cs.GR) [pdf, html, other]: Title: EditP23: 3D Editing via Propagation of Image Prompts to Multi-View

Roi Bar-On, Dana Cohen-Bar, Daniel Cohen-Or

Comments: Code, supplementary videos, interactive 3D visualizations, and additional results are available at this https URL

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[3017] arXiv:2506.20683 (cross-list from eess.IV) [pdf, html, other]: Title: Global and Local Contrastive Learning for Joint Representations from Cardiac MRI and ECG

Alexander Selivanov, Philip Müller, Özgün Turgut, Nil Stolt-Ansó, Daniel Rückert

Comments: accepted to MICCAI 2025 (Springer LNCS)

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)
[3018] arXiv:2506.20689 (cross-list from eess.IV) [pdf, other]: Title: U-R-VEDA: Integrating UNET, Residual Links, Edge and Dual Attention, and Vision Transformer for Accurate Semantic Segmentation of CMRs

Racheal Mukisa, Arvind K. Bansal

Comments: 15 pages, 3 figures

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3019] arXiv:2506.20703 (cross-list from cs.GR) [pdf, html, other]: Title: Generative Blocks World: Moving Things Around in Pictures

Vaibhav Vavilala, Seemandhar Jain, Rahul Vasanth, D.A. Forsyth, Anand Bhattad

Comments: 23 pages, 16 figures, 2 tables

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[3020] arXiv:2506.20812 (cross-list from cs.RO) [pdf, html, other]: Title: Model-Based Real-Time Pose and Sag Estimation of Overhead Power Lines Using LiDAR for Drone Inspection

Alexandre Girard, Steven A. Parkison, Philippe Hamelin

Comments: Submitted to IEEE case 2025

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[3021] arXiv:2506.20816 (cross-list from cs.LG) [pdf, html, other]: Title: Universal and Efficient Detection of Adversarial Data through Nonuniform Impact on Network Layers

Furkan Mumcu, Yasin Yilmaz

Comments: arXiv admin note: substantial text overlap with arXiv:2410.17442

Journal-ref: Transactions on Machine Learning Research, June 2025

Subjects: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[3022] arXiv:2506.20875 (cross-list from cs.GR) [pdf, html, other]: Title: 3DGH: 3D Head Generation with Composable Hair and Face

Chengan He, Junxuan Li, Tobias Kirschstein, Artem Sevastopolsky, Shunsuke Saito, Qingyang Tan, Javier Romero, Chen Cao, Holly Rushmeier, Giljoo Nam

Comments: Accepted to SIGGRAPH 2025. Project page: this https URL

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[3023] arXiv:2506.20897 (cross-list from eess.IV) [pdf, html, other]: Title: Development of MR spectral analysis method robust against static magnetic field inhomogeneity

Shuki Maruyama, Hidenori Takeshima

Comments: 11 pages, 6 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3024] arXiv:2506.20946 (cross-list from cs.GR) [pdf, html, other]: Title: Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models

Donggoo Kang, Jangyeong Kim, Dasol Jeong, Junyoung Choi, Jeonga Wi, Hyunmin Lee, Joonho Gwon, Joonki Paik

Subjects: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3025] arXiv:2506.20969 (cross-list from cs.RO) [pdf, html, other]: Title: ThermalDiffusion: Visual-to-Thermal Image-to-Image Translation for Autonomous Navigation

Shruti Bansal, Wenshan Wang, Yifei Liu, Parv Maheshwari

Comments: Accepted at Thermal Infrared in Robotics (TIRO) Workshop, ICRA 2025

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[3026] arXiv:2506.20990 (cross-list from cs.LG) [pdf, html, other]: Title: SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes

Yifan Yang, Zhen Zhang, Rupak Vignesh Swaminathan, Jing Liu, Nathan Susanj, Zheng Zhang

Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[3027] arXiv:2506.21037 (cross-list from cs.LG) [pdf, html, other]: Title: RL-Selector: Reinforcement Learning-Guided Data Selection via Redundancy Assessment

Suorong Yang, Peijia Li, Furao Shen, Jian Zhao

Comments: ICCV 2025

Journal-ref: ICCV 2025

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3028] arXiv:2506.21041 (cross-list from cs.RO) [pdf, html, other]: Title: SEAL: Vision-Language Model-Based Safe End-to-End Cooperative Autonomous Driving with Adaptive Long-Tail Modeling

Junwei You, Pei Li, Zhuoyu Jiang, Zilin Huang, Rui Gan, Haotian Shi, Bin Ran

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3029] arXiv:2506.21144 (cross-list from cs.LG) [pdf, html, other]: Title: Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion

Yuguang Zhang, Kuangpu Guo, Zhihe Lu, Yunbo Wang, Jian Liang

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3030] arXiv:2506.21171 (cross-list from eess.IV) [pdf, other]: Title: Uncover Treasures in DCT: Advancing JPEG Quality Enhancement by Exploiting Latent Correlations

Jing Yang, Qunliang Xing, Mai Xu, Minglang Qiao

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3031] arXiv:2506.21245 (cross-list from eess.IV) [pdf, html, other]: Title: GANet-Seg: Adversarial Learning for Brain Tumor Segmentation with Hybrid Generative Models

Qifei Cui, Xinyu Lu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3032] arXiv:2506.21272 (cross-list from cs.GR) [pdf, html, other]: Title: FairyGen: Storied Cartoon Video from a Single Child-Drawn Character

Jiayi Zheng, Xiaodong Cun

Comments: Project Page: this https URL ; Code: this https URL

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[3033] arXiv:2506.21319 (cross-list from cs.HC) [pdf, html, other]: Title: SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding

Can Liu, Chunlin Da, Xiaoxiao Long, Yuxiao Yang, Yu Zhang, Yong Wang

Subjects: Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)
[3034] arXiv:2506.21331 (cross-list from cs.DL) [pdf, html, other]: Title: Automatic Reviewers Assignment to a Research Paper Based on Allied References and Publications Weight

Tamim Al Mahmud, B M Mainul Hossain, Dilshad Ara

Comments: IEEE Conference Proceedings (5 Pages)

Journal-ref: 2018 4th International Conference on Computing, Communication and Automation (ICCCA), Greater Noida, India, 2018, pp. 1-5

Subjects: Digital Libraries (cs.DL); Computer Vision and Pattern Recognition (cs.CV)
[3035] arXiv:2506.21448 (cross-list from eess.AS) [pdf, html, other]: Title: ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

Huadai Liu, Jialei Wang, Kaicheng Luo, Wen Wang, Qian Chen, Zhou Zhao, Wei Xue

Subjects: Audio and Speech Processing (eess.AS); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
[3036] arXiv:2506.21458 (cross-list from cs.AI) [pdf, other]: Title: Spatial Mental Modeling from Limited Views

Baiqiao Yin, Qineng Wang, Pingyue Zhang, Jianshu Zhang, Kangrui Wang, Zihan Wang, Jieyu Zhang, Keshigeyan Chandrasegaran, Han Liu, Ranjay Krishna, Saining Xie, Manling Li, Jiajun Wu, Li Fei-Fei

Comments: Preprint version

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[3037] arXiv:2506.21499 (cross-list from eess.IV) [pdf, html, other]: Title: Lightweight Physics-Informed Zero-Shot Ultrasound Plane Wave Denoising

Hojat Asgariandehkordi, Mostafa Sharifzadeh, Hassan Rivaz

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3038] arXiv:2506.21535 (cross-list from eess.IV) [pdf, html, other]: Title: Exploring the Design Space of 3D MLLMs for CT Report Generation

Mohammed Baharoon, Jun Ma, Congyu Fang, Augustin Toma, Bo Wang

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3039] arXiv:2506.21537 (cross-list from quant-ph) [pdf, html, other]: Title: ResQ: A Novel Framework to Implement Residual Neural Networks on Analog Rydberg Atom Quantum Computers

Nicholas S. DiBrita, Jason Han, Tirthak Patel

Comments: ResQ will appear in the Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2025

Subjects: Quantum Physics (quant-ph); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET)
[3040] arXiv:2506.21586 (cross-list from cs.CL) [pdf, html, other]: Title: Can Vision Language Models Understand Mimed Actions?

Hyundong Cho, Spencer Lin, Tejas Srinivasan, Michael Saxon, Deuksin Kwon, Natali T. Chavez, Jonathan May

Comments: ACL 2025 Findings

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3041] arXiv:2506.21592 (cross-list from cs.CL) [pdf, html, other]: Title: SignBart -- New approach with the skeleton sequence for Isolated Sign language Recognition

Tinh Nguyen, Minh Khue Phan Tran

Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[3042] arXiv:2506.21601 (cross-list from cs.IR) [pdf, html, other]: Title: Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization

Duong Bach

Comments: 9 pages

Subjects: Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV)
[3043] arXiv:2506.21604 (cross-list from cs.IR) [pdf, html, other]: Title: Evaluating VisualRAG: Quantifying Cross-Modal Performance in Enterprise Document Understanding

Varun Mannam, Fang Wang, Xin Chen

Comments: Conference: KDD conference workshop: this https URL

Subjects: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
[3044] arXiv:2506.21629 (cross-list from cs.GR) [pdf, html, other]: Title: ICP-3DGS: SfM-free 3D Gaussian Splatting for Large-scale Unbounded Scenes

Chenhao Zhang, Yezhi Shen, Fengqing Zhu

Comments: 6 pages, Source code is available at this https URL. To appear at ICIP 2025

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[3045] arXiv:2506.21630 (cross-list from cs.RO) [pdf, html, other]: Title: TOMD: A Trail-based Off-road Multimodal Dataset for Traversable Pathway Segmentation under Challenging Illumination Conditions

Yixin Sun, Li Li, Wenke E, Amir Atapour-Abarghouei, Toby P. Breckon

Comments: 8 pages, 9 figures, 2025 IJCNN

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3046] arXiv:2506.21635 (cross-list from cs.RO) [pdf, html, other]: Title: AeroLite-MDNet: Lightweight Multi-task Deviation Detection Network for UAV Landing

Haiping Yang, Huaxing Liu, Wei Wu, Zuohui Chen, Ning Wu

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3047] arXiv:2506.21655 (cross-list from cs.LG) [pdf, html, other]: Title: APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization

Minjie Hong, Zirun Guo, Yan Xia, Zehan Wang, Ziang Zhang, Tao Jin, Zhou Zhao

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3048] arXiv:2506.21680 (cross-list from eess.IV) [pdf, html, other]: Title: PhotonSplat: 3D Scene Reconstruction and Colorization from SPAD Sensors

Sai Sri Teja, Sreevidya Chintalapati, Vinayak Gupta, Mukund Varma T, Haejoon Lee, Aswin Sankaranarayanan, Kaushik Mitra

Comments: Accepted at the International Conference on Computational Photography(ICCP) 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3049] arXiv:2506.21714 (cross-list from cs.LG) [pdf, html, other]: Title: ODE$_t$(ODE$_l$): Shortcutting the Time and Length in Diffusion and Flow Models for Faster Sampling

Denis Gudovskiy, Wenzhao Zheng, Tomoyuki Okuno, Yohei Nakata, Kurt Keutzer

Comments: Preprint. Github page: this http URL

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3050] arXiv:2506.21732 (cross-list from cs.RO) [pdf, html, other]: Title: Experimental investigation of pose informed reinforcement learning for skid-steered visual navigation

Ameya Salvi, Venkat Krovi

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Systems and Control (eess.SY)
[3051] arXiv:2506.21748 (cross-list from physics.optics) [pdf, html, other]: Title: Inverse Design of Diffractive Metasurfaces Using Diffusion Models

Liav Hen, Erez Yosef, Dan Raviv, Raja Giryes, Jacob Scheuer

Subjects: Optics (physics.optics); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3052] arXiv:2506.21765 (cross-list from eess.IV) [pdf, html, other]: Title: TUS-REC2024: A Challenge to Reconstruct 3D Freehand Ultrasound Without External Tracker

Qi Li, Shaheer U. Saeed, Yuliang Huang, Mingyuan Luo, Zhongnuo Yan, Jiongquan Chen, Xin Yang, Dong Ni, Nektarios Winter, Phuc Nguyen, Lucas Steinberger, Caelan Haney, Yuan Zhao, Mingjie Jiang, Bowen Ren, SiYeoul Lee, Seonho Kim, MinKyung Seo, MinWoo Kim, Yimeng Dou, Zhiwei Zhang, Yin Li, Tomy Varghese, Dean C. Barratt, Matthew J. Clarkson, Tom Vercauteren, Yipeng Hu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3053] arXiv:2506.21812 (cross-list from cs.CL) [pdf, html, other]: Title: Towards Transparent AI: A Survey on Explainable Large Language Models

Avash Palikhe, Zhenyu Yu, Zichong Wang, Wenbin Zhang

Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[3054] arXiv:2506.21860 (cross-list from cs.RO) [pdf, html, other]: Title: Embodied Domain Adaptation for Object Detection

Xiangyu Shi, Yanyuan Qiao, Lingqiao Liu, Feras Dayoub

Comments: Accepted by IROS 2025

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[3055] arXiv:2506.21876 (cross-list from cs.CL) [pdf, html, other]: Title: Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

Qiyue Gao, Xinyu Pi, Kevin Liu, Junrong Chen, Ruolan Yang, Xinqi Huang, Xinyu Fang, Lu Sun, Gautham Kishore, Bo Ai, Stone Tao, Mengyang Liu, Jiaxi Yang, Chao-Jung Lai, Chuanyang Jin, Jiannan Xiang, Benhao Huang, Zeming Chen, David Danks, Hao Su, Tianmin Shu, Ziqiao Ma, Lianhui Qin, Zhiting Hu

Comments: ACL 2025 (Findings)

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3056] arXiv:2506.21880 (cross-list from eess.IV) [pdf, html, other]: Title: Physical Degradation Model-Guided Interferometric Hyperspectral Reconstruction with Unfolding Transformer

Yuansheng Li, Yunhao Zou, Linwei Chen, Ying Fu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3057] arXiv:2506.21884 (cross-list from eess.IV) [pdf, html, other]: Title: UnMix-NeRF: Spectral Unmixing Meets Neural Radiance Fields

Fabian Perez, Sara Rojas, Carlos Hinojosa, Hoover Rueda-Chacón, Bernard Ghanem

Comments: Paper accepted at ICCV 2025 main conference

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Signal Processing (eess.SP)
[3058] arXiv:2506.21934 (cross-list from cs.IR) [pdf, html, other]: Title: CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design

Najmeh Forouzandehmehr, Reza Yousefi Maragheh, Sriram Kollipara, Kai Zhao, Topojoy Biswas, Evren Korpeoglu, Kannan Achan

Subjects: Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV)
[3059] arXiv:2506.21976 (cross-list from cs.LG) [pdf, html, other]: Title: SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model

Shuhan Tan, John Lambert, Hong Jeon, Sakshum Kulshrestha, Yijing Bai, Jing Luo, Dragomir Anguelov, Mingxing Tan, Chiyu Max Jiang

Comments: Accepted to CVPR 2025

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multiagent Systems (cs.MA); Robotics (cs.RO)
[3060] arXiv:2506.21977 (cross-list from eess.IV) [pdf, other]: Title: StableCodec: Taming One-Step Diffusion for Extreme Image Compression

Tianyu Zhang, Xin Luo, Li Li, Dong Liu

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3061] arXiv:2506.22012 (cross-list from eess.IV) [pdf, html, other]: Title: Noise-Inspired Diffusion Model for Generalizable Low-Dose CT Reconstruction

Qi Gao, Zhihao Chen, Dong Zeng, Junping Zhang, Jianhua Ma, Hongming Shan

Comments: Accepted for publication in Medical Image Analysis, 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3062] arXiv:2506.22041 (cross-list from eess.IV) [pdf, html, other]: Title: Towards Scalable and Robust White Matter Lesion Localization via Multimodal Deep Learning

Julia Machnio, Sebastian Nørgaard Llambias, Mads Nielsen, Mostafa Mehdipour Ghazi

Comments: 2nd Sorbonne-Heidelberg Workshop on AI in medicine: Machine Learning for multi-modal data

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3063] arXiv:2506.22116 (cross-list from cs.RO) [pdf, html, other]: Title: Evaluating Pointing Gestures for Target Selection in Human-Robot Collaboration

Noora Sassali, Roel Pieters

Comments: Accepted by the 2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). Preprint

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[3064] arXiv:2506.22156 (cross-list from cs.AR) [pdf, html, other]: Title: Hardware acceleration for ultra-fast Neural Network training on FPGA for MRF map reconstruction

Mattia Ricchi, Fabrizio Alfonsi, Camilla Marella, Marco Barbieri, Alessandra Retico, Leonardo Brizi, Alessandro Gabrielli, Claudia Testa

Comments: 8 pages, 2 figures, to be published in conference proceedings of SDPS 2024: 2024 International Conference of the Society for Design and Process Science on Advances and Challenges of Applying AI/GenAI in Design and Process Science

Subjects: Hardware Architecture (cs.AR); Computer Vision and Pattern Recognition (cs.CV); Instrumentation and Detectors (physics.ins-det)
[3065] arXiv:2506.22176 (cross-list from cs.RO) [pdf, html, other]: Title: KnotDLO: Toward Interpretable Knot Tying

Holly Dinkel, Raghavendra Navaratna, Jingyi Xiang, Brian Coltin, Trey Smith, Timothy Bretl

Comments: 4 pages, 5 figures, presented at the Workshop on 3D Visual Representations for Manipulation at the 2023 IEEE International Conference on Robotics and Automation in Yokohama, Japan. Video presentation [this https URL]. Poster [this https URL] 3DVRM Workshop [this https URL]

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[3066] arXiv:2506.22222 (cross-list from eess.IV) [pdf, html, other]: Title: Advanced Deep Learning Techniques for Automated Segmentation of Type B Aortic Dissections

Hao Xu, Ruth Lim, Brian E. Chapman

Comments: 9 pages, 5 figures, 3 tables

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3067] arXiv:2506.22226 (cross-list from eess.IV) [pdf, html, other]: Title: Cardiovascular disease classification using radiomics and geometric features from cardiac CT

Ajay Mittal, Raghav Mehta, Omar Todd, Philipp Seeböck, Georg Langs, Ben Glocker

Comments: Under Review at STACOM 2025 with MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3068] arXiv:2506.22280 (cross-list from eess.IV) [pdf, html, other]: Title: DIGS: Dynamic CBCT Reconstruction using Deformation-Informed 4D Gaussian Splatting and a Low-Rank Free-Form Deformation Model

Yuliang Huang, Imraj Singh, Thomas Joyce, Kris Thielemans, Jamie R. McClelland

Comments: Accepted by MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3069] arXiv:2506.22304 (cross-list from cs.LG) [pdf, html, other]: Title: Unfolding Generative Flows with Koopman Operators: Fast and Interpretable Sampling

Erkan Turan, Aristotelis Siozopoulos, Maks Ovsjanikov

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3070] arXiv:2506.22340 (cross-list from quant-ph) [pdf, html, other]: Title: QuKAN: A Quantum Circuit Born Machine approach to Quantum Kolmogorov Arnold Networks

Yannick Werner, Akash Malemath, Mengxi Liu, Vitor Fortes Rey, Nikolaos Palaiodimopoulos, Paul Lukowicz, Maximilian Kiefer-Emmanouilidis

Subjects: Quantum Physics (quant-ph); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3071] arXiv:2506.22397 (cross-list from eess.IV) [pdf, other]: Title: Dehazing Light Microscopy Images with Guided Conditional Flow Matching: finding a sweet spot between fidelity and realism

Anirban Ray, Ashesh, Florian Jug

Comments: 4 figures, 10 pages + refs, 40 pages total (including supplement), 24 supplementary figures

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3072] arXiv:2506.22426 (cross-list from eess.IV) [pdf, html, other]: Title: Single-shot HDR using conventional image sensor shutter functions and optical randomization

Xiang Dai, Kyrollos Yanny, Kristina Monakhova, Nicholas Antipa

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Signal Processing (eess.SP); Optics (physics.optics)
[3073] arXiv:2506.22467 (cross-list from eess.SP) [pdf, other]: Title: SegmentAnyMuscle: A universal muscle segmentation model across different locations in MRI

Roy Colglazier, Jisoo Lee, Haoyu Dong, Hanxue Gu, Yaqian Chen, Joseph Cao, Zafer Yildiz, Zhonghao Liu, Nicholas Konz, Jichen Yang, Jikai Zhang, Yuwen Chen, Lin Li, Adrian Camarena, Maciej A. Mazurowski

Comments: 24 pages, 6 figures

Subjects: Signal Processing (eess.SP); Computer Vision and Pattern Recognition (cs.CV)
[3074] arXiv:2506.22482 (cross-list from cs.NI) [pdf, other]: Title: Wireless Home Automation Using Social Networking Websites

Divya Alok Gupta, Dwith Chenna, B. Aditya Vighnesh Ramakanth

Comments: 20th Annual International Conference on Advanced Computing and Communications (ADCOM) 2014

Subjects: Networking and Internet Architecture (cs.NI); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[3075] arXiv:2506.22494 (cross-list from cs.RO) [pdf, html, other]: Title: DriveBLIP2: Attention-Guided Explanation Generation for Complex Driving Scenarios

Shihong Ling, Yue Wan, Xiaowei Jia, Na Du

Comments: Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025. 7 pages, 3 figures

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3076] arXiv:2506.22532 (cross-list from eess.IV) [pdf, other]: Title: High Resolution Isotropic 3D Cine imaging with Automated Segmentation using Concatenated 2D Real-time Imaging and Deep Learning

Mark Wrobel (1), Michele Pascale (1), Tina Yao (1), Ruaraidh Campbell (1), Elena Milano (2), Michael Quail (1 and 2), Jennifer Steeden (1), Vivek Muthurangu (1) ((1) UCL Centre for Translational Cardiovascular Imaging, University College London, (2) Great Ormond Street Hospital)

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3077] arXiv:2506.22568 (cross-list from math.OC) [pdf, html, other]: Title: Maximum Dispersion, Maximum Concentration: Enhancing the Quality of MOP Solutions

Gladston Moreira, Ivan Meneghini, Elizabeth Wanner

Comments: 11 pages

Subjects: Optimization and Control (math.OC); Computer Vision and Pattern Recognition (cs.CV)
[3078] arXiv:2506.22580 (cross-list from eess.IV) [pdf, html, other]: Title: FedCLAM: Client Adaptive Momentum with Foreground Intensity Matching for Federated Medical Image Segmentation

Vasilis Siomos, Jonathan Passerat-Palmbach, Giacomo Tarroni

Comments: 10 pages, 2 figures, Accepted at MICCAI 2025

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3079] arXiv:2506.22593 (cross-list from cs.RO) [pdf, html, other]: Title: Pixels-to-Graph: Real-time Integration of Building Information Models and Scene Graphs for Semantic-Geometric Human-Robot Understanding

Antonello Longo, Chanyoung Chung, Matteo Palieri, Sung-Kyun Kim, Ali Agha, Cataldo Guaragnella, Shehryar Khattak

Comments: Paper accepted to 2025 IEEE International Conference on Automation Science and Engineering (CASE)

Subjects: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3080] arXiv:2506.22706 (cross-list from cs.CR) [pdf, other]: Title: General Autonomous Cybersecurity Defense: Learning Robust Policies for Dynamic Topologies and Diverse Attackers

Arun Ramamurthy, Neil Dhir

Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
[3081] arXiv:2506.22790 (cross-list from eess.IV) [pdf, html, other]: Title: ICME 2025 Generalizable HDR and SDR Video Quality Measurement Grand Challenge

Yixu Chen, Bowen Chen, Hai Wei, Alan C. Bovik, Baojun Li, Wei Sun, Linhan Cao, Kang Fu, Dandan Zhu, Jun Jia, Menghan Hu, Xiongkuo Min, Guangtao Zhai, Dounia Hammou, Fei Yin, Rafal Mantiuk, Amritha Premkumar, Prajit T Rajendran, Vignesh V Menon

Comments: ICME 2025 Grand Challenges

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
[3082] arXiv:2506.22799 (cross-list from cs.GR) [pdf, html, other]: Title: VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding

Minchao Jiang, Shunyu Jia, Jiaming Gu, Xiaoyuan Lu, Guangming Zhu, Anqi Dong, Liang Zhang

Comments: Accepted to ICCV 2025

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3083] arXiv:2506.22802 (cross-list from cs.LG) [pdf, html, other]: Title: Riemannian-Geometric Fingerprints of Generative Models

Hae Jin Song, Laurent Itti

Subjects: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[3084] arXiv:2506.22826 (cross-list from math.OC) [pdf, html, other]: Title: Denoising Multi-Color QR Codes and Stiefel-Valued Data by Relaxed Regularizations

Robert Beinert, Jonas Bresch

Comments: 9 pages, 2 figures, 3 algorithms

Subjects: Optimization and Control (math.OC); Computer Vision and Pattern Recognition (cs.CV); Numerical Analysis (math.NA)
[3085] arXiv:2506.22882 (cross-list from eess.IV) [pdf, html, other]: Title: CA-Diff: Collaborative Anatomy Diffusion for Brain Tissue Segmentation

Qilong Xing, Zikai Song, Yuteng Ye, Yuke Chen, Youjia Zhang, Na Feng, Junqing Yu, Wei Yang

Comments: ICME 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3086] arXiv:2506.22952 (cross-list from eess.IV) [pdf, html, other]: Title: Hierarchical Characterization of Brain Dynamics via State Space-based Vector Quantization

Yanwu Yang, Thomas Wolfers

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Neurons and Cognition (q-bio.NC)
[3087] arXiv:2506.22973 (cross-list from cs.GR) [pdf, html, other]: Title: Confident Splatting: Confidence-Based Compression of 3D Gaussian Splatting via Learnable Beta Distributions

AmirHossein Naghi Razlighi, Elaheh Badali Golezani, Shohreh Kasaei

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[3088] arXiv:2506.22992 (cross-list from cs.AI) [pdf, html, other]: Title: MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning

Yulun Jiang, Yekun Chai, Maria Brbić, Michael Moor

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[3089] arXiv:2506.23016 (cross-list from cs.HC) [pdf, html, other]: Title: Deep Learning in Mild Cognitive Impairment Diagnosis using Eye Movements and Image Content in Visual Memory Tasks

Tomás Silva Santos Rocha, Anastasiia Mikhailova, Moreno I. Coco, José Santos-Victor

Comments: 13 pages, 5 figures

Subjects: Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)
[3090] arXiv:2506.23041 (cross-list from cs.LG) [pdf, html, other]: Title: ReMem: Mutual Information-Aware Fine-tuning of Pretrained Vision Transformers for Effective Knowledge Distillation

Chengyu Dong, Huan Gui, Noveen Sachdeva, Long Jin, Ke Yin, Jingbo Shang, Lichan Hong, Ed H.Chi, Zhe Zhao

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3091] arXiv:2506.23046 (cross-list from cs.CL) [pdf, html, other]: Title: SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions

Xianzhe Fan, Xuhui Zhou, Chuanyang Jin, Kolby Nottingham, Hao Zhu, Maarten Sap

Comments: 23 pages, 6 figures

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
[3092] arXiv:2506.23102 (cross-list from eess.IV) [pdf, html, other]: Title: MedRegion-CT: Region-Focused Multimodal LLM for Comprehensive 3D CT Report Generation

Sunggu Kyung, Jinyoung Seo, Hyunseok Lim, Dongyeong Kim, Hyungbin Park, Jimin Sung, Jihyun Kim, Wooyoung Jo, Yoojin Nam, Namkug Kim

Comments: 14 pages, 5 figures, submitted to ICCV 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3093] arXiv:2506.23121 (cross-list from eess.IV) [pdf, html, other]: Title: CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation

Xinlei Yu, Changmiao Wang, Hui Jin, Ahmed Elazab, Gangyong Jia, Xiang Wan, Changqing Zou, Ruiquan Ge

Comments: Accepted By ACMMM25

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3094] arXiv:2506.23145 (cross-list from cs.LG) [pdf, html, other]: Title: Forget-MI: Machine Unlearning for Forgetting Multimodal Information in Healthcare Settings

Shahad Hardan, Darya Taratynova, Abdelmajid Essofi, Karthik Nandakumar, Mohammad Yaqub

Subjects: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
[3095] arXiv:2506.23147 (cross-list from cs.LG) [pdf, html, other]: Title: maneuverRecognition -- A Python package for Timeseries Classification in the domain of Vehicle Telematics

Jonathan Schuster, Fabian Transchel

Comments: 6 pages, 2 figures

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3096] arXiv:2506.23184 (cross-list from eess.IV) [pdf, html, other]: Title: Score-based Diffusion Model for Unpaired Virtual Histology Staining

Anran Liu, Xiaofei Wang, Jing Cai, Chao Li

Comments: 11 pages, 3 figures

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3097] arXiv:2506.23208 (cross-list from eess.IV) [pdf, html, other]: Title: Multi-Source COVID-19 Detection via Variance Risk Extrapolation

Runtian Yuan, Qingqiu Li, Junlin Hou, Jilan Xu, Yuejie Zhang, Rui Feng, Hao Chen

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3098] arXiv:2506.23221 (cross-list from cs.LG) [pdf, html, other]: Title: Single Image Inpainting and Super-Resolution with Simultaneous Uncertainty Guarantees by Universal Reproducing Kernels

Bálint Horváth, Balázs Csanád Csáji

Comments: 23 pages, 8 figures, 6 tables

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3099] arXiv:2506.23259 (cross-list from eess.IV) [pdf, html, other]: Title: Improving Myocardial Infarction Detection via Synthetic ECG Pretraining

Lachin Naghashyar

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3100] arXiv:2506.23298 (cross-list from eess.IV) [pdf, html, other]: Title: Exposing and Mitigating Calibration Biases and Demographic Unfairness in MLLM Few-Shot In-Context Learning for Medical Image Classification

Xing Shen, Justin Szeto, Mingyang Li, Hengguan Huang, Tal Arbel

Comments: Preprint version. The peer-reviewed version of this paper has been accepted to MICCAI 2025 main conference

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3101] arXiv:2506.23305 (cross-list from eess.IV) [pdf, html, other]: Title: BPD-Neo: An MRI Dataset for Lung-Trachea Segmentation with Clinical Data for Neonatal Bronchopulmonary Dysplasia

Rachit Saluja, Arzu Kovanlikaya, Candace Chien, Lauren Kathryn Blatt, Jeffrey M. Perlman, Stefan Worgall, Mert R. Sabuncu, Jonathan P. Dyke

Comments: Adding link to Zenodo repo for dataset

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3102] arXiv:2506.23309 (cross-list from eess.IV) [pdf, html, other]: Title: SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting

Yiming Huang, Long Bai, Beilei Cui, Kun Yuan, Guankun Wang, Mobarak I. Hoque, Nicolas Padoy, Nassir Navab, Hongliang Ren

Comments: MICCAI 2025. Project Page: this https URL

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3103] arXiv:2506.23316 (cross-list from cs.RO) [pdf, html, other]: Title: InfGen: Scenario Generation as Next Token Group Prediction

Zhenghao Peng, Yuxin Liu, Bolei Zhou

Subjects: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
[3104] arXiv:2506.23334 (cross-list from eess.IV) [pdf, html, other]: Title: Federated Breast Cancer Detection Enhanced by Synthetic Ultrasound Image Augmentation

Hongyi Pan, Ziliang Hong, Gorkem Durak, Ziyue Xu, Ulas Bagci

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3105] arXiv:2506.23466 (cross-list from eess.IV) [pdf, other]: Title: FD-DiT: Frequency Domain-Directed Diffusion Transformer for Low-Dose CT Reconstruction

Qiqing Liu, Guoquan Wei, Zekun Zhou, Yiyang Wen, Liu Shi, Qiegen Liu

Comments: 11pages, 11 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Medical Physics (physics.med-ph)
[3106] arXiv:2506.23471 (cross-list from cs.IR) [pdf, html, other]: Title: KiseKloset: Comprehensive System For Outfit Retrieval, Recommendation, And Try-On

Thanh-Tung Phan-Nguyen, Khoi-Nguyen Nguyen-Ngoc, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

Subjects: Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV)
[3107] arXiv:2506.23484 (cross-list from cs.MM) [pdf, html, other]: Title: TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion Sensitivity

Yuzhuo Chen, Zehua Ma, Han Fang, Weiming Zhang, Nenghai Yu

Comments: Camera-ready version for ICCV 2025. Adds GitHub link; acknowledgments; appendix. Abstract and Figure 1 updated for clarity

Subjects: Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
[3108] arXiv:2506.23490 (cross-list from eess.IV) [pdf, html, other]: Title: UltraTwin: Towards Cardiac Anatomical Twin Generation from Multi-view 2D Ultrasound

Junxuan Yu, Yaofei Duan, Yuhao Huang, Yu Wang, Rongbo Ling, Weihao Luo, Ang Zhang, Jingxian Xu, Qiongying Ni, Yongsong Zhou, Binghan Li, Haoran Dou, Liping Liu, Yanfen Chu, Feng Geng, Zhe Sheng, Zhifeng Ding, Dingxin Zhang, Rui Huang, Yuhang Zhang, Xiaowei Xu, Tao Tan, Dong Ni, Zhongshan Gou, Xin Yang

Comments: accepted by miccai 2025

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3109] arXiv:2506.23492 (cross-list from cs.LG) [pdf, html, other]: Title: Sample Margin-Aware Recalibration of Temperature Scaling

Haolan Guo, Linwei Tao, Haoyang Luo, Minjing Dong, Chang Xu

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3110] arXiv:2506.23506 (cross-list from eess.IV) [pdf, other]: Title: Artificial Intelligence-assisted Pixel-level Lung (APL) Scoring for Fast and Accurate Quantification in Ultra-short Echo-time MRI

Bowen Xin, Rohan Hickey, Tamara Blake, Jin Jin, Claire E Wainwright, Thomas Benkert, Alto Stemmer, Peter Sly, David Coman, Jason Dowling

Comments: Oral presentation in ISMRM2025

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Medical Physics (physics.med-ph)
[3111] arXiv:2506.23516 (cross-list from cs.LG) [pdf, html, other]: Title: FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization

Seung-Wook Kim, Seongyeol Kim, Jiah Kim, Seowon Ji, Se-Ho Lee

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3112] arXiv:2506.23537 (cross-list from eess.IV) [pdf, html, other]: Title: AFUNet: Cross-Iterative Alignment-Fusion Synergy for HDR Reconstruction via Deep Unfolding Paradigm

Xinyue Li, Zhangkai Ni, Wenhan Yang

Comments: Accepted to International Conference on Computer Vision (ICCV) 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3113] arXiv:2506.23563 (cross-list from cs.AI) [pdf, html, other]: Title: MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

Huanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xikun Zhang, Jingyi Zhang, Yuxin Song, Wenhao Wu, Dacheng Tao

Comments: Technical report

Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
[3114] arXiv:2506.23584 (cross-list from eess.IV) [pdf, html, other]: Title: A Clinically-Grounded Two-Stage Framework for Renal CT Report Generation

Renjie Liang, Zhengkang Fan, Jinqian Pan, Chenkun Sun, Russell Terry, Jie Xu

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3115] arXiv:2506.23664 (cross-list from eess.IV) [pdf, html, other]: Title: Diffusion Model-based Data Augmentation Method for Fetal Head Ultrasound Segmentation

Fangyijie Wang, Kevin Whelan, Félix Balado, Kathleen M. Curran, Guénolé Silvestre

Comments: Accepted at Irish Machine Vision and Image Processing Conference (IMVIP) 2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3116] arXiv:2506.23700 (cross-list from eess.IV) [pdf, html, other]: Title: MedSAM-CA: A CNN-Augmented ViT with Attention-Enhanced Multi-Scale Fusion for Medical Image Segmentation

Peiting Tian, Xi Chen, Haixia Bi, Fan Li

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3117] arXiv:2506.23701 (cross-list from eess.IV) [pdf, html, other]: Title: MDPG: Multi-domain Diffusion Prior Guidance for MRI Reconstruction

Lingtong Zhang, Mengdie Song, Xiaohan Hao, Huayu Mai, Bensheng Qiu

Comments: Accept by MICCAI2025

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3118] arXiv:2506.23717 (cross-list from cs.NE) [pdf, html, other]: Title: Towards Efficient and Accurate Spiking Neural Networks via Adaptive Bit Allocation

Xingting Yao, Qinghao Hu, Fei Zhou, Tielong Liu, Gang Li, Peisong Wang, Jian Cheng

Subjects: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3119] arXiv:2506.23721 (cross-list from eess.IV) [pdf, html, other]: Title: Deep Learning-Based Semantic Segmentation for Real-Time Kidney Imaging and Measurements with Augmented Reality-Assisted Ultrasound

Gijs Luijten, Roberto Maria Scardigno, Lisle Faray de Paiva, Peter Hoyer, Jens Kleesiek, Domenico Buongiorno, Vitoantonio Bevilacqua, Jan Egger

Subjects: Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
[3120] arXiv:2506.23731 (cross-list from cs.LG) [pdf, html, other]: Title: Radioactive Watermarks in Diffusion and Autoregressive Image Generative Models

Michel Meintz, Jan Dubiński, Franziska Boenisch, Adam Dziedzic

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3121] arXiv:2506.23759 (cross-list from eess.IV) [pdf, html, other]: Title: Spatio-Temporal Representation Decoupling and Enhancement for Federated Instrument Segmentation in Surgical Videos

Zheng Fang, Xiaoming Qi, Chun-Mei Feng, Jialun Pei, Weixin Si, Yueming Jin

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3122] arXiv:2506.23824 (cross-list from cs.LG) [pdf, html, other]: Title: Supercm: Revisiting Clustering for Semi-Supervised Learning

Durgesh Singh, Ahcene Boubekki, Robert Jenssen, Michael C. Kampffmeyer

Journal-ref: 10.1109/ICASSP49357.2023.10095856

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3123] arXiv:2506.23957 (cross-list from cs.GR) [pdf, html, other]: Title: GaVS: 3D-Grounded Video Stabilization via Temporally-Consistent Local Reconstruction and Rendering

Zinuo You, Stamatios Georgoulis, Anpei Chen, Siyu Tang, Dengxin Dai

Comments: siggraph 2025, project website: this https URL. version 2, update discussion

Subjects: Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
[3124] arXiv:2506.24000 (cross-list from cs.LG) [pdf, html, other]: Title: The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models

Lijun Sheng, Jian Liang, Ran He, Zilei Wang, Tieniu Tan

Comments: Github link: this https URL

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
[3125] arXiv:2506.24003 (cross-list from eess.IV) [pdf, html, other]: Title: ShapeKit

Junqi Liu, Dongli He, Wenxuan Li, Ningyu Wang, Alan L. Yuille, Zongwei Zhou

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3126] arXiv:2506.24016 (cross-list from cs.CL) [pdf, html, other]: Title: EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations

Hyunjong Kim, Sangyeop Kim, Jongheon Jeong, Yeongjae Cho, Sungzoon Cho

Comments: Accepted at ACL 2025 Findings

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
[3127] arXiv:2506.24034 (cross-list from physics.med-ph) [pdf, html, other]: Title: Supervised Diffusion-Model-Based PET Image Reconstruction

George Webber, Alexander Hammers, Andrew P King, Andrew J Reader

Comments: 12 pages, 6 figures. Submitted to MICCAI 2025, not peer-reviewed

Subjects: Medical Physics (physics.med-ph); Computer Vision and Pattern Recognition (cs.CV)
[3128] arXiv:2506.24074 (cross-list from eess.IV) [pdf, html, other]: Title: C3VDv2 -- Colonoscopy 3D video dataset with enhanced realism

Mayank V. Golhar, Lucas Sebastian Galeano Fretes, Loren Ayers, Venkata S. Akshintala, Taylor L. Bobrow, Nicholas J. Durr

Comments: 19 pages, 7 figures

Subjects: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
[3129] arXiv:2506.24108 (cross-list from cs.GR) [pdf, html, other]: Title: Navigating with Annealing Guidance Scale in Diffusion Space

Shai Yehezkel, Omer Dahary, Andrey Voynov, Daniel Cohen-Or

Comments: Project page: this https URL

Subjects: Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
[3130] arXiv:2506.24124 (cross-list from cs.LG) [pdf, html, other]: Title: Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives

Sixun Dong, Wei Fan, Teresa Wu, Yanjie Fu

Comments: Code: this https URL

Subjects: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

Total of 3130 entries : 1-2000 2001-3130 2701-3130

Showing up to 2000 entries per page: fewer | more | all