-
Measurement of the time-integrated $CP$ asymmetry in $D^0 \to K^0_{\rm S} K^0_{\rm S}$ decays using opposite-side flavor tagging at Belle and Belle II
Authors:
Belle,
Belle II Collaborations,
:,
I. Adachi,
Y. Ahn,
N. Akopov,
S. Alghamdi,
M. Alhakami,
A. Aloisio,
N. Althubiti,
K. Amos,
M. Angelsmark,
N. Anh Ky,
C. Antonioli,
D. M. Asner,
H. Atmacan,
T. Aushev,
M. Aversano,
R. Ayad,
V. Babu,
H. Bae,
N. K. Baghel,
S. Bahinipati,
P. Bambade,
Sw. Banerjee
, et al. (356 additional authors not shown)
Abstract:
We measure the time-integrated $CP$ asymmetry in $D^0 \to K^0_{\rm S} K^0_{\rm S}$ decays reconstructed in $e^+e^-\to c{\overline c}$ events collected by the Belle and Belle II experiments. The corresponding data samples have integrated luminosities of 980 and 428 fb${}^{-1}$, respectively. To infer the flavor of the $D^0$ meson, we exploit the correlation between the flavor of the reconstructed d…
▽ More
We measure the time-integrated $CP$ asymmetry in $D^0 \to K^0_{\rm S} K^0_{\rm S}$ decays reconstructed in $e^+e^-\to c{\overline c}$ events collected by the Belle and Belle II experiments. The corresponding data samples have integrated luminosities of 980 and 428 fb${}^{-1}$, respectively. To infer the flavor of the $D^0$ meson, we exploit the correlation between the flavor of the reconstructed decay and the electric charges of particles reconstructed in the rest of the $e^+e^-\to c{\overline c}$ event. This results in a sample which is independent from any other previously used at Belle or Belle II. The result, $A_{CP}(D^0 \to K^0_{\rm S} K^0_{\rm S}) = (1.3 \pm 2.0 \pm 0.2)\%$, where the first uncertainty is statistical and the second systematic, is consistent with previous determinations and with $CP$ symmetry.
△ Less
Submitted 22 April, 2025;
originally announced April 2025.
-
Search for lepton-flavor-violating $τ^- \to \ell^- K_s^0$ decays at Belle and Belle II
Authors:
Belle,
Belle II Collaborations,
:,
I. Adachi,
L. Aggarwal,
H. Ahmed,
Y. Ahn,
H. Aihara,
N. Akopov,
S. Alghamdi,
M. Alhakami,
A. Aloisio,
N. Althubiti,
K. Amos,
M. Angelsmark,
N. Anh Ky,
C. Antonioli,
D. M. Asner,
H. Atmacan,
V. Aushev,
M. Aversano,
R. Ayad,
V. Babu,
N. K. Baghel,
S. Bahinipati
, et al. (397 additional authors not shown)
Abstract:
We present the results of a search for charged-lepton-flavor violating decays $τ^{-} \rightarrow \ell^{-}K_{S}^{0}$, where $\ell^{-}$ is either an electron or a muon. We combine $e^+e^-$ data samples recorded by the Belle II experiment at the SuperKEKB collider (428 fb$^{-1}$) with samples recorded by the Belle experiment at the KEKB collider (980 fb$^{-1}$) to obtain a sample of 1.3 billion…
▽ More
We present the results of a search for charged-lepton-flavor violating decays $τ^{-} \rightarrow \ell^{-}K_{S}^{0}$, where $\ell^{-}$ is either an electron or a muon. We combine $e^+e^-$ data samples recorded by the Belle II experiment at the SuperKEKB collider (428 fb$^{-1}$) with samples recorded by the Belle experiment at the KEKB collider (980 fb$^{-1}$) to obtain a sample of 1.3 billion $e^+e^-\toτ^+τ^-$ events. We observe 0 and 1 events and set $90\%$ confidence level upper limits of $0.8 \times 10^{-8}$ and $1.2 \times 10^{-8}$ on the branching fractions of the decay modes $τ^{-} \rightarrow e^{-}K_{S}^{0}$ and $τ^{-} \rightarrow μ^{-}K_{S}^{0}$, respectively. These are the most stringent upper limits to date.
△ Less
Submitted 22 April, 2025;
originally announced April 2025.
-
Twin Co-Adaptive Dialogue for Progressive Image Generation
Authors:
Jianhui Wang,
Yangfan He,
Yan Zhong,
Xinyuan Song,
Jiayi Su,
Yuheng Feng,
Hongyang He,
Wenyu Zhu,
Xinhang Yuan,
Kuan Lu,
Menghao Huo,
Miao Zhang,
Keqin Li,
Jiaqi Chen,
Tianyu Shi,
Xueqian Wang
Abstract:
Modern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in user prompts. In this work, we present Twin-Co, a framework that leverages synchronized, co-adaptive dialogue to progressively refine image generation. Instead of a static generation process, Twin-Co employs a dynamic, i…
▽ More
Modern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in user prompts. In this work, we present Twin-Co, a framework that leverages synchronized, co-adaptive dialogue to progressively refine image generation. Instead of a static generation process, Twin-Co employs a dynamic, iterative workflow where an intelligent dialogue agent continuously interacts with the user. Initially, a base image is generated from the user's prompt. Then, through a series of synchronized dialogue exchanges, the system adapts and optimizes the image according to evolving user feedback. The co-adaptive process allows the system to progressively narrow down ambiguities and better align with user intent. Experiments demonstrate that Twin-Co not only enhances user experience by reducing trial-and-error iterations but also improves the quality of the generated images, streamlining the creative process across various applications.
△ Less
Submitted 21 April, 2025;
originally announced April 2025.
-
Correction for nonignorable nonresponse bias in the estimation of turnout using callback data
Authors:
Xinyu Li,
Naiwen Ying,
Kendrick Qijun Li,
Xu Shi,
Wang Miao
Abstract:
Overestimation of turnout has long been an issue in election surveys, with nonresponse bias or voter overrepresentation regarded as one of the major sources of bias. However, the adjustment for nonignorable nonresponse bias is substantially challenging. Based on the ANES Non-Response Follow-Up Study concerning the 2020 U.S. presidential election, we investigate the role of callback data in adjusti…
▽ More
Overestimation of turnout has long been an issue in election surveys, with nonresponse bias or voter overrepresentation regarded as one of the major sources of bias. However, the adjustment for nonignorable nonresponse bias is substantially challenging. Based on the ANES Non-Response Follow-Up Study concerning the 2020 U.S. presidential election, we investigate the role of callback data in adjusting for nonresponse bias in the estimation of turnout. Callback data are the records of contact attempts in the survey course, available in many modern large-scale surveys. We propose a stableness of resistance assumption to account for the nonignorable missingness in the outcome, which states that the impact of the missing outcome on the response propensity is stable in the first two call attempts. Under this assumption and by leveraging covariates information from the census data, we establish the identifiability and develop estimation methods for turnout, including a doubly robust estimator. Our methods produce estimates very close to the official turnout and successfully capture the trend of declining willingness to vote as response reluctance or contact difficulty increases. This work hints at the importance of adjusting for nonignorable nonresponse bias and exhibits the promise of callback data for political surveys.
△ Less
Submitted 19 April, 2025;
originally announced April 2025.
-
ViMo: A Generative Visual GUI World Model for App Agent
Authors:
Dezhao Luo,
Bohan Tang,
Kang Li,
Georgios Papoudakis,
Jifei Song,
Shaogang Gong,
Jianye Hao,
Jun Wang,
Kun Shao
Abstract:
App agents, which autonomously operate mobile Apps through Graphical User Interfaces (GUIs), have gained significant interest in real-world applications. Yet, they often struggle with long-horizon planning, failing to find the optimal actions for complex tasks with longer steps. To address this, world models are used to predict the next GUI observation based on user actions, enabling more effectiv…
▽ More
App agents, which autonomously operate mobile Apps through Graphical User Interfaces (GUIs), have gained significant interest in real-world applications. Yet, they often struggle with long-horizon planning, failing to find the optimal actions for complex tasks with longer steps. To address this, world models are used to predict the next GUI observation based on user actions, enabling more effective agent planning. However, existing world models primarily focus on generating only textual descriptions, lacking essential visual details. To fill this gap, we propose ViMo, the first visual world model designed to generate future App observations as images. For the challenge of generating text in image patches, where even minor pixel errors can distort readability, we decompose GUI generation into graphic and text content generation. We propose a novel data representation, the Symbolic Text Representation~(STR) to overlay text content with symbolic placeholders while preserving graphics. With this design, ViMo employs a STR Predictor to predict future GUIs' graphics and a GUI-text Predictor for generating the corresponding text. Moreover, we deploy ViMo to enhance agent-focused tasks by predicting the outcome of different action options. Experiments show ViMo's ability to generate visually plausible and functionally effective GUIs that enable App agents to make more informed decisions.
△ Less
Submitted 15 April, 2025;
originally announced April 2025.
-
Search for $J/ψ\rightarrow K^{0}_{S}K^{0}_{S}$ and $ψ(3686)\rightarrow K^{0}_{S}K^{0}_{S}$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (680 additional authors not shown)
Abstract:
Using data samples of $(10087\pm 44)\times10^{6}$ $J/ψ$ events and $(2712.4\pm 14.3)\times10^{6}$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, we search for the CP violating decays $J/ψ\rightarrow K^{0}_{S}K^{0}_{S}$ and $ψ(3686)\rightarrow K^{0}_{S}K^{0}_{S}$. No significant signals are observed over the expected background yields. The upper limits on their branchin…
▽ More
Using data samples of $(10087\pm 44)\times10^{6}$ $J/ψ$ events and $(2712.4\pm 14.3)\times10^{6}$ $ψ(3686)$ events collected with the BESIII detector at the BEPCII collider, we search for the CP violating decays $J/ψ\rightarrow K^{0}_{S}K^{0}_{S}$ and $ψ(3686)\rightarrow K^{0}_{S}K^{0}_{S}$. No significant signals are observed over the expected background yields. The upper limits on their branching fractions are set as $\mathcal{B}(J/ψ\rightarrow K^{0}_{S}K^{0}_{S}) <4.7\times 10^{-9}$ and $\mathcal{B}(ψ(3686)\rightarrow K^{0}_{S}K^{0}_{S}) <1.1\times 10^{-8}$ at the 90% confidence level. These results improve the previous limits by a factor of three for $J/ψ\rightarrow K^{0}_{S} K^{0}_{S}$ and two orders of magnitude for $ψ(3686)\rightarrow K^{0}_{S} K^{0}_{S}$.
△ Less
Submitted 18 April, 2025;
originally announced April 2025.
-
Search for $1^{-+}$ charmonium-like hybrid via $e^{+}e^{-}\rightarrow γη^{(\prime)} η_{c}$ at center-of-mass energies between 4.258 and 4.681 GeV
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (696 additional authors not shown)
Abstract:
Using $e^{+}e^{-}$ collision data corresponding to an integrated luminosity of 10.6 fb$^{-1}$ collected at center-of-mass energies between 4.258 and 4.681 GeV with the BESIII detector at the BEPCII collider, we search for the $1^{- +}$ charmonium-like hybrid via $e^{+}e^{-}\rightarrowγηη_{c}$ and $e^{+}e^{-}\rightarrowγη^{\prime}η_{c}$ decays for the first time. No significant signal is observed a…
▽ More
Using $e^{+}e^{-}$ collision data corresponding to an integrated luminosity of 10.6 fb$^{-1}$ collected at center-of-mass energies between 4.258 and 4.681 GeV with the BESIII detector at the BEPCII collider, we search for the $1^{- +}$ charmonium-like hybrid via $e^{+}e^{-}\rightarrowγηη_{c}$ and $e^{+}e^{-}\rightarrowγη^{\prime}η_{c}$ decays for the first time. No significant signal is observed and the upper limits on the Born cross sections for both processes are set at the 90% confidence level.
△ Less
Submitted 18 April, 2025;
originally announced April 2025.
-
Response to recent comments on Phys. Rev. B 107, 245423 (2023) and Subsection S4.3 of the Supp. Info. for Nature 638, 651-655 (2025)
Authors:
Morteza Aghaee,
Zulfi Alam,
Mariusz Andrzejczuk,
Andrey E. Antipov,
Mikhail Astafev,
Amin Barzegar,
Bela Bauer,
Jonathan Becker,
Umesh Kumar Bhaskar,
Alex Bocharov,
Srini Boddapati,
David Bohn,
Jouri Bommer,
Leo Bourdet,
Samuel Boutin,
Benjamin J. Chapman,
Sohail Chatoor,
Anna Wulff Christensen,
Patrick Codd,
William S. Cole,
Paul Cooper,
Fabiano Corsetti,
Ajuan Cui,
Andreas Ekefjärd,
Saeed Fallahi
, et al. (105 additional authors not shown)
Abstract:
The topological gap protocol (TGP) is a statistical test designed to identify a topological phase with high confidence and without human bias. It is used to determine a promising parameter regime for operating topological qubits. The protocol's key metric is the probability of incorrectly identifying a trivial region as topological, referred to as the false discovery rate (FDR). Two recent manuscr…
▽ More
The topological gap protocol (TGP) is a statistical test designed to identify a topological phase with high confidence and without human bias. It is used to determine a promising parameter regime for operating topological qubits. The protocol's key metric is the probability of incorrectly identifying a trivial region as topological, referred to as the false discovery rate (FDR). Two recent manuscripts [arXiv:2502.19560, arXiv:2503.08944] engage with the topological gap protocol and its use in Phys. Rev. B 107, 245423 (2023) and Subsection S4.3 of the Supplementary Information for Nature 638, 651-655 (2025), although they do not explicitly dispute the main results of either one. We demonstrate that the objections in arXiv:2502.19560 and arXiv:2503.08944 are unfounded, and we uphold the conclusions of Phys. Rev. B 107, 245423 (2023) and Nature 638, 651-655 (2025). Specifically, we show that no flaws have been identified in our estimate of the false discovery rate (FDR). We provide a point-by-point rebuttal of the comments in arXiv:2502.19560 and arXiv:2503.08944.
△ Less
Submitted 17 April, 2025;
originally announced April 2025.
-
The Tenth NTIRE 2025 Image Denoising Challenge Report
Authors:
Lei Sun,
Hang Guo,
Bin Ren,
Luc Van Gool,
Radu Timofte,
Yawei Li,
Xiangyu Kong,
Hyunhee Park,
Xiaoxuan Yu,
Suejin Han,
Hakjae Jeon,
Jia Li,
Hyung-Ju Chun,
Donghun Ryou,
Inju Ha,
Bohyung Han,
Jingyu Ma,
Zhijuan Huang,
Huiyuan Fu,
Hongyuan Yu,
Boqi Zhang,
Jiawei Shi,
Heng Zhang,
Huadong Ma,
Deepak Kumar Tyagi
, et al. (69 additional authors not shown)
Abstract:
This paper presents an overview of the NTIRE 2025 Image Denoising Challenge (σ = 50), highlighting the proposed methodologies and corresponding results. The primary objective is to develop a network architecture capable of achieving high-quality denoising performance, quantitatively evaluated using PSNR, without constraints on computational complexity or model size. The task assumes independent ad…
▽ More
This paper presents an overview of the NTIRE 2025 Image Denoising Challenge (σ = 50), highlighting the proposed methodologies and corresponding results. The primary objective is to develop a network architecture capable of achieving high-quality denoising performance, quantitatively evaluated using PSNR, without constraints on computational complexity or model size. The task assumes independent additive white Gaussian noise (AWGN) with a fixed noise level of 50. A total of 290 participants registered for the challenge, with 20 teams successfully submitting valid results, providing insights into the current state-of-the-art in image denoising.
△ Less
Submitted 16 April, 2025;
originally announced April 2025.
-
RESPLE: Recursive Spline Estimation for LiDAR-Based Odometry
Authors:
Ziyu Cao,
William Talbot,
Kailai Li
Abstract:
We present a novel recursive Bayesian estimation framework for continuous-time six-DoF dynamic motion estimation using B-splines. The state vector consists of a recurrent set of position control points and orientation control point increments, enabling a straightforward modification of the iterated extended Kalman filter without involving the error-state formulation. The resulting recursive spline…
▽ More
We present a novel recursive Bayesian estimation framework for continuous-time six-DoF dynamic motion estimation using B-splines. The state vector consists of a recurrent set of position control points and orientation control point increments, enabling a straightforward modification of the iterated extended Kalman filter without involving the error-state formulation. The resulting recursive spline estimator (RESPLE) provides a versatile, pragmatic and lightweight solution for motion estimation and is further exploited for direct LiDAR-based odometry, supporting integration of one or multiple LiDARs and an IMU. We conduct extensive real-world benchmarking based on public datasets and own experiments, covering aerial, wheeled, legged, and wearable platforms operating in indoor, urban, wild environments with diverse LiDARs. RESPLE-based solutions achieve superior estimation accuracy and robustness over corresponding state-of-the-art systems, while attaining real-time performance. Notably, our LiDAR-only variant outperforms existing LiDAR-inertial systems in scenarios without significant LiDAR degeneracy, and showing further improvements when additional LiDAR and inertial sensors are incorporated for more challenging conditions. We release the source code and own experimental datasets at https://github.com/ASIG-X/RESPLE .
△ Less
Submitted 15 April, 2025;
originally announced April 2025.
-
DeepSelective: Feature Gating and Representation Matching for Interpretable Clinical Prediction
Authors:
Ruochi Zhang,
Qian Yang,
Xiaoyang Wang,
Haoran Wu,
Qiong Zhou,
Yu Wang,
Kewei Li,
Yueying Wang,
Yusi Fan,
Jiale Zhang,
Lan Huang,
Chang Liu,
Fengfeng Zhou
Abstract:
The rapid accumulation of Electronic Health Records (EHRs) has transformed healthcare by providing valuable data that enhance clinical predictions and diagnoses. While conventional machine learning models have proven effective, they often lack robust representation learning and depend heavily on expert-crafted features. Although deep learning offers powerful solutions, it is often criticized for i…
▽ More
The rapid accumulation of Electronic Health Records (EHRs) has transformed healthcare by providing valuable data that enhance clinical predictions and diagnoses. While conventional machine learning models have proven effective, they often lack robust representation learning and depend heavily on expert-crafted features. Although deep learning offers powerful solutions, it is often criticized for its lack of interpretability. To address these challenges, we propose DeepSelective, a novel end to end deep learning framework for predicting patient prognosis using EHR data, with a strong emphasis on enhancing model interpretability. DeepSelective combines data compression techniques with an innovative feature selection approach, integrating custom-designed modules that work together to improve both accuracy and interpretability. Our experiments demonstrate that DeepSelective not only enhances predictive accuracy but also significantly improves interpretability, making it a valuable tool for clinical decision-making. The source code is freely available at http://www.healthinformaticslab.org/supp/resources.php .
△ Less
Submitted 15 April, 2025;
originally announced April 2025.
-
Test of lepton flavor universality with measurements of $R(D^{+})$ and $R(D^{*+})$ using semileptonic $B$ tagging at the Belle II experiment
Authors:
Belle II Collaboration,
I. Adachi,
K. Adamczyk,
L. Aggarwal,
H. Ahmed,
H. Aihara,
N. Akopov,
S. Alghamdi,
M. Alhakami,
A. Aloisio,
N. Althubiti,
K. Amos,
M. Angelsmark,
N. Anh Ky,
C. Antonioli,
D. M. Asner,
H. Atmacan,
T. Aushev,
V. Aushev,
M. Aversano,
R. Ayad,
V. Babu,
H. Bae,
N. K. Baghel,
S. Bahinipati
, et al. (428 additional authors not shown)
Abstract:
We report measurements of the ratios of branching fractions $\mathcal{R}(D^{(*)+}) = \mathcal{B}(\overline{B}{}^0 \to D^{(*)+} \,τ^- \, \overlineν_τ) / \mathcal{B}(\overline{B}{}^0 \to D^{(*)+} \, \ell^- \, \overlineν_\ell)$, where $\ell$ denotes either an electron or a muon. These ratios test the universality of the charged-current weak interaction. The results are based on a…
▽ More
We report measurements of the ratios of branching fractions $\mathcal{R}(D^{(*)+}) = \mathcal{B}(\overline{B}{}^0 \to D^{(*)+} \,τ^- \, \overlineν_τ) / \mathcal{B}(\overline{B}{}^0 \to D^{(*)+} \, \ell^- \, \overlineν_\ell)$, where $\ell$ denotes either an electron or a muon. These ratios test the universality of the charged-current weak interaction. The results are based on a $365\, \mathrm{fb}^{-1}$ data sample collected with the Belle II detector at the SuperKEKB $e^+e^-$ collider, which operates at a center-of-mass energy corresponding to the $Υ(4S)$ resonance, just above the threshold for $B\overline{B}{}$ production. Signal candidates are reconstructed by selecting events in which the companion $B$ meson from the $Υ(4S) \to B\overline{B}{}$ decay is identified in semileptonic modes. The $τ$ lepton is reconstructed via its leptonic decays. We obtain $\mathcal{R}(D^+) = 0.418 \pm 0.074 ~({\mathrm{stat}}) \pm 0.051 ~({\mathrm{syst}})$ and $\mathcal{R}(D^{*+}) = 0.306 \pm 0.034 ~({\mathrm{stat}}) \pm 0.018 ~({\mathrm{syst}})$, which are consistent with world average values. Accounting for the correlation between them, these values differ from the Standard Model expectation by a collective significance of $1.7$ standard deviations.
△ Less
Submitted 15 April, 2025;
originally announced April 2025.
-
Benchmarking Next-Generation Reasoning-Focused Large Language Models in Ophthalmology: A Head-to-Head Evaluation on 5,888 Items
Authors:
Minjie Zou,
Sahana Srinivasan,
Thaddaeus Wai Soon Lo,
Ke Zou,
Gabriel Dawei Yang,
Xuguang Ai,
Hyunjae Kim,
Maxwell Singer,
Fares Antaki,
Kelvin Li,
Robert Chang,
Marcus Tan,
David Ziyou Chen,
Dianbo Liu,
Qingyu Chen,
Yih Chung Tham
Abstract:
Recent advances in reasoning-focused large language models (LLMs) mark a shift from general LLMs toward models designed for complex decision-making, a crucial aspect in medicine. However, their performance in specialized domains like ophthalmology remains underexplored. This study comprehensively evaluated and compared the accuracy and reasoning capabilities of four newly developed reasoning-focus…
▽ More
Recent advances in reasoning-focused large language models (LLMs) mark a shift from general LLMs toward models designed for complex decision-making, a crucial aspect in medicine. However, their performance in specialized domains like ophthalmology remains underexplored. This study comprehensively evaluated and compared the accuracy and reasoning capabilities of four newly developed reasoning-focused LLMs, namely DeepSeek-R1, OpenAI o1, o3-mini, and Gemini 2.0 Flash-Thinking. Each model was assessed using 5,888 multiple-choice ophthalmology exam questions from the MedMCQA dataset in zero-shot setting. Quantitative evaluation included accuracy, Macro-F1, and five text-generation metrics (ROUGE-L, METEOR, BERTScore, BARTScore, and AlignScore), computed against ground-truth reasonings. Average inference time was recorded for a subset of 100 randomly selected questions. Additionally, two board-certified ophthalmologists qualitatively assessed clarity, completeness, and reasoning structure of responses to differential diagnosis questions.O1 (0.902) and DeepSeek-R1 (0.888) achieved the highest accuracy, with o1 also leading in Macro-F1 (0.900). The performance of models across the text-generation metrics varied: O3-mini excelled in ROUGE-L (0.151), o1 in METEOR (0.232), DeepSeek-R1 and o3-mini tied for BERTScore (0.673), DeepSeek-R1 (-4.105) and Gemini 2.0 Flash-Thinking (-4.127) performed best in BARTScore, while o3-mini (0.181) and o1 (0.176) led AlignScore. Inference time across the models varied, with DeepSeek-R1 being slowest (40.4 seconds) and Gemini 2.0 Flash-Thinking fastest (6.7 seconds). Qualitative evaluation revealed that DeepSeek-R1 and Gemini 2.0 Flash-Thinking tended to provide detailed and comprehensive intermediate reasoning, whereas o1 and o3-mini displayed concise and summarized justifications.
△ Less
Submitted 15 April, 2025;
originally announced April 2025.
-
Mathematical Analysis of the PDE Model for the Consensus-based Optimization
Authors:
Jinhuan Wang,
Keyu Li,
Hui Huang
Abstract:
In this paper, we develop an analytical framework for the partial differential equation underlying the consensus-based optimization model. The main challenge arises from the nonlinear, nonlocal nature of the consensus point, coupled with a diffusion term that is both singular and degenerate. By employing a regularization procedure in combination with a compactness argument, we establish the global…
▽ More
In this paper, we develop an analytical framework for the partial differential equation underlying the consensus-based optimization model. The main challenge arises from the nonlinear, nonlocal nature of the consensus point, coupled with a diffusion term that is both singular and degenerate. By employing a regularization procedure in combination with a compactness argument, we establish the global existence and uniqueness of weak solutions in $L^\infty(0,T;L^1\cap L^\infty(\mathbb{R}^d))$. Furthermore, we show that the weak solutions exhibit improved $H^2$-regularity when the initial data is regular.
△ Less
Submitted 15 April, 2025;
originally announced April 2025.
-
Precise measurement of the form factors in $D^0\rightarrow K^*(892)^-μ^+ν_μ$ and test of lepton universality with $D^0\rightarrow K^*(892)^-\ell^+ν_{\ell}$ decays
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (696 additional authors not shown)
Abstract:
We report a study of the semileptonic decay $D^0 \rightarrow \bar{K}^0π^-μ^+ν_μ$ based on a sample of $7.9~\mathrm{fb}^{-1}$ of $e^+e^-$ annihilation data collected at a center-of-mass energy of 3.773~GeV with the BESIII detector at the BEPCII collider. The branching fraction of the decay is measured for the first time to be…
▽ More
We report a study of the semileptonic decay $D^0 \rightarrow \bar{K}^0π^-μ^+ν_μ$ based on a sample of $7.9~\mathrm{fb}^{-1}$ of $e^+e^-$ annihilation data collected at a center-of-mass energy of 3.773~GeV with the BESIII detector at the BEPCII collider. The branching fraction of the decay is measured for the first time to be $\mathcal{B}(D^0\rightarrow \bar{K}^0π^-μ^+ν_μ) = (1.373 \pm 0.020_{\rm stat} \pm 0.023_{\rm syst})\%$, where the first uncertainty is statistical and the second is systematic. Based on the investigation of the decay dynamics, we find that the decay is dominated by the $K^{*}(892)^-$ resonance with the branching fraction measured to be $\mathcal{B}(D^0\rightarrow K^{*}(892)^-μ^+ν_μ) = (1.948 \pm 0.033_{\rm stat} \pm 0.036_{\rm syst})\%$. We also determine the hadronic form factors for the $D^0\rightarrow K^{*}(892)^-μ^+ν_μ$ decay to be $r_{V} = V(0)/A_1(0) = 1.46 \pm 0.11_{\rm stat} \pm 0.04_{\rm syst}$, $r_{2} = A_2(0)/A_1(0) = 0.71 \pm 0.08_{\rm stat} \pm 0.03_{\rm syst}$, and $A_1(0)=0.609 \pm 0.008_{\rm stat} \pm 0.008_{\rm syst}$, where $V(0)$ is the vector form factor and $A_{1,2}(0)$ are the axial form factors evaluated at $q^2=0$. The $A_1(0)$ is measured for the first time in $D^0\rightarrow K^{*}(892)^-μ^+ν_μ$ decay. Averaging the form-factor parameters that we reported previously in $D^0\rightarrow K^*(892)^-(\rightarrow \bar{K}^0π^-)e^+ν_{e}$ and $D^0\rightarrow K^*(892)^-(\rightarrow K^-π^0)μ^+ν_μ$ decays, we obtain $r_{V}=1.456\pm0.040_{\rm stat}\pm0.016_{\rm syst}$, $r_{2}=0.715\pm0.031_{\rm stat}\pm0.014_{\rm stat}$, and $A_1(0)=0.614\pm0.005_{\rm stat}\pm0.004_{\rm syst}$. This is the most precise determination of the form-factor parameters to date measured in $D\rightarrow K^*(892)$ transition, which provide the most stringent test on various theoretical models.
△ Less
Submitted 15 April, 2025;
originally announced April 2025.
-
The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report
Authors:
Bin Ren,
Hang Guo,
Lei Sun,
Zongwei Wu,
Radu Timofte,
Yawei Li,
Yao Zhang,
Xinning Chai,
Zhengxue Cheng,
Yingsheng Qin,
Yucai Yang,
Li Song,
Hongyuan Yu,
Pufan Xu,
Cheng Wan,
Zhijuan Huang,
Peng Guo,
Shuyuan Cui,
Chenjun Li,
Xuehai Hu,
Pan Pan,
Xin Zhang,
Heng Zhang,
Qing Luo,
Linyan Jiang
, et al. (122 additional authors not shown)
Abstract:
This paper presents a comprehensive review of the NTIRE 2025 Challenge on Single-Image Efficient Super-Resolution (ESR). The challenge aimed to advance the development of deep models that optimize key computational metrics, i.e., runtime, parameters, and FLOPs, while achieving a PSNR of at least 26.90 dB on the $\operatorname{DIV2K\_LSDIR\_valid}$ dataset and 26.99 dB on the…
▽ More
This paper presents a comprehensive review of the NTIRE 2025 Challenge on Single-Image Efficient Super-Resolution (ESR). The challenge aimed to advance the development of deep models that optimize key computational metrics, i.e., runtime, parameters, and FLOPs, while achieving a PSNR of at least 26.90 dB on the $\operatorname{DIV2K\_LSDIR\_valid}$ dataset and 26.99 dB on the $\operatorname{DIV2K\_LSDIR\_test}$ dataset. A robust participation saw \textbf{244} registered entrants, with \textbf{43} teams submitting valid entries. This report meticulously analyzes these methods and results, emphasizing groundbreaking advancements in state-of-the-art single-image ESR techniques. The analysis highlights innovative approaches and establishes benchmarks for future research in the field.
△ Less
Submitted 14 April, 2025;
originally announced April 2025.
-
Undermining Federated Learning Accuracy in EdgeIoT via Variational Graph Auto-Encoders
Authors:
Kai Li,
Shuyan Hu,
Bochun Wu,
Sai Zou,
Wei Ni,
Falko Dressler
Abstract:
EdgeIoT represents an approach that brings together mobile edge computing with Internet of Things (IoT) devices, allowing for data processing close to the data source. Sending source data to a server is bandwidth-intensive and may compromise privacy. Instead, federated learning allows each device to upload a shared machine-learning model update with locally processed data. However, this technique,…
▽ More
EdgeIoT represents an approach that brings together mobile edge computing with Internet of Things (IoT) devices, allowing for data processing close to the data source. Sending source data to a server is bandwidth-intensive and may compromise privacy. Instead, federated learning allows each device to upload a shared machine-learning model update with locally processed data. However, this technique, which depends on aggregating model updates from various IoT devices, is vulnerable to attacks from malicious entities that may inject harmful data into the learning process. This paper introduces a new attack method targeting federated learning in EdgeIoT, known as data-independent model manipulation attack. This attack does not rely on training data from the IoT devices but instead uses an adversarial variational graph auto-encoder (AV-GAE) to create malicious model updates by analyzing benign model updates intercepted during communication. AV-GAE identifies and exploits structural relationships between benign models and their training data features. By manipulating these structural correlations, the attack maximizes the training loss of the federated learning system, compromising its overall effectiveness.
△ Less
Submitted 14 April, 2025;
originally announced April 2025.
-
Search for $B^0 \to K^{\ast 0} τ^+ τ^-$ decays at the Belle II experiment
Authors:
Belle II Collaboration,
I. Adachi,
K. Adamczyk,
L. Aggarwal,
H. Ahmed,
H. Aihara,
N. Akopov,
M. Alhakami,
A. Aloisio,
N. Althubiti,
M. Angelsmark,
N. Anh Ky,
D. M. Asner,
H. Atmacan,
V. Aushev,
M. Aversano,
R. Ayad,
V. Babu,
H. Bae,
N. K. Baghel,
S. Bahinipati,
P. Bambade,
Sw. Banerjee,
S. Bansal,
M. Barrett
, et al. (424 additional authors not shown)
Abstract:
We present a search for the rare flavor-changing neutral-current decay $B^0 \to K^{\ast 0} τ^+ τ^-$ with data collected by the Belle II experiment at the SuperKEKB electron-positron collider. The analysis uses a 365 fb$^{-1}$ data sample recorded at the center-of-mass energy of the $Υ(4S)$ resonance. One of the $B$ mesons produced in the $Υ(4S)\to B^0 \bar{B}^0$ process is fully reconstructed in a…
▽ More
We present a search for the rare flavor-changing neutral-current decay $B^0 \to K^{\ast 0} τ^+ τ^-$ with data collected by the Belle II experiment at the SuperKEKB electron-positron collider. The analysis uses a 365 fb$^{-1}$ data sample recorded at the center-of-mass energy of the $Υ(4S)$ resonance. One of the $B$ mesons produced in the $Υ(4S)\to B^0 \bar{B}^0$ process is fully reconstructed in a hadronic decay mode, while its companion $B$ meson is required to decay into a $K^{\ast 0}$ and two $τ$ leptons of opposite charge. The $τ$ leptons are reconstructed in final states with a single electron, muon, charged pion or charged $ρ$ meson, and additional neutrinos. We set an upper limit on the branching ratio of $BR(B^0 \to K^{\ast 0} τ^+ τ^-) < 1.8 \times 10^{-3}$ at the 90% confidence level, which is the most stringent constraint reported to date.
△ Less
Submitted 14 April, 2025;
originally announced April 2025.
-
SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
Authors:
Kaiyu Li,
Zepeng Xin,
Li Pang,
Chao Pang,
Yupeng Deng,
Jing Yao,
Guisong Xia,
Deyu Meng,
Zhi Wang,
Xiangyong Cao
Abstract:
Remote sensing has become critical for understanding environmental dynamics, urban planning, and disaster management. However, traditional remote sensing workflows often rely on explicit segmentation or detection methods, which struggle to handle complex, implicit queries that require reasoning over spatial context, domain knowledge, and implicit user intent. Motivated by this, we introduce a new…
▽ More
Remote sensing has become critical for understanding environmental dynamics, urban planning, and disaster management. However, traditional remote sensing workflows often rely on explicit segmentation or detection methods, which struggle to handle complex, implicit queries that require reasoning over spatial context, domain knowledge, and implicit user intent. Motivated by this, we introduce a new task, \ie, geospatial pixel reasoning, which allows implicit querying and reasoning and generates the mask of the target region. To advance this task, we construct and release the first large-scale benchmark dataset called EarthReason, which comprises 5,434 manually annotated image masks with over 30,000 implicit question-answer pairs. Moreover, we propose SegEarth-R1, a simple yet effective language-guided segmentation baseline that integrates a hierarchical visual encoder, a large language model (LLM) for instruction parsing, and a tailored mask generator for spatial correlation. The design of SegEarth-R1 incorporates domain-specific adaptations, including aggressive visual token compression to handle ultra-high-resolution remote sensing images, a description projection module to fuse language and multi-scale features, and a streamlined mask prediction pipeline that directly queries description embeddings. Extensive experiments demonstrate that SegEarth-R1 achieves state-of-the-art performance on both reasoning and referring segmentation tasks, significantly outperforming traditional and LLM-based segmentation methods. Our data and code will be released at https://github.com/earth-insights/SegEarth-R1.
△ Less
Submitted 13 April, 2025;
originally announced April 2025.
-
Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large Images
Authors:
Jiuchen Chen,
Xinyu Yan,
Qizhi Xu,
Kaiqi Li
Abstract:
Global contextual information and local detail features are essential for haze removal tasks. Deep learning models perform well on small, low-resolution images, but they encounter difficulties with large, high-resolution ones due to GPU memory limitations. As a compromise, they often resort to image slicing or downsampling. The former diminishes global information, while the latter discards high-f…
▽ More
Global contextual information and local detail features are essential for haze removal tasks. Deep learning models perform well on small, low-resolution images, but they encounter difficulties with large, high-resolution ones due to GPU memory limitations. As a compromise, they often resort to image slicing or downsampling. The former diminishes global information, while the latter discards high-frequency details. To address these challenges, we propose DehazeXL, a haze removal method that effectively balances global context and local feature extraction, enabling end-to-end modeling of large images on mainstream GPU hardware. Additionally, to evaluate the efficiency of global context utilization in haze removal performance, we design a visual attribution method tailored to the characteristics of haze removal tasks. Finally, recognizing the lack of benchmark datasets for haze removal in large images, we have developed an ultra-high-resolution haze removal dataset (8KDehaze) to support model training and testing. It includes 10000 pairs of clear and hazy remote sensing images, each sized at 8192 $\times$ 8192 pixels. Extensive experiments demonstrate that DehazeXL can infer images up to 10240 $\times$ 10240 pixels with only 21 GB of memory, achieving state-of-the-art results among all evaluated methods. The source code and experimental dataset are available at https://github.com/CastleChen339/DehazeXL.
△ Less
Submitted 13 April, 2025;
originally announced April 2025.
-
Optimal Transport-Based Generative Models for Bayesian Posterior Sampling
Authors:
Ke Li,
Wei Han,
Yuexi Wang,
Yun Yang
Abstract:
We investigate the problem of sampling from posterior distributions with intractable normalizing constants in Bayesian inference. Our solution is a new generative modeling approach based on optimal transport (OT) that learns a deterministic map from a reference distribution to the target posterior through constrained optimization. The method uses structural constraints from OT theory to ensure uni…
▽ More
We investigate the problem of sampling from posterior distributions with intractable normalizing constants in Bayesian inference. Our solution is a new generative modeling approach based on optimal transport (OT) that learns a deterministic map from a reference distribution to the target posterior through constrained optimization. The method uses structural constraints from OT theory to ensure uniqueness of the solution and allows efficient generation of many independent, high-quality posterior samples. The framework supports both continuous and mixed discrete-continuous parameter spaces, with specific adaptations for latent variable models and near-Gaussian posteriors. Beyond computational benefits, it also enables new inferential tools based on OT-derived multivariate ranks and quantiles for Bayesian exploratory analysis and visualization. We demonstrate the effectiveness of our approach through multiple simulation studies and a real-world data analysis.
△ Less
Submitted 10 April, 2025;
originally announced April 2025.
-
On the Practice of Deep Hierarchical Ensemble Network for Ad Conversion Rate Prediction
Authors:
Jinfeng Zhuang,
Yinrui Li,
Runze Su,
Ke Xu,
Zhixuan Shao,
Kungang Li,
Ling Leng,
Han Sun,
Meng Qi,
Yixiong Meng,
Yang Tang,
Zhifang Liu,
Qifei Shen,
Aayush Mudgal,
Caleb Lu,
Jie Liu,
Hongda Shen
Abstract:
The predictions of click through rate (CTR) and conversion rate (CVR) play a crucial role in the success of ad-recommendation systems. A Deep Hierarchical Ensemble Network (DHEN) has been proposed to integrate multiple feature crossing modules and has achieved great success in CTR prediction. However, its performance for CVR prediction is unclear in the conversion ads setting, where an ad bids for…
▽ More
The predictions of click through rate (CTR) and conversion rate (CVR) play a crucial role in the success of ad-recommendation systems. A Deep Hierarchical Ensemble Network (DHEN) has been proposed to integrate multiple feature crossing modules and has achieved great success in CTR prediction. However, its performance for CVR prediction is unclear in the conversion ads setting, where an ad bids for the probability of a user's off-site actions on a third party website or app, including purchase, add to cart, sign up, etc. A few challenges in DHEN: 1) What feature-crossing modules (MLP, DCN, Transformer, to name a few) should be included in DHEN? 2) How deep and wide should DHEN be to achieve the best trade-off between efficiency and efficacy? 3) What hyper-parameters to choose in each feature-crossing module? Orthogonal to the model architecture, the input personalization features also significantly impact model performance with a high degree of freedom. In this paper, we attack this problem and present our contributions biased to the applied data science side, including:
First, we propose a multitask learning framework with DHEN as the single backbone model architecture to predict all CVR tasks, with a detailed study on how to make DHEN work effectively in practice; Second, we build both on-site real-time user behavior sequences and off-site conversion event sequences for CVR prediction purposes, and conduct ablation study on its importance; Last but not least, we propose a self-supervised auxiliary loss to predict future actions in the input sequence, to help resolve the label sparseness issue in CVR prediction.
Our method achieves state-of-the-art performance compared to previous single feature crossing modules with pre-trained user personalization features.
△ Less
Submitted 23 April, 2025; v1 submitted 10 April, 2025;
originally announced April 2025.
-
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
Authors:
Kaixin Li,
Ziyang Meng,
Hongzhan Lin,
Ziyang Luo,
Yuchen Tian,
Jing Ma,
Zhiyong Huang,
Tat-Seng Chua
Abstract:
Recent advancements in Multi-modal Large Language Models (MLLMs) have led to significant progress in developing GUI agents for general tasks such as web browsing and mobile phone use. However, their application in professional domains remains under-explored. These specialized workflows introduce unique challenges for GUI perception models, including high-resolution displays, smaller target sizes,…
▽ More
Recent advancements in Multi-modal Large Language Models (MLLMs) have led to significant progress in developing GUI agents for general tasks such as web browsing and mobile phone use. However, their application in professional domains remains under-explored. These specialized workflows introduce unique challenges for GUI perception models, including high-resolution displays, smaller target sizes, and complex environments. In this paper, we introduce ScreenSpot-Pro, a new benchmark designed to rigorously evaluate the grounding capabilities of MLLMs in high-resolution professional settings. The benchmark comprises authentic high-resolution images from a variety of professional domains with expert annotations. It spans 23 applications across five industries and three operating systems. Existing GUI grounding models perform poorly on this dataset, with the best model achieving only 18.9%. Our experiments reveal that strategically reducing the search area enhances accuracy. Based on this insight, we propose ScreenSeekeR, a visual search method that utilizes the GUI knowledge of a strong planner to guide a cascaded search, achieving state-of-the-art performance with 48.1% without any additional training. We hope that our benchmark and findings will advance the development of GUI agents for professional applications. Code, data and leaderboard can be found at https://gui-agent.github.io/grounding-leaderboard.
△ Less
Submitted 4 April, 2025;
originally announced April 2025.
-
Search for the baryon and lepton number violating decay $J/ψ\to pe^-$ + c.c
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (664 additional authors not shown)
Abstract:
Based on $(2712.4\pm 14.3) \times 10^{6} $ ${ψ(3686)}$ events collected by the BESIII detector operating at the BEPCII storage ring, we perform a search for the baryon- and lepton-number violating decay $J/ψ\to pe^{-}+c.c.$ via $ψ(3686) \to π^{+}π^{-}J/ψ$. No significant signal is found. An upper limit on the branching fraction of $\mathcal{B}(J/ψ\to p e^{-}+ c.c.) < 3.1 \times 10^{-8}$ at 90\% co…
▽ More
Based on $(2712.4\pm 14.3) \times 10^{6} $ ${ψ(3686)}$ events collected by the BESIII detector operating at the BEPCII storage ring, we perform a search for the baryon- and lepton-number violating decay $J/ψ\to pe^{-}+c.c.$ via $ψ(3686) \to π^{+}π^{-}J/ψ$. No significant signal is found. An upper limit on the branching fraction of $\mathcal{B}(J/ψ\to p e^{-}+ c.c.) < 3.1 \times 10^{-8}$ at 90\% confidence level.
△ Less
Submitted 10 April, 2025;
originally announced April 2025.
-
Microlensing at Cosmological Distances: Event Rate Predictions in the Warhol Arc of MACS 0416
Authors:
J. M. Palencia,
J. M. Diego,
L. Dai,
M. Pascale,
R. Windhorst,
A. M. Koekemoer,
Sung Kei Li,
B. J. Kavanagh,
Fengwu Sun,
Amruth Alfred,
Ashish K. Meena,
Thomas J. Broadhurst,
Patrick L. Kelly,
Derek Perera,
Hayley Williams,
Adi Zitrin
Abstract:
Highly magnified stars ($μ$ $>$ 100) are now outinely identified as transient events at cosmological distances thanks to microlensing by intra-cluster stars near the critical curves of galaxy clusters. Using the {\it James Webb} Space Telescope (JWST) in combination with the {\it Hubble} Space Telescope (HST), we outline here an analytical framework that is applied to the Warhol arc (at $z=0.94$)…
▽ More
Highly magnified stars ($μ$ $>$ 100) are now outinely identified as transient events at cosmological distances thanks to microlensing by intra-cluster stars near the critical curves of galaxy clusters. Using the {\it James Webb} Space Telescope (JWST) in combination with the {\it Hubble} Space Telescope (HST), we outline here an analytical framework that is applied to the Warhol arc (at $z=0.94$) in the MACS 0416 galaxy cluster (at $z=0.396)$ where over a dozen microlensed stars have been detected to date. This method is general and can be applied to other lensed arcs. Within this lensed galaxy we fit the spatially resolved SED spanned by eight JWST-NIRCam filters combined with three ACS filters, for accurate lensed star predictions in 2D. With this tool we can generate 2D maps of microlensed stars for well resolved arcs in general, including dependence on wavelength and limiting apparent magnitude, for comparison with with planned cadenced campaigns for JWST and Hubble, for constraining directly the IMF and the level of dark matter substructure.
△ Less
Submitted 28 April, 2025; v1 submitted 9 April, 2025;
originally announced April 2025.
-
Constraining the z $\sim$ 1 Initial Mass Function with {\it HST} and {\it JWST} Lensed Stars in MACS J0416.1-2403
Authors:
Sung Kei Li,
Jose M. Diego,
Ashish K. Meena,
Jeremy Lim,
Leo W. H. Fung,
Arsen Levitskiy,
James Nianias,
Jose M. Palencia,
Hayley Williams,
Jiashuo Zhang,
Alfred Amruth,
Thomas J. Broadhurst,
Wenlei Chen,
Alexei V. Filippenko,
Patrick L. Kelly,
Anton M. Koekemoer,
Derek Perera,
Bangzheng Sun,
Liliya L. R. Williams,
Rogier A. Windhorst,
Haojin Yan,
Adi Zitrin
Abstract:
Our understanding of galaxy properties and evolution is contingent on knowing the initial mass function (IMF), and yet to date, the IMF is constrained only to local galaxies. Individual stars are now becoming routinely detected at cosmological distances, where luminous stars such as supergiants in background galaxies strongly lensed by galaxy clusters are temporarily further magnified by huge fact…
▽ More
Our understanding of galaxy properties and evolution is contingent on knowing the initial mass function (IMF), and yet to date, the IMF is constrained only to local galaxies. Individual stars are now becoming routinely detected at cosmological distances, where luminous stars such as supergiants in background galaxies strongly lensed by galaxy clusters are temporarily further magnified by huge factors (up to $10^{4}$) by intracluster stars, thus being detected as transients. The detection rate of these events depends on the abundance of luminous stars in the background galaxy and is thus sensitive to the IMF and the star-formation history (SFH), especially for the blue supergiants detected as transients in the rest-frame ultraviolet/optical filters. As a proof of concept, we use simple SFH and IMF models constrained by spectral energy distributions (SEDs) to see how well we can predict the {\it HST} and {\it JWST} transient detection rate in a lensed arc dubbed ``Spock'' (redshift $z = 1.0054$). We find that demanding a simultaneous fit of the SED and the rest-frame ultraviolet/optical transient detection rate places constraints on the IMF, independent of the assumed simple SFH model. We conclude our Bayesian likelihood analysis indicates that the data definitively prefers the ``Spock'' galaxy to have a Salpeter IMF ($α= 2.35$) rather than a Top-heavy IMF ($α= 1$) -- which is thought to be the case in the early universe -- with no clear excess of supergiants above the standard IMF.
△ Less
Submitted 28 April, 2025; v1 submitted 9 April, 2025;
originally announced April 2025.
-
CHIME: A Compressive Framework for Holistic Interest Modeling
Authors:
Yong Bai,
Rui Xiang,
Kaiyuan Li,
Yongxiang Tang,
Yanhua Cheng,
Xialong Liu,
Peng Jiang,
Kun Gai
Abstract:
Modeling holistic user interests is important for improving recommendation systems but is challenged by high computational cost and difficulty in handling diverse information with full behavior context. Existing search-based methods might lose critical signals during behavior selection. To overcome these limitations, we propose CHIME: A Compressive Framework for Holistic Interest Modeling. It uses…
▽ More
Modeling holistic user interests is important for improving recommendation systems but is challenged by high computational cost and difficulty in handling diverse information with full behavior context. Existing search-based methods might lose critical signals during behavior selection. To overcome these limitations, we propose CHIME: A Compressive Framework for Holistic Interest Modeling. It uses adapted large language models to encode complete user behaviors with heterogeneous inputs. We introduce multi-granular contrastive learning objectives to capture both persistent and transient interest patterns and apply residual vector quantization to generate compact embeddings. CHIME demonstrates superior ranking performance across diverse datasets, establishing a robust solution for scalable holistic interest modeling in recommendation systems.
△ Less
Submitted 9 April, 2025;
originally announced April 2025.
-
BBQRec: Behavior-Bind Quantization for Multi-Modal Sequential Recommendation
Authors:
Kaiyuan Li,
Rui Xiang,
Yong Bai,
Yongxiang Tang,
Yanhua Cheng,
Xialong Liu,
Peng Jiang,
Kun Gai
Abstract:
Multi-modal sequential recommendation systems leverage auxiliary signals (e.g., text, images) to alleviate data sparsity in user-item interactions. While recent methods exploit large language models to encode modalities into discrete semantic IDs for autoregressive prediction, we identify two critical limitations: (1) Existing approaches adopt fragmented quantization, where modalities are independ…
▽ More
Multi-modal sequential recommendation systems leverage auxiliary signals (e.g., text, images) to alleviate data sparsity in user-item interactions. While recent methods exploit large language models to encode modalities into discrete semantic IDs for autoregressive prediction, we identify two critical limitations: (1) Existing approaches adopt fragmented quantization, where modalities are independently mapped to semantic spaces misaligned with behavioral objectives, and (2) Over-reliance on semantic IDs disrupts inter-modal semantic coherence, thereby weakening the expressive power of multi-modal representations for modeling diverse user preferences.
To address these challenges, we propose a Behavior-Bind multi-modal Quantization for Sequential Recommendation (BBQRec for short) featuring dual-aligned quantization and semantics-aware sequence modeling. First, our behavior-semantic alignment module disentangles modality-agnostic behavioral patterns from noisy modality-specific features through contrastive codebook learning, ensuring semantic IDs are inherently tied to recommendation tasks. Second, we design a discretized similarity reweighting mechanism that dynamically adjusts self-attention scores using quantized semantic relationships, preserving multi-modal synergies while avoiding invasive modifications to the sequence modeling architecture. Extensive evaluations across four real-world benchmarks demonstrate BBQRec's superiority over the state-of-the-art baselines.
△ Less
Submitted 9 April, 2025;
originally announced April 2025.
-
Angular analysis of the decay $B_{s}^{0}\toφe^+e^-$
Authors:
LHCb collaboration,
R. Aaij,
A. S. W. Abdelmotteleb,
C. Abellan Beteta,
F. Abudinén,
T. Ackernley,
A. A. Adefisoye,
B. Adeva,
M. Adinolfi,
P. Adlarson,
C. Agapopoulou,
C. A. Aidala,
Z. Ajaltouni,
S. Akar,
K. Akiba,
P. Albicocco,
J. Albrecht,
F. Alessio,
Z. Aliouche,
P. Alvarez Cartelle,
R. Amalric,
S. Amato,
J. L. Amey,
Y. Amhis,
L. An
, et al. (1129 additional authors not shown)
Abstract:
An angular analysis of the decay $B^0_s\toφe^+e^-$ is presented, using proton-proton collision data collected with the LHCb detector between 2011 and 2018 at centre-of-mass energies of 7, 8 and 13$\,\mathrm{Te\kern -0.1em V}$. The combined dataset corresponds to an integrated luminosity of $9\,\mathrm{fb}^{-1}$. Observables are determined by fitting time-integrated projections of the angular distr…
▽ More
An angular analysis of the decay $B^0_s\toφe^+e^-$ is presented, using proton-proton collision data collected with the LHCb detector between 2011 and 2018 at centre-of-mass energies of 7, 8 and 13$\,\mathrm{Te\kern -0.1em V}$. The combined dataset corresponds to an integrated luminosity of $9\,\mathrm{fb}^{-1}$. Observables are determined by fitting time-integrated projections of the angular distribution in three bins of dielectron mass squared, $q^2$, corresponding to $[0.1,1.1]$, $[1.1,6.0]$ and $[15.0,19.0]\,\mathrm{Ge\kern -0.1em V}^2\!/c^4$. The results are compatible with predictions based on the Standard Model of particle physics.
△ Less
Submitted 8 April, 2025;
originally announced April 2025.
-
Transfer between Modalities with MetaQueries
Authors:
Xichen Pan,
Satya Narayan Shukla,
Aashu Singh,
Zhuokai Zhao,
Shlok Kumar Mishra,
Jialiang Wang,
Zhiyang Xu,
Jiuhai Chen,
Kunpeng Li,
Felix Juefei-Xu,
Ji Hou,
Saining Xie
Abstract:
Unified multimodal models aim to integrate understanding (text output) and generation (pixel output), but aligning these different modalities within a single architecture often demands complex training recipes and careful data balancing. We introduce MetaQueries, a set of learnable queries that act as an efficient interface between autoregressive multimodal LLMs (MLLMs) and diffusion models. MetaQ…
▽ More
Unified multimodal models aim to integrate understanding (text output) and generation (pixel output), but aligning these different modalities within a single architecture often demands complex training recipes and careful data balancing. We introduce MetaQueries, a set of learnable queries that act as an efficient interface between autoregressive multimodal LLMs (MLLMs) and diffusion models. MetaQueries connects the MLLM's latents to the diffusion decoder, enabling knowledge-augmented image generation by leveraging the MLLM's deep understanding and reasoning capabilities. Our method simplifies training, requiring only paired image-caption data and standard diffusion objectives. Notably, this transfer is effective even when the MLLM backbone remains frozen, thereby preserving its state-of-the-art multimodal understanding capabilities while achieving strong generative performance. Additionally, our method is flexible and can be easily instruction-tuned for advanced applications such as image editing and subject-driven generation.
△ Less
Submitted 8 April, 2025;
originally announced April 2025.
-
Observation of the very rare $Σ^+ \to p μ^+ μ^-$ decay
Authors:
LHCb collaboration,
R. Aaij,
A. S. W. Abdelmotteleb,
C. Abellan Beteta,
F. Abudinén,
T. Ackernley,
A. A. Adefisoye,
B. Adeva,
M. Adinolfi,
P. Adlarson,
C. Agapopoulou,
C. A. Aidala,
Z. Ajaltouni,
S. Akar,
K. Akiba,
P. Albicocco,
J. Albrecht,
F. Alessio,
Z. Aliouche,
P. Alvarez Cartelle,
R. Amalric,
S. Amato,
J. L. Amey,
Y. Amhis,
L. An
, et al. (1128 additional authors not shown)
Abstract:
The first observation of the $Σ^+ \to p μ^+ μ^-$ decay is reported with high significance using proton-proton collision data, corresponding to an integrated luminosity of $5.4\,\rm{fb}^{-1}$, collected with the LHCb detector at a centre-of-mass energy of 13~TeV. A yield of $237\pm 16$ $Σ^+ \to p μ^+ μ^-$ decays is obtained, where the uncertainty is statistical only. A branching fraction of…
▽ More
The first observation of the $Σ^+ \to p μ^+ μ^-$ decay is reported with high significance using proton-proton collision data, corresponding to an integrated luminosity of $5.4\,\rm{fb}^{-1}$, collected with the LHCb detector at a centre-of-mass energy of 13~TeV. A yield of $237\pm 16$ $Σ^+ \to p μ^+ μ^-$ decays is obtained, where the uncertainty is statistical only. A branching fraction of $(1.08 \pm 0.17) \times 10^{-8}$ is measured, where the uncertainty includes statistical and systematic sources. No evidence of resonant structures is found in the dimuon invariant-mass distribution. All results are compatible with Standard Model expectations. This represents the rarest decay of a baryon ever observed.
△ Less
Submitted 8 April, 2025;
originally announced April 2025.
-
The Quantum Double of Hopf Algebras Realized via Partial Dualization and the Tensor Category of Its Representations
Authors:
Ji-Wei He,
Xiaojie Kong,
Kangqiao Li
Abstract:
In this paper, we aim to study the (generalized) quantum double $K^{\ast\mathrm{cop}}\bowtie_σH$ determined by a (skew) pairing between finite-dimensional Hopf algebras $K^{\ast\mathrm{cop}}$ and $H$, especially the tensor category $\mathsf{Rep}(K^{\ast\mathrm{cop}}\bowtie_σH)$ of its finite-dimensional representations. Specifically, we show that $K^{\ast\mathrm{cop}}\bowtie_σH$ is a left partiall…
▽ More
In this paper, we aim to study the (generalized) quantum double $K^{\ast\mathrm{cop}}\bowtie_σH$ determined by a (skew) pairing between finite-dimensional Hopf algebras $K^{\ast\mathrm{cop}}$ and $H$, especially the tensor category $\mathsf{Rep}(K^{\ast\mathrm{cop}}\bowtie_σH)$ of its finite-dimensional representations. Specifically, we show that $K^{\ast\mathrm{cop}}\bowtie_σH$ is a left partially dualized (quasi-)Hopf algebra of $K^\mathrm{op}\otimes H$, and use this formulation to establish tensor equivalences from $\mathsf{Rep}(K^{\ast\mathrm{cop}}\bowtie_σH)$ to the categories ${}^K_K\mathcal{M}^K_H$ and ${}^{K^\ast}_{K^\ast}\mathcal{M}^{H^\ast}_{K^\ast}$ of two-sided two-cosided relative Hopf modules, as well as the category ${}_H\mathfrak{YD}^K$ of relative Yetter-Drinfeld modules.
△ Less
Submitted 8 April, 2025;
originally announced April 2025.
-
Observation of Transverse Polarization and Determination of Electromagnetic Form Factor of $Λ$ Hyperon at $\sqrt{s}= 3.773$ GeV
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (697 additional authors not shown)
Abstract:
Using a 20.3 fb$^{-1}$ of $e^{+}e^{-}$ collision data sample collected by the BESIII detector at the BEPCII collider, we present an observation of transverse polarization and a complete determination of the electromagnetic form factor of the $Λ$ hyperon in $e^{+}e^{-}\toΛ\barΛ$ decay with the entangled $Λ-\barΛ$ pair at $\sqrt{s}=3.773$ GeV. The relative phase between the electric and magnetic for…
▽ More
Using a 20.3 fb$^{-1}$ of $e^{+}e^{-}$ collision data sample collected by the BESIII detector at the BEPCII collider, we present an observation of transverse polarization and a complete determination of the electromagnetic form factor of the $Λ$ hyperon in $e^{+}e^{-}\toΛ\barΛ$ decay with the entangled $Λ-\barΛ$ pair at $\sqrt{s}=3.773$ GeV. The relative phase between the electric and magnetic form factors is determined to be $ΔΦ=(1.53\pm0.36\pm0.03)$ rad with a significance of 5.5$σ$ taking into account systematic uncertainty. This result indicates a non-zero phase between the transition amplitudes of the $Λ\barΛ$ helicity states. Additionally, we measure the angular distribution parameter and the modulus of the ratio between the electric and the magnetic form factor is found to be $η=0.86\pm0.05\pm0.03$ and $R(s)=|G_{E}(s)/G_{M}(s)|=0.47\pm0.08\pm0.05$, where the first uncertainty is statistical and the second systematic.
△ Less
Submitted 7 April, 2025;
originally announced April 2025.
-
Observation of the doubly-charmed-baryon decay ${\it Ξ}_{cc}^{++}\to{\it Ξ}_{c}^{0}π^{+}π^{+}$
Authors:
LHCb collaboration,
R. Aaij,
A. S. W. Abdelmotteleb,
C. Abellan Beteta,
F. Abudinén,
T. Ackernley,
A. A. Adefisoye,
B. Adeva,
M. Adinolfi,
P. Adlarson,
C. Agapopoulou,
C. A. Aidala,
Z. Ajaltouni,
S. Akar,
K. Akiba,
P. Albicocco,
J. Albrecht,
F. Alessio,
Z. Aliouche,
P. Alvarez Cartelle,
R. Amalric,
S. Amato,
J. L. Amey,
Y. Amhis,
L. An
, et al. (1130 additional authors not shown)
Abstract:
A search for the doubly-charmed-baryon decay ${\it Ξ}_{cc}^{++}\to{\it Ξ}_{c}^{0}(\to p K^{-}K^{-}π^{+})π^{+}π^{+}$ is performed using proton-proton collision data collected by the LHCb experiment at a centre-of-mass energy of 13 $\text{TeV}$ and corresponding to an integrated luminosity of 5.4 $\text{fb}^{-1}$. A significant structure consistent with the ${\it Ξ}_{cc}^{++}$ baryon is observed in…
▽ More
A search for the doubly-charmed-baryon decay ${\it Ξ}_{cc}^{++}\to{\it Ξ}_{c}^{0}(\to p K^{-}K^{-}π^{+})π^{+}π^{+}$ is performed using proton-proton collision data collected by the LHCb experiment at a centre-of-mass energy of 13 $\text{TeV}$ and corresponding to an integrated luminosity of 5.4 $\text{fb}^{-1}$. A significant structure consistent with the ${\it Ξ}_{cc}^{++}$ baryon is observed in the ${\it Ξ}_{c}^{0} π^{+} π^{+}$ invariant-mass spectrum. Using the ${\it Ξ}_{cc}^{++}\to{\it Λ}_{c}^{+}(\to pK^{-}π^{+})K^{-}π^{+}π^{+}$ decay as the normalisation channel, the branching fraction ratio \begin{equation*}
\frac{{\cal B}({\it Ξ}_{cc}^{++} \to {\it Ξ}_{c}^{0}π^{+}π^{+})\times{\cal B}({\it Ξ}_{c}^{0}\to pK^{-}K^{-}π^{+})} {{\cal B}({\it Ξ}_{cc}^{++}\to{\it Λ}_{c}^{+} K^{-}π^{+}π^{+})\times{\cal B}({\it Λ}_{c}^{+}\to pK^{-}π^{+})} \end{equation*} is measured to be $0.105\pm 0.014\text{\,(stat)} \pm 0.007\text{\,(syst)}$.
△ Less
Submitted 7 April, 2025;
originally announced April 2025.
-
The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models?
Authors:
Weichen Zhang,
Ruiying Peng,
Chen Gao,
Jianjie Fang,
Xin Zeng,
Kaiyuan Li,
Ziyou Wang,
Jinqiang Cui,
Xin Wang,
Xinlei Chen,
Yong Li
Abstract:
3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the role of point clouds in 3D spatial reasoning remains under-explored. In this work, we comprehensively evaluate and analyze these models to answer the research question: \textit{Does point cloud truly boost the spatial reasoning capacit…
▽ More
3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the role of point clouds in 3D spatial reasoning remains under-explored. In this work, we comprehensively evaluate and analyze these models to answer the research question: \textit{Does point cloud truly boost the spatial reasoning capacities of 3D LLMs?} We first evaluate the spatial reasoning capacity of LLMs with different input modalities by replacing the point cloud with the visual and text counterparts. We then propose a novel 3D QA (Question-answering) benchmark, ScanReQA, that comprehensively evaluates models' understanding of binary spatial relationships. Our findings reveal several critical insights: 1) LLMs without point input could even achieve competitive performance even in a zero-shot manner; 2) existing 3D LLMs struggle to comprehend the binary spatial relationships; 3) 3D LLMs exhibit limitations in exploiting the structural coordinates in point clouds for fine-grained spatial reasoning. We think these conclusions can help the next step of 3D LLMs and also offer insights for foundation models in other modalities. We release datasets and reproducible codes in the anonymous project page: https://3d-llm.xyz.
△ Less
Submitted 6 April, 2025;
originally announced April 2025.
-
Observation of $ψ(3686) \to Ξ^- K^0_S \barΩ^+ $+c.c
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (680 additional authors not shown)
Abstract:
Using a sample of $(2.712\pm0.014) \times 10^{9}$ $ψ(3686)$ events collected with the BESIII detector at the electron positron collider BEPCII, the decay $ψ(3686) \to Ξ^- K^0_S \barΩ^+ +c.c.$ is observed for the first time, which has a significance of 5.9 standard deviations. The branching fraction of this decay is measured to be $(2.91\pm0.47\pm0.33)\times 10^{-6}$, where the first and second unc…
▽ More
Using a sample of $(2.712\pm0.014) \times 10^{9}$ $ψ(3686)$ events collected with the BESIII detector at the electron positron collider BEPCII, the decay $ψ(3686) \to Ξ^- K^0_S \barΩ^+ +c.c.$ is observed for the first time, which has a significance of 5.9 standard deviations. The branching fraction of this decay is measured to be $(2.91\pm0.47\pm0.33)\times 10^{-6}$, where the first and second uncertainties are statistical and systematic, respectively. The ratio between $\mathcal{B}_{ψ(3686) \to Ξ^- K^0_S \barΩ^+ +c.c.}$ and $\mathcal{B}_{ψ(3686) \to Ω^- K^+ \barΞ^0 +c.c.}$ is determined to be $1.05\pm0.23\pm0.14 $, which deviates with the isospin symmetry conservation predicted value of 0.5 by $2.1σ$.
△ Less
Submitted 6 April, 2025;
originally announced April 2025.
-
Observation of a Three-Resonance Structure in the Cross Section of $e^+e^-\toπ^+π^- h_c$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (697 additional authors not shown)
Abstract:
Using $e^+e^-$ collision data collected with the BESIII detector operating at the Beijing Electron Positron Collider, the cross section of $e^+e^-\to π^+π^- h_c$ is measured at 59 points with center-of-mass energy $\sqrt{s}$ ranging from $4.009$ to $4.950~\mathrm{GeV}$ with a total integrated luminosity of $22.2~\mathrm{fb}^{-1}$. The cross section between $4.3$ and $4.45~\mathrm{GeV}$ exhibits a…
▽ More
Using $e^+e^-$ collision data collected with the BESIII detector operating at the Beijing Electron Positron Collider, the cross section of $e^+e^-\to π^+π^- h_c$ is measured at 59 points with center-of-mass energy $\sqrt{s}$ ranging from $4.009$ to $4.950~\mathrm{GeV}$ with a total integrated luminosity of $22.2~\mathrm{fb}^{-1}$. The cross section between $4.3$ and $4.45~\mathrm{GeV}$ exhibits a plateau-like shape and drops sharply around $4.5~\mathrm{GeV}$, which cannot be described by two resonances only. Three coherent Breit-Wigner functions are used to parameterize the $\sqrt{s}$-dependent cross section line shape. The masses and widths are determined to be $M_1=(4223.6_{-3.7-2.9}^{+3.6+2.6})~\mathrm{MeV}/c^2$, $Γ_1=(58.5_{-11.4-6.5}^{+10.8+6.7})~\mathrm{MeV}$, $M_2=(4327.4_{-18.8-9.3}^{+20.1+10.7})~\mathrm{MeV}/c^2$, $Γ_2=(244.1_{-27.1-18.0}^{+34.0+23.9})~\mathrm{MeV}$, and $M_3=(4467.4_{-5.4-2.7}^{+7.2+3.2})~\mathrm{MeV}/c^2$, $Γ_3=(62.8_{-14.4-6.6}^{+19.2+9.8})~\mathrm{MeV}$. The first uncertainties are statistical and the other two are systematic. The statistical significance of the three Breit-Wigner assumption over the two Breit-Wigner assumption is greater than $5σ$.
△ Less
Submitted 5 April, 2025;
originally announced April 2025.
-
Mapping at First Sense: A Lightweight Neural Network-Based Indoor Structures Prediction Method for Robot Autonomous Exploration
Authors:
Haojia Gao,
Haohua Que,
Kunrong Li,
Weihao Shan,
Mingkai Liu,
Rong Zhao,
Lei Mu,
Xinghua Yang,
Qi Wei,
Fei Qiao
Abstract:
Autonomous exploration in unknown environments is a critical challenge in robotics, particularly for applications such as indoor navigation, search and rescue, and service robotics. Traditional exploration strategies, such as frontier-based methods, often struggle to efficiently utilize prior knowledge of structural regularities in indoor spaces. To address this limitation, we propose Mapping at F…
▽ More
Autonomous exploration in unknown environments is a critical challenge in robotics, particularly for applications such as indoor navigation, search and rescue, and service robotics. Traditional exploration strategies, such as frontier-based methods, often struggle to efficiently utilize prior knowledge of structural regularities in indoor spaces. To address this limitation, we propose Mapping at First Sense, a lightweight neural network-based approach that predicts unobserved areas in local maps, thereby enhancing exploration efficiency. The core of our method, SenseMapNet, integrates convolutional and transformerbased architectures to infer occluded regions while maintaining computational efficiency for real-time deployment on resourceconstrained robots. Additionally, we introduce SenseMapDataset, a curated dataset constructed from KTH and HouseExpo environments, which facilitates training and evaluation of neural models for indoor exploration. Experimental results demonstrate that SenseMapNet achieves an SSIM (structural similarity) of 0.78, LPIPS (perceptual quality) of 0.68, and an FID (feature distribution alignment) of 239.79, outperforming conventional methods in map reconstruction quality. Compared to traditional frontier-based exploration, our method reduces exploration time by 46.5% (from 2335.56s to 1248.68s) while maintaining a high coverage rate (88%) and achieving a reconstruction accuracy of 88%. The proposed method represents a promising step toward efficient, learning-driven robotic exploration in structured environments.
△ Less
Submitted 5 April, 2025;
originally announced April 2025.
-
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
Authors:
NVIDIA,
:,
Aaron Blakeman,
Aarti Basant,
Abhinav Khattar,
Adithya Renduchintala,
Akhiad Bercovich,
Aleksander Ficek,
Alexis Bjorlin,
Ali Taghibakhshi,
Amala Sanjay Deshmukh,
Ameya Sunil Mahabaleshwarkar,
Andrew Tao,
Anna Shors,
Ashwath Aithal,
Ashwin Poojary,
Ayush Dattagupta,
Balaram Buddharaju,
Bobby Chen,
Boris Ginsburg,
Boxin Wang,
Brandon Norick,
Brian Butterfield,
Bryan Catanzaro,
Carlo del Mundo
, et al. (176 additional authors not shown)
Abstract:
As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer models designed to reduce inference cost for a given accuracy level. To achieve this goal, we replace the majority of self-attention layers in the common Transf…
▽ More
As inference-time scaling becomes critical for enhanced reasoning capabilities, it is increasingly becoming important to build models that are efficient to infer. We introduce Nemotron-H, a family of 8B and 56B/47B hybrid Mamba-Transformer models designed to reduce inference cost for a given accuracy level. To achieve this goal, we replace the majority of self-attention layers in the common Transformer model architecture with Mamba layers that perform constant computation and require constant memory per generated token. We show that Nemotron-H models offer either better or on-par accuracy compared to other similarly-sized state-of-the-art open-sourced Transformer models (e.g., Qwen-2.5-7B/72B and Llama-3.1-8B/70B), while being up to 3$\times$ faster at inference. To further increase inference speed and reduce the memory required at inference time, we created Nemotron-H-47B-Base from the 56B model using a new compression via pruning and distillation technique called MiniPuzzle. Nemotron-H-47B-Base achieves similar accuracy to the 56B model, but is 20% faster to infer. In addition, we introduce an FP8-based training recipe and show that it can achieve on par results with BF16-based training. This recipe is used to train the 56B model. We are releasing Nemotron-H base model checkpoints with support in Hugging Face and NeMo.
△ Less
Submitted 15 April, 2025; v1 submitted 4 April, 2025;
originally announced April 2025.
-
PF3Det: A Prompted Foundation Feature Assisted Visual LiDAR 3D Detector
Authors:
Kaidong Li,
Tianxiao Zhang,
Kuan-Chuan Peng,
Guanghui Wang
Abstract:
3D object detection is crucial for autonomous driving, leveraging both LiDAR point clouds for precise depth information and camera images for rich semantic information. Therefore, the multi-modal methods that combine both modalities offer more robust detection results. However, efficiently fusing LiDAR points and images remains challenging due to the domain gaps. In addition, the performance of ma…
▽ More
3D object detection is crucial for autonomous driving, leveraging both LiDAR point clouds for precise depth information and camera images for rich semantic information. Therefore, the multi-modal methods that combine both modalities offer more robust detection results. However, efficiently fusing LiDAR points and images remains challenging due to the domain gaps. In addition, the performance of many models is limited by the amount of high quality labeled data, which is expensive to create. The recent advances in foundation models, which use large-scale pre-training on different modalities, enable better multi-modal fusion. Combining the prompt engineering techniques for efficient training, we propose the Prompted Foundational 3D Detector (PF3Det), which integrates foundation model encoders and soft prompts to enhance LiDAR-camera feature fusion. PF3Det achieves the state-of-the-art results under limited training data, improving NDS by 1.19% and mAP by 2.42% on the nuScenes dataset, demonstrating its efficiency in 3D detection.
△ Less
Submitted 4 April, 2025;
originally announced April 2025.
-
FontGuard: A Robust Font Watermarking Approach Leveraging Deep Font Knowledge
Authors:
Kahim Wong,
Jicheng Zhou,
Kemou Li,
Yain-Whar Si,
Xiaowei Wu,
Jiantao Zhou
Abstract:
The proliferation of AI-generated content brings significant concerns on the forensic and security issues such as source tracing, copyright protection, etc, highlighting the need for effective watermarking technologies. Font-based text watermarking has emerged as an effective solution to embed information, which could ensure copyright, traceability, and compliance of the generated text content. Ex…
▽ More
The proliferation of AI-generated content brings significant concerns on the forensic and security issues such as source tracing, copyright protection, etc, highlighting the need for effective watermarking technologies. Font-based text watermarking has emerged as an effective solution to embed information, which could ensure copyright, traceability, and compliance of the generated text content. Existing font watermarking methods usually neglect essential font knowledge, which leads to watermarked fonts of low quality and limited embedding capacity. These methods are also vulnerable to real-world distortions, low-resolution fonts, and inaccurate character segmentation. In this paper, we introduce FontGuard, a novel font watermarking model that harnesses the capabilities of font models and language-guided contrastive learning. Unlike previous methods that focus solely on the pixel-level alteration, FontGuard modifies fonts by altering hidden style features, resulting in better font quality upon watermark embedding. We also leverage the font manifold to increase the embedding capacity of our proposed method by generating substantial font variants closely resembling the original font. Furthermore, in the decoder, we employ an image-text contrastive learning to reconstruct the embedded bits, which can achieve desirable robustness against various real-world transmission distortions. FontGuard outperforms state-of-the-art methods by +5.4%, +7.4%, and +5.8% in decoding accuracy under synthetic, cross-media, and online social network distortions, respectively, while improving the visual quality by 52.7% in terms of LPIPS. Moreover, FontGuard uniquely allows the generation of watermarked fonts for unseen fonts without re-training the network. The code and dataset are available at https://github.com/KAHIMWONG/FontGuard.
△ Less
Submitted 3 April, 2025;
originally announced April 2025.
-
Superconductivity in the Medium-Entropy/High-Entropy Re-based Alloys with a Non-Centrosymmetric $α$-Mn Lattice
Authors:
Kuan Li,
Longfu Li,
Lingyong Zenga,
Yucheng Li,
Rui Chen,
Peifeng Yu,
Kangwang Wang,
Zaichen Xiang,
Tian Shang,
Huixia Luo
Abstract:
Medium or high-entropy alloys (MEAs-HEAs) and rhenium-based compounds with a non-centrosymmetric (NC) structure have received a lot of attention for offering a fertile soil in search for unconventional superconductivity. Here, five previously unreported NC Re-based MEA-HEA superconductors with an $α$-Mn lattice are successfully synthesized, with their superconducting transition temperatures (Tcs)…
▽ More
Medium or high-entropy alloys (MEAs-HEAs) and rhenium-based compounds with a non-centrosymmetric (NC) structure have received a lot of attention for offering a fertile soil in search for unconventional superconductivity. Here, five previously unreported NC Re-based MEA-HEA superconductors with an $α$-Mn lattice are successfully synthesized, with their superconducting transition temperatures (Tcs) ranging from 4 to 5 K. An increase in the superconducting transition temperature (Tc) can be achieved by modulating the valence electron count (VEC) through compositional adjustments. Magnetization measurements confirm that all the synthesized Re-based MEA-HEAs are bulk type-II superconductors. Specific heat analysis reveals that the superconducting state of these HEAs can be well described by a single-gap s-wave model. Our results show that the Kadowaki-Woods ratio of these $α$-Mn MEA/HEA superconductors are close to the typical value of heavy fermion compounds, suggesting the existence of strong electronic correlation. These findings provide promising material platforms to study the role of high disorder in the origin of superconductivity in the NC MEAs-HEAs.
△ Less
Submitted 3 April, 2025;
originally announced April 2025.
-
Evidence of doubly OZI-suppressed decay $η_{c} \to ωφ$ in the radiative decay $J/ψ\to γη_{c}$
Authors:
BESIII Collaboration,
M. Ablikim,
M. N. Achasov,
P. Adlarson,
X. C. Ai,
R. Aliberti,
A. Amoroso,
Q. An,
Y. Bai,
O. Bakina,
Y. Ban,
H. -R. Bao,
V. Batozskaya,
K. Begzsuren,
N. Berger,
M. Berlowski,
M. Bertani,
D. Bettoni,
F. Bianchi,
E. Bianco,
A. Bortone,
I. Boyko,
R. A. Briere,
A. Brueggemann,
H. Cai
, et al. (680 additional authors not shown)
Abstract:
Using a sample of $(10087\pm44) \times 10^{6}$ $J/ψ$ events collected with the BESIII detector at the BEPCII collider, the first evidence for the doubly OZI-suppressed decay $η_{c} \to ωφ$ is reported with a significance of 4.0$σ$. The branching fraction of $η_{c} \to ωφ$ is measured to be $\mathcal{B}(η_{c} \to ωφ) = (3.86 \pm 0.92 \pm 0.62) \times 10^{-5}$, where the first uncertainty is statist…
▽ More
Using a sample of $(10087\pm44) \times 10^{6}$ $J/ψ$ events collected with the BESIII detector at the BEPCII collider, the first evidence for the doubly OZI-suppressed decay $η_{c} \to ωφ$ is reported with a significance of 4.0$σ$. The branching fraction of $η_{c} \to ωφ$ is measured to be $\mathcal{B}(η_{c} \to ωφ) = (3.86 \pm 0.92 \pm 0.62) \times 10^{-5}$, where the first uncertainty is statistical and the second is systematic. This result provides valuable insights into the underlying mechanisms of charmonium decays, particularly for processes such as $η_{c} \to VV$ (where $V$ represents a vector meson).
△ Less
Submitted 2 April, 2025;
originally announced April 2025.
-
A 0.3\% calibration of the W UMa-type contact binary luminosity based on Gaia DR3
Authors:
Jing Li,
Xiaodian Chen,
Shu Wang,
Kun Wang,
Kai Li,
Licai Deng
Abstract:
W Ursa Majoris (W UMa)-type contact binary systems (CBs) with period--luminosity (PL) relations are valuable distance indicators. The PL relations of CBs are affected by metallicity. Here, we establish PL relations and period--luminosity--metallicity (PLZ) relations in nine bands from optical to mid-infrared ($BP$, $G$, $RP$, $J$, $H$, $K_S$, $W1$, $W2$, $W3$) and in five Wesenheit bands based on…
▽ More
W Ursa Majoris (W UMa)-type contact binary systems (CBs) with period--luminosity (PL) relations are valuable distance indicators. The PL relations of CBs are affected by metallicity. Here, we establish PL relations and period--luminosity--metallicity (PLZ) relations in nine bands from optical to mid-infrared ($BP$, $G$, $RP$, $J$, $H$, $K_S$, $W1$, $W2$, $W3$) and in five Wesenheit bands based on Gaia DR3 parallaxes. The dispersion of PLZ relations gradually decreases from the optical to mid-infrared bands, with the minimum dispersion of 0.138 mag. We fit the best PL relations for three bands ($W1$, $W_{G,BP,RP}$, $W_{W1,BP,RP}$) under different parallax uncertainty criteria and determine a parallax (after correction) zero point of $zp_\varpi=24\pm4$ $μ$as. After fixing the parallax zero point, we find that the total zero errors of the PL and PLZ relation are smallest when the parallax uncertainty is less than 2\%, resulting in a calibrated CB luminosity with an accuracy of 0.33\%, which is more accurate than the classical Cepheids. Furthermore, by examining the absolute magnitude residuals in different metallicity intervals, we find that the use of a linear metallicity effect is appropriate for CBs with different metallicities. These results indicate that CBs are excellent standard candles.
△ Less
Submitted 2 April, 2025;
originally announced April 2025.
-
From Easy to Hard: Building a Shortcut for Differentially Private Image Synthesis
Authors:
Kecen Li,
Chen Gong,
Xiaochen Li,
Yuzhong Zhao,
Xinwen Hou,
Tianhao Wang
Abstract:
Differentially private (DP) image synthesis aims to generate synthetic images from a sensitive dataset, alleviating the privacy leakage concerns of organizations sharing and utilizing synthetic images. Although previous methods have significantly progressed, especially in training diffusion models on sensitive images with DP Stochastic Gradient Descent (DP-SGD), they still suffer from unsatisfacto…
▽ More
Differentially private (DP) image synthesis aims to generate synthetic images from a sensitive dataset, alleviating the privacy leakage concerns of organizations sharing and utilizing synthetic images. Although previous methods have significantly progressed, especially in training diffusion models on sensitive images with DP Stochastic Gradient Descent (DP-SGD), they still suffer from unsatisfactory performance. In this work, inspired by curriculum learning, we propose a two-stage DP image synthesis framework, where diffusion models learn to generate DP synthetic images from easy to hard. Unlike existing methods that directly use DP-SGD to train diffusion models, we propose an easy stage in the beginning, where diffusion models learn simple features of the sensitive images. To facilitate this easy stage, we propose to use `central images', simply aggregations of random samples of the sensitive dataset. Intuitively, although those central images do not show details, they demonstrate useful characteristics of all images and only incur minimal privacy costs, thus helping early-phase model training. We conduct experiments to present that on the average of four investigated image datasets, the fidelity and utility metrics of our synthetic images are 33.1% and 2.1% better than the state-of-the-art method.
△ Less
Submitted 2 April, 2025;
originally announced April 2025.
-
Towards Resilient Federated Learning in CyberEdge Networks: Recent Advances and Future Trends
Authors:
Kai Li,
Zhengyang Zhang,
Azadeh Pourkabirian,
Wei Ni,
Falko Dressler,
Ozgur B. Akan
Abstract:
In this survey, we investigate the most recent techniques of resilient federated learning (ResFL) in CyberEdge networks, focusing on joint training with agglomerative deduction and feature-oriented security mechanisms. We explore adaptive hierarchical learning strategies to tackle non-IID data challenges, improving scalability and reducing communication overhead. Fault tolerance techniques and agg…
▽ More
In this survey, we investigate the most recent techniques of resilient federated learning (ResFL) in CyberEdge networks, focusing on joint training with agglomerative deduction and feature-oriented security mechanisms. We explore adaptive hierarchical learning strategies to tackle non-IID data challenges, improving scalability and reducing communication overhead. Fault tolerance techniques and agglomerative deduction mechanisms are studied to detect unreliable devices, refine model updates, and enhance convergence stability. Unlike existing FL security research, we comprehensively analyze feature-oriented threats, such as poisoning, inference, and reconstruction attacks that exploit model features. Moreover, we examine resilient aggregation techniques, anomaly detection, and cryptographic defenses, including differential privacy and secure multi-party computation, to strengthen FL security. In addition, we discuss the integration of 6G, large language models (LLMs), and interoperable learning frameworks to enhance privacy-preserving and decentralized cross-domain training. These advancements offer ultra-low latency, artificial intelligence (AI)-driven network management, and improved resilience against adversarial attacks, fostering the deployment of secure ResFL in CyberEdge networks.
△ Less
Submitted 1 April, 2025;
originally announced April 2025.
-
Using machine learning method for variable star classification using the TESS Sectors 1-57 data
Authors:
Li-Heng Wang,
Kai Li,
Xiang Gao,
Ya-Ni Guo,
Guo-You Sun
Abstract:
The Transiting Exoplanet Survey Satellite (TESS) is a wide-field all-sky survey mission designed to detect Earth-sized exoplanets. After over four years photometric surveys, data from sectors 1-57, including approximately 1,050,000 light curves with a 2-minute cadence, were collected. By cross-matching the data with Gaia's variable star catalogue, we obtained labeled datasets for further analysis.…
▽ More
The Transiting Exoplanet Survey Satellite (TESS) is a wide-field all-sky survey mission designed to detect Earth-sized exoplanets. After over four years photometric surveys, data from sectors 1-57, including approximately 1,050,000 light curves with a 2-minute cadence, were collected. By cross-matching the data with Gaia's variable star catalogue, we obtained labeled datasets for further analysis. Using a random forest classifier, we performed classification of variable stars and designed distinct classification processes for each subclass, 6770 EA, 2971 EW, 980 CEP, 8347 DSCT, 457 RRab, 404 RRc and 12348 ROT were identified. Each variable star was visually inspected to ensure the reliability and accuracy of the compiled catalog. Subsequently, we ultimately obtained 6046 EA, 3859 EW, 2058 CEP, 8434 DSCT, 482 RRab, 416 RRc, and 9694 ROT, and a total of 14092 new variable stars were discovered.
△ Less
Submitted 31 March, 2025;
originally announced April 2025.
-
DOMAC: Differentiable Optimization for High-Speed Multipliers and Multiply-Accumulators
Authors:
Chenhao Xue,
Yi Ren,
Jinwei Zhou,
Kezhi Li,
Chen Zhang,
Yibo Lin,
Lining Zhang,
Qiang Xu,
Guangyu Sun
Abstract:
Multipliers and multiply-accumulators (MACs) are fundamental building blocks for compute-intensive applications such as artificial intelligence. With the diminishing returns of Moore's Law, optimizing multiplier performance now necessitates process-aware architectural innovations rather than relying solely on technology scaling. In this paper, we introduce DOMAC, a novel approach that employs diff…
▽ More
Multipliers and multiply-accumulators (MACs) are fundamental building blocks for compute-intensive applications such as artificial intelligence. With the diminishing returns of Moore's Law, optimizing multiplier performance now necessitates process-aware architectural innovations rather than relying solely on technology scaling. In this paper, we introduce DOMAC, a novel approach that employs differentiable optimization for designing multipliers and MACs at specific technology nodes. DOMAC establishes an analogy between optimizing multi-staged parallel compressor trees and training deep neural networks. Building on this insight, DOMAC reformulates the discrete optimization challenge into a continuous problem by incorporating differentiable timing and area objectives. This formulation enables us to utilize existing deep learning toolkit for highly efficient implementation of the differentiable solver. Experimental results demonstrate that DOMAC achieves significant enhancements in both performance and area efficiency compared to state-of-the-art baselines and commercial IPs in multiplier and MAC designs.
△ Less
Submitted 31 March, 2025;
originally announced March 2025.
-
A high-fidelity surrogate model for the ion temperature gradient (ITG) instability using a small expensive simulation dataset
Authors:
Chenguang Wan,
Youngwoo Cho,
Zhisong Qu,
Yann Camenen,
Robin Varennes,
Kyungtak Lim,
Kunpeng Li,
Jiangang Li,
Yanlong Li,
Xavier Garbet
Abstract:
One of the main challenges in building high-fidelity surrogate models of tokamak turbulence is the substantial demand for high-quality data. Typically, producing high-quality data involves simulating complex physical processes, which requires extensive computing resources. In this work, we propose a fine tuning-based approach to develop the surrogate model that reduces the amount of high-quality d…
▽ More
One of the main challenges in building high-fidelity surrogate models of tokamak turbulence is the substantial demand for high-quality data. Typically, producing high-quality data involves simulating complex physical processes, which requires extensive computing resources. In this work, we propose a fine tuning-based approach to develop the surrogate model that reduces the amount of high-quality data required by 80\%. We demonstrate the effectiveness of this approach by constructing a proof-of-principle ITG surrogate model using datasets generated from two gyrokinetic codes, GKW and GX. GX needs in terms of computing resources are much lighter than GKW. Remarkably, the surrogate models' performance remain nearly the same whether trained on 798 GKW results alone or 159 GKW results plus an additional 11979 GX results. These encouraging outcomes indicate that fine tuning methods can significantly decrease the high-quality data needed to develop the simulation-driven surrogate model. Moreover, the approach presented here has the potential to facilitate surrogate model development for heavy codes and may ultimately pave the way for digital twin systems of tokamaks.
△ Less
Submitted 30 March, 2025;
originally announced March 2025.
-
SCORE: Story Coherence and Retrieval Enhancement for AI Narratives
Authors:
Qiang Yi,
Yangfan He,
Jianhui Wang,
Xinyuan Song,
Shiyao Qian,
Xinhang Yuan,
Miao Zhang,
Li Sun,
Keqin Li,
Kuan Lu,
Menghao Huo,
Jiaqi Chen,
Tianyu Shi
Abstract:
Large Language Models (LLMs) can generate creative and engaging narratives from user-specified input, but maintaining coherence and emotional depth throughout these AI-generated stories remains a challenge. In this work, we propose SCORE, a framework for Story Coherence and Retrieval Enhancement, designed to detect and resolve narrative inconsistencies. By tracking key item statuses and generating…
▽ More
Large Language Models (LLMs) can generate creative and engaging narratives from user-specified input, but maintaining coherence and emotional depth throughout these AI-generated stories remains a challenge. In this work, we propose SCORE, a framework for Story Coherence and Retrieval Enhancement, designed to detect and resolve narrative inconsistencies. By tracking key item statuses and generating episode summaries, SCORE uses a Retrieval-Augmented Generation (RAG) approach, incorporating TF-IDF and cosine similarity to identify related episodes and enhance the overall story structure. Results from testing multiple LLM-generated stories demonstrate that SCORE significantly improves the consistency and stability of narrative coherence compared to baseline GPT models, providing a more robust method for evaluating and refining AI-generated narratives.
△ Less
Submitted 21 April, 2025; v1 submitted 30 March, 2025;
originally announced March 2025.