Self-Improving Robust Preference Optimization

Choi, Eugene; Ahmadian, Arash; Geist, Matthieu; Pietquin, Oilvier; Azar, Mohammad Gheshlaghi

Computer Science > Machine Learning

arXiv:2406.01660 (cs)

[Submitted on 3 Jun 2024 (v1), last revised 11 Apr 2025 (this version, v4)]

Title:Self-Improving Robust Preference Optimization

Authors:Eugene Choi, Arash Ahmadian, Matthieu Geist, Oilvier Pietquin, Mohammad Gheshlaghi Azar

View PDF HTML (experimental)

Abstract:Online and offline RLHF methods, such as PPO and DPO, have been highly successful in aligning AI with human preferences. Despite their success, however, these methods suffer from fundamental limitations: (a) Models trained with RLHF can learn from mistakes or negative examples through RL mechanism or contrastive loss during training. However, at inference time, they lack an innate self-improvement mechanism for error corrections. (b) The optimal solution of existing methods is highly task-dependent, making it difficult for them to generalize to new tasks. To address these challenges, we propose Self-Improving Robust Preference Optimization (SRPO), a practical and mathematically principled offline RLHF framework. The key idea behind SRPO is to cast the problem of learning from human preferences as a self-improvement process, mathematically formulated as a min-max objective that jointly optimizes a self-improvement policy and a generative policy in an adversarial fashion. Crucially, the solution for this optimization problem is independent of the training task, which makes it robust to its changes. We then show that this objective can be reformulated as a non-adversarial offline loss, which can be efficiently optimized using standard supervised learning techniques at scale. To demonstrate SRPO's effectiveness, we evaluate it using AI Win-Rate (WR) against human (GOLD) completions. When tested on the XSum dataset, SRPO outperforms DPO by a margin of 15% after 5 self revisions, achieving an impressive 90% WR. Moreover, on the challenging Arena-Hard prompts, SRPO outperforms both DPO and IPO (by 4% without revision and 6% after a single revision), reaching a 56% WR against against Llama-3.1-8B-Instruct.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:2406.01660 [cs.LG]
	(or arXiv:2406.01660v4 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2406.01660

Submission history

From: Mohammad Gheshlaghi Azar [view email]
[v1] Mon, 3 Jun 2024 17:53:25 UTC (124 KB)
[v2] Wed, 5 Jun 2024 01:25:34 UTC (137 KB)
[v3] Fri, 7 Jun 2024 17:25:12 UTC (137 KB)
[v4] Fri, 11 Apr 2025 23:24:37 UTC (149 KB)

Computer Science > Machine Learning

Title:Self-Improving Robust Preference Optimization

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Self-Improving Robust Preference Optimization

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators