0% Complete
صفحه اصلی
/
دومین همایش بین المللی هوش مصنوعی
From RLHF to DPO and Beyond: A Review of Preference Optimization Methods for Aligning Large Language Models
نویسندگان :
Poorya Saneei
1
Parham Rahimi
2
Behrouz Minaei-Bidgoli
3
1- دانشگاه علم و صنعت ایران
2- دانشگاه علم و صنعت ایران
3- دانشگاه علم و صنعت ایران
کلمات کلیدی :
Large Language Models،LLM Alignment،Preference Optimization،Reinforcement Learning from Human Feedback (RLHF)،Direct Preference Optimization (DPO)،Reward Model،Reference-Free Algorithms
چکیده :
LLMs are not aligned with human values by default, and this challenge has fostered the development of various alignment methods to make models helpful, honest, and harmless. One of the first practical methods for solving this problem was Reinforcement Learning from Human Feedback (RLHF), but its complexity, instability, and high computational cost have prevented it from large-scale application. Since then, many alignment techniques have been suggested, but a consistent review of their evolution is missing. This paper traces the evolution of the alignment methods, focusing on how it evolved from RLHF to the family of algorithms called Direct Preference Optimization, or simply DPO. It elegantly reparametrizes the alignment problem into a supervised learning problem. Further, we consider how the new generation of algorithms overcame the limitations of DPO along three major axes: better computational efficiency by removing the reference model, flexibility of formats and objectives, and optimization for complex goals such as creativity. This paper provides an overview of how alignment research has moved from complexity to simple, efficient, and modular algorithms and shows what the future of alignment will be like.
لیست مقالات
لیست مقالات بایگانی شده
Efficient DL Model for Voice Pathology Detection in Healthcare Applications using Sustained Vowels
Sahar Farazi - Yasser Shekofteh
Beyond Sequences: A Benchmark for Atomic Hand-Object Interaction Using a Static RNN Encoder
Yousef Azizi Movahed - Fatemeh Ziaeetabar
Soil Shear Strength Prediction Using Genetic-Optimized Boosting Machine Learning Models
Mohammadreza Ghadami - Ali Noorzad - Hamid Mohammadnezhad
Interpretable Machine Learning for Rocking-Induced Settlement Prediction Using SHAP Analysis
Seyed Emad Miri - Hamid Mohammadnezhad
Technical Analyst Attention Network (TAAN): An Interpretable Deep Learning Model for Algorithmic Trading in Crypto and Forex Markets
Ali Tavassolian - Fatemeh Yousefloei - Mojtaba Vahidi Asl - Monireh Abdoos
Development of 3D Neural Cellular Automata for Learning Spatio-Temporal Patterns
Yashar Rezazadeh Shahir - َُAkram Beigi
AI-Powered Beauty: Innovations, Transformations, and Ethical Considerations
Rana Poureskandar - Abbas Mirzaei - Babak Nouri-Moghaddam
A Thorough Analysis of How Chatbots Engage, with Aspects of Customer Experience; An In depth Review
Omid Noori
Cross-domain prediction of ultimate strength in corroded reinforced concrete columns using domain adaptation techniques
Melina Sadeghi - Pooria Poorahad - Mahmoud R. Shiravand
ChatGPT 4 and personality prediction
Zahra Eslami - Marcus Cheetham - Seyed Abolfazl Valizadeh
بیشتر
ثمین همایش، سامانه مدیریت کنفرانس ها و جشنواره ها - نگارش 44.5.0