Published December 2025
| Version v1
Dissertation
Open
Self-Play Methods in Reinforcement Learning for Language Models
Description
In this thesis we develop a series of practical algorithms for language model to self-train, by actively and strategically creating and controlling learning experiences themselves, enabling better model generalizability and training efficiency compared to traditional approaches.
Files
thesis_self_play_final.pdf
Files
(1.5 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:05271764d744a6518eafb9d1f056cfc3
|
1.5 MB | Preview Download |
Additional details
Identifiers
- Other
- oai:uchicago.tind.io:16624