Published December 2025 | Version v1
Dissertation Open

Self-Play Methods in Reinforcement Learning for Language Models

Creators

  • 1. University of Chicago

Contributors

Advisor:

Committee members:

Description

In this thesis we develop a series of practical algorithms for language model to self-train, by actively and strategically creating and controlling learning experiences themselves, enabling better model generalizability and training efficiency compared to traditional approaches.

Files

thesis_self_play_final.pdf

Files (1.5 MB)

Name Size Download all
md5:05271764d744a6518eafb9d1f056cfc3
1.5 MB Preview Download

Additional details

Identifiers

Other
oai:uchicago.tind.io:16624

UChicago Information

Division(s)
Physical Sciences Division
Department(s)
Computer Science