Aligning AI System with Human Preference: Personalization, Robustification, and Evaluation
Description
This dissertation studies the problem of aligning artificial intelligence systems with human preferences through three fundamentally connected dimensions: personalization, robustification, and evaluation. Modern AI systems increasingly operate in interactive environments where they must adapt to diverse users, learn from imperfect feedback, and continuously assess the quality and alignment of their own behaviors. These challenges are intrinsically sequential: effective personalization requires learning from human interactions; such interaction-driven systems must remain reliable under noisy or adversarial feedback, motivating robustification; and understanding whether these systems truly align with human preferences ultimately requires principled evaluation frameworks. Addressing this pipeline demands both rigorous modeling and learning algorithms with strong theoretical guarantees.
First, from the perspective of personalization, this dissertation develops personalized decision-making frameworks based on contextual reinforcement learning, with applications such as auto-bidding systems in digital advertising. A central challenge in these environments is learning from delayed and cumulative user responses while adapting to heterogeneous preferences across individuals and contexts. We propose learning algorithms that effectively balance exploration and exploitation under delayed feedback and state-dependent effects, achieving near-optimal theoretical guarantees and strong empirical performance. These results demonstrate how AI systems can adapt their behaviors to individual users in complex interactive environments.
However, personalization necessarily relies on feedback collected from human interactions, which is often noisy, inconsistent, strategic, or even adversarial. This motivates the second part of the dissertation on robustification. We investigate learning under imperfect and corrupted human feedback and establish an intrinsic efficiency--robustness trade-off showing that algorithms with faster statistical efficiency are generally more vulnerable to adversarial or corrupted observations. This vulnerability becomes even more pronounced in multi-agent systems, where the manipulation of a single agent can substantially alter global interaction dynamics and collective behavior. Our results reveal how localized attacks may propagate system-wide and motivate the design of defense-aware and corruption-resilient learning algorithms.
Finally, once AI systems become adaptive and robust, a fundamental question remains: how should we evaluate whether their behaviors are truly informative, diverse, and aligned with human preferences? To address this challenge, the final part of this dissertation studies information conveyance in the generation space of large language models and other generative AI systems. We introduce principled information-theoretic measures and statistically consistent estimators, grounded in a tree-based view of language generation, to characterize how uncertainty and semantic diversity are distributed across model outputs. Our analysis provides new insights into how fine-tuning reshapes generation uncertainty into more informative, semantically meaningful, and human-aligned behaviors, rather than merely reducing diversity. Together, these contributions establish a unified perspective on aligning AI systems with human preferences through adaptive personalization, robust learning under imperfect feedback, and principled evaluation of generative behaviors.
Files
Yuwei_Dissertation_Revised_Final_Draft (1).pdf
Files
(3.7 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:178a2228088a06a644bd4d4f0b629fc7
|
3.7 MB | Preview Download |
Additional details
Dates
- Submitted
-
2026-08Dissertation