Published March 2026
| Version v1
Dissertation
Open
Steering Model Robustness via Minimal Training Data Modification
Description
Data has been the core for training machine learning models because it provides the foundation where the model learns patterns, features and structures.Recent advances in machine learning have resulted in models of increasing size and complexity. For example, recent generative models are trained on billions of samples and feature billions of parameters. Given the massive amounts of training data, many assume that these large models are inherently robust to training-time attacks, because it would require modifying a significant portion of the training data to compromise the model's robustness. This dissertation challenges this prevailing assumption, and asks: "Is it possible to steer a model's security behavior by injecting minimal yet strategically optimized data to its training data?" This dissertation addresses this question by empirically validating and analytically verifying its feasibility through defenses on deep neural networks (DNNs) and attacks on large text-to-image generative models. For DNNs, existing work proposes to include intentionally modified samples in the training data to defend against inference-time attacks. However, the fundamental tension between robustness and accuracy remained unsolved. Moreover, practitioners do not have any practical tool to flexibly control the model's security property once a deployed model is breached. Towards these, my work develops a theoretical understanding on the robustness-accuracy tradeoff in training-time defenses by characterizing the optimal loss in training robust models. To further improve model robustness post deployment, I propose a fast and robust model versioning mechanism through injecting minimal task-irrelevant data to the training data, which allows model owners to recover from model breaches. My work also explores the feasibility of manipulating performance of generative models through poisoning attacks against large text-to-image models. My work shows that large text-to-image models, although trained on billions of samples, are surprisingly vulnerable to low-volume optimized attacks against specific prompt during training. My research analyzes the culprit of such vulnerability, which lies in the tension between model's architecture design and data complexity. This line of work has turned into a practical protection tool for human creatives against training on unauthorized data. This dissertation provides an understanding of model robustness under optimized training data modification from both empirical studies and theoretical analysis. The feasibility to steer model robustness with minimal data enables continued control for model owners and proactive protection for data contributors.
Files
DingWenxin_dissertation.pdf
Files
(29.8 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:56d99d9570bbe9b95100fc2833ebcc19
|
29.8 MB | Preview Download |
Additional details
Identifiers
- Other
- oai:uchicago.tind.io:16767