Defining and Measuring Qualities of Language for Human-Centered AI
Description
Natural language is a fundamental medium of communication, across humans and machines. It is often subjective and contextual, meaning that many language-based tasks lack singular ground truth answers. However, reliable analysis and evaluation of such data and tasks are still crucial, especially in high-stakes settings with direct impact on humans. This thesis approaches these challenges by grounding the measurement of subjective concepts in human signals of domain knowledge and building frameworks that leverage multiple, complementary dimensions for evaluation. We first present a suite of metrics to quantify rhetorical distinctiveness and divisiveness in political discourse. By combining large language models, lexical resources, and graph-based analyses, we characterize how U.S. presidents differ in linguistic style and use of antagonistic language in their public speeches. The second study introduces an approach for distilling large-scale clinician feedback into structured, LLM-enforceable checklist items, in order to define criteria for clinical note quality. Used with LLM-as-a-Judge, these checklists yield interpretable and auditable criteria that demonstrate alignment with expert judgment and outperform a zero-shot baseline in coverage, diversity, and predictive power. We incorporate this pipeline into an open-source library that unifies multiple approaches for automatic checklist generation and scoring from the literature, to support extensibility of rubric-based evaluation. Finally, we assess the impact of human-AI conversation on voter engagement and reasoning around ballot measures. Through controlled user studies, we observe considerable differences in vote distributions across treatment groups, with discourse-oriented chatbot conditions yielding user-written rationales that are preferred under Elo-style ratings, deliberative of tradeoffs, and self-articulated. Together, these studies demonstrate how subjective and abstract qualities can be quantified for natural language tasks, by incorporating the domain expertise of humans, scalability of NLP tools and LLMs, and multi-faceted interpretations of quality.
Files
karen-zhou_dissertation-final.pdf
Files
(6.6 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:c21854d430c2998024a9ced719755b3a
|
6.6 MB | Preview Download |