Published August 2026
| Version v1
Dissertation
Open
Towards Reliable AI Scientists
Description
Scientific discovery often follows a loop. Scientists form hypotheses and ideas, decide which are worth pursuing, and then test them. The challenging part to scale is the decision step, because "good" can mean many things, including explanatory value, novelty, feasibility, and potential impact.
My early work, Literature meets Data, introduced an algorithm for hypothesis generation. By combining information from scientific literature and observed data, we can generate hypotheses that explain real-world phenomena. Building on this, we developed HypoBench, to provide principled evaluations for hypothesis generation with a focus on explanatory power. HypoBench helps identify what current methods do well and where they fall short, while highlighting an open question: interestingness and novelty are crucial but difficult to evaluate consistently. This evaluation challenge also applies to scientific publishing, where increasing submissions make it harder for reviewers to reliably differentiate higher- and lower-quality work. We argue that improving review quality requires stronger evidence than reading and summarizing papers alone.
Motivated by these limitations, we develop two complementary systems. NeuriCo aims to enable AI agents to help explore and test research ideas in practice. Veritas asks how AI systems can help verify scientific claims rigorously, grounded by execution. Together, they cover two key components for building reliable AI scientists: producing evidence for new research ideas and checking evidence for existing claims. This thesis takes one step toward AI systems that can reliably help humans make meaningful progress in scientific discovery.
Files
haokun-thesis.pdf
Files
(1.6 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:780084692543888bd3d231206748db83
|
1.6 MB | Preview Download |