This repository contains the code for the paper:
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
Bingyang Ye*, Shan Chen*, Jingxuan Tu, Chen Liu, Zidi Xiong, Samuel Schmidgall, Danielle S. Bitterman
Harvard University, Mass General Brigham, Boston Children's Hospital, Google DeepMind
(*Co-first authors)
If you use this codebase in your research, please cite our work:
@misc{ye2026prooftimebenchmarkevaluating,
title={Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments},
author={Bingyang Ye and Shan Chen and Jingxuan Tu and Chen Liu and Zidi Xiong and Samuel Schmidgall and Danielle S. Bitterman},
year={2026},
eprint={2601.07606},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2601.07606},
}Preprint: https://arxiv.org/abs/2601.07606
The benchmark datasets are available on HuggingFace:
@dataset{proof-of-time-dataset-2026,
title={Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments},
author={Ye, Bingyang and Chen, Shan and Tu, Jingxuan and Liu, Chen and Xiong, Zidi and Schmidgall, Samuel and Bitterman, Danielle S.},
year={2026},
publisher={HuggingFace},
url={https://huggingface.co/datasets/AIM-Harvard/proof-of-time}
}