Scientific reasoning with AI co-investigators
When does AI assistance improve the diversity, testability, and calibration of scientific hypotheses?
Proposed design
Proposed randomized within-subject study with domain researchers completing matched hypothesis-generation and critique tasks with and without AI support.
Planned public output
Preregistered protocol, de-identified task materials, scoring rubric, and analysis code where consent and licensing allow.
Primary measures
- Hypothesis diversity
- Testability
- Expert-rated plausibility
- Confidence calibration
- Error detection
