Elicit
Record the learner's initial hypothesis, mechanism, predictions, confidence, and evidence before assistance.
Research questionCan structured AI-assisted instruction improve how graduate learners generate, test, and revise scientific hypotheses?
SL-001 public status
PLANNED STUDY CASE FILE
This case file treats scientific reasoning as a teachable practice rather than a fluency contest. The proposed intervention would prompt learners to state mechanisms, generate discriminating predictions, identify rival explanations, and revise confidence after critique. The study would test learning and transfer, not merely whether learners produce longer or more polished answers.
Estimate the contrast in rubric-scored scientific reasoning after structured AI-assisted instruction.
Test whether any improvement transfers to a new topic without access to the AI tutor.
Measure whether learner confidence becomes better calibrated to answer quality.
Identify failure modes, including deference to plausible but weak AI critiques.
DESIGN CANDIDATE
Every element remains provisional until the protocol is registered. Unknowns are shown as unknowns instead of being filled with unsupported precision.
Record the learner's initial hypothesis, mechanism, predictions, confidence, and evidence before assistance.
Present structured prompts for rival explanations, falsifiers, hidden assumptions, and measurement limits.
Require a traceable revision that states what changed, what did not change, and why.
Assess the same reasoning moves on a new scientific problem without the original scaffolding.
MEASUREMENT
Candidate roles may change during protocol review. Any primary outcome will be fixed before data collection or access to relevant outcome data.
| Outcome | Role | Operational definition | Timing |
|---|---|---|---|
| Scientific reasoning score | Candidate primary | Blinded analytic-rubric score covering mechanism, discriminating predictions, alternatives, testability, and evidence alignment. | Immediate post-instruction |
| Delayed transfer score | Candidate primary | Rubric score on a novel domain-adjacent problem completed without access to the study tutor. | Delay to be specified |
| Calibration error | Candidate secondary | Absolute difference between normalized confidence and observed rubric performance. | Baseline, post-instruction, and transfer |
| Unsupported uptake | Safety measure | Count and severity of revisions that adopt an AI suggestion without adequate evidence or valid reasoning. | During assisted lesson |
| Instructional time | Implementation | Active time required to complete the lesson, with idle-time rules prespecified. | Each lesson |
ANALYSIS DISCIPLINE
Final estimands, models, exclusions, missing-data rules, multiplicity decisions, and stopping conditions will be specified in the registered protocol where applicable.
Specify one primary contrast before outcome data are inspected and distinguish confirmatory from exploratory analyses.
Use a model that accounts for repeated observations and any cohort, course, or instructor clustering supported by the final design.
Report effect estimates with uncertainty intervals, score distributions, missingness, and inter-rater reliability.
Test sensitivity to prior domain knowledge, prior AI use, task order, rubric specification, and exclusion decisions.
Report null, inconclusive, and unfavorable results with the same outcome definitions used for favorable results.
RESEARCH INTEGRITY
The record must be detailed enough to audit what learners experienced, what the AI system could do, and where qualified humans remained responsible.
VALIDITY REGISTER
These responses reduce specific risks. They do not eliminate uncertainty or guarantee that the final design will support a causal claim.
Use newly authored transfer tasks, record search exposure where feasible, and keep scorers separate from lesson design.
Blind condition labels, standardize response formatting, train raters, and report reliability before adjudication.
Measure prior AI use and learner experience, and avoid interpreting engagement as learning.
Standardize essential lesson components and record permissible instructor adaptation.
Include delayed transfer and avoid claims about durable learning without follow-up evidence.
STUDY GATES
A stage label is a public claim. The record moves forward only when its stated exit condition is documented.
Question, intended inference, learning context, and initial validity risks are recorded.
Primary outcomes, allocation, power or precision rationale, and analysis plan are fixed.
Scientific, ethics, privacy, accessibility, and security determinations are documented.
A time-stamped protocol is public before enrollment or outcome inspection.
Results, uncertainty, deviations, validity limits, artifacts, and corrections are released.
EVIDENCE CONTEXT
These external sources provide context for design decisions. They are not Santaros Labs outputs, endorsements, or evidence that this proposed study has been completed.
Scientific Reports · 2025
Provides a real higher-education comparison and shows why instructional design, time, and learning outcomes should be measured separately.
Open sourceInstitute of Education Sciences · Living standard
Informs falsifiable hypotheses, outcome relevance, open science, implementation records, generalizability, and reporting of uncertainty.
Open sourceUNESCO · 2023
Frames human-centered pedagogical design, privacy, capacity, and institutional responsibility.
Open sourceOPEN SCIENCE PLAN
Availability will depend on consent, ethics review, licensing, privacy, security, and institutional requirements. A restriction will be explained rather than presented as open access.
Next portfolio record
SL-002Reproducible analysis training with AI assistanceMETHODS COLLABORATION
We welcome educators, learning scientists, domain researchers, statisticians, research software specialists, and governance reviewers who can strengthen the protocol.
A research inquiry does not imply study enrollment, institutional approval, funding, or authorship.