Decompose
Turn the review question into explicit claims, populations, exposures or interventions, comparators, outcomes, and time frames.
Research questionCan structured AI-assisted instruction improve how learners represent consensus, uncertainty, contradiction, and evidence quality?
SL-003 public status
PLANNED STUDY CASE FILE
This case file focuses on disciplined synthesis rather than summary fluency. Learners would map claims to sources, separate absence of evidence from evidence of absence, distinguish inconsistency from imprecision, and communicate how strongly a conclusion is supported. AI assistance would be evaluated as an instructional scaffold whose own omissions and unsupported claims must be detected.
Estimate the contrast in source-supported conclusion accuracy between instructional conditions.
Measure detection and representation of material scientific disagreement.
Assess whether learners calibrate conclusion strength to evidence certainty.
Test transfer to a new evidence packet without the original prompts or examples.
DESIGN CANDIDATE
Every element remains provisional until the protocol is registered. Unknowns are shown as unknowns instead of being filled with unsupported precision.
Turn the review question into explicit claims, populations, exposures or interventions, comparators, outcomes, and time frames.
Link each material statement to supporting, conflicting, indirect, or missing source evidence.
Evaluate risk of bias, inconsistency, indirectness, imprecision, and publication or selection concerns.
Write a conclusion whose scope and confidence match the available evidence, then test transfer on a new packet.
MEASUREMENT
Candidate roles may change during protocol review. Any primary outcome will be fixed before data collection or access to relevant outcome data.
| Outcome | Role | Operational definition | Timing |
|---|---|---|---|
| Atomic claim support | Candidate primary | Proportion of material conclusion claims correctly supported by the cited source passages. | Final synthesis |
| Contradiction representation | Candidate co-primary | Accuracy and completeness when identifying material disagreement and its likely sources. | Final synthesis |
| Certainty calibration | Candidate secondary | Agreement between expressed conclusion strength and expert-adjudicated evidence certainty. | Final synthesis |
| Selective omission | Safety measure | Rate at which evidence that could materially change the conclusion is omitted. | Final synthesis |
| Delayed transfer | Candidate secondary | Performance on a new packet without access to the instructional scaffold. | Delay to be specified |
ANALYSIS DISCIPLINE
Final estimands, models, exclusions, missing-data rules, multiplicity decisions, and stopping conditions will be specified in the registered protocol where applicable.
Define atomic-claim and contradiction units before scoring begins.
Model repeated evidence packets and raters, with clustering by course or learning group where applicable.
Report agreement before and after adjudication and disclose rubric changes.
Analyze unsupported claims, omissions, and overconfident conclusions as separate error classes.
Test sensitivity to evidence certainty, domain familiarity, packet length, and alternative scoring thresholds.
RESEARCH INTEGRITY
The record must be detailed enough to audit what learners experienced, what the AI system could do, and where qualified humans remained responsible.
VALIDITY REGISTER
These responses reduce specific risks. They do not eliminate uncertainty or guarantee that the final design will support a causal claim.
Collect independent annotations, publish the adjudication process, and retain disagreement where consensus is not justified.
Prespecify selection rules and include supporting, null, conflicting, indirect, and lower-certainty evidence.
Use controlled source access and newly assembled packets where licensing and review permit.
Score evidence representation separately from style and surface fluency.
Require passage-level verification and score citation presence separately from citation support.
STUDY GATES
A stage label is a public claim. The record moves forward only when its stated exit condition is documented.
Learning objective, target inference, candidate packets, and error classes are recorded.
Domains, reference standard, primary outcomes, and analysis decisions are fixed.
Scientific, education, licensing, privacy, and accessibility reviews are documented.
Protocol, packets, prompts, and scoring rules are time-stamped before data collection.
All result directions, uncertainty, disagreement, deviations, and artifact access are reported.
EVIDENCE CONTEXT
These external sources provide context for design decisions. They are not Santaros Labs outputs, endorsements, or evidence that this proposed study has been completed.
The BMJ · 2021
Informs transparent reporting of review questions, selection, synthesis methods, results, limitations, and availability.
Open sourceCochrane · Current handbook
Frames structured presentation of findings and assessment of certainty across risk of bias, inconsistency, indirectness, imprecision, and publication bias.
Open sourceUNESCO · 2023
Supports human-centered pedagogical validation and attention to privacy, agency, inclusion, and institutional readiness.
Open sourceOPEN SCIENCE PLAN
Availability will depend on consent, ethics review, licensing, privacy, security, and institutional requirements. A restriction will be explained rather than presented as open access.
Next portfolio record
SL-004Mentored team learning in shared AI research workspacesMETHODS COLLABORATION
We welcome educators, learning scientists, domain researchers, statisticians, research software specialists, and governance reviewers who can strengthen the protocol.
A research inquiry does not imply study enrollment, institutional approval, funding, or authorship.