Study portfolio
SL-002Stage 02: Protocol draftingPlanned intervention study

Reproducible analysis training with AI assistance

Research questionCan learners produce a more reproducible scientific analysis when AI assistance is embedded inside explicit provenance instruction?

SL-002 public status

Current stage
Protocol drafting
Stage 02 of 05
Recruitment
Not started
No participants are being enrolled
Results
None
This case file contains no study findings
Last reviewed
28 August 2026
Public planning record

PLANNED STUDY CASE FILE

A Learning Question With a Testable Decision

This case file connects computational training with research integrity. Learners would not be scored only on obtaining the expected answer. They would also be evaluated on whether an independent analyst can reconstruct data selection, transformations, code execution, model use, judgment calls, and the computing environment from the submitted record.

Learning context
A graduate statistics, data science, laboratory methods, or research-software course in which learners complete an end-to-end analysis from an open scientific dataset.
Decision this study should support
Determine whether provenance-first AI instruction improves independent reproducibility enough to justify integration into quantitative research training.

Protocol objectives

  1. 01

    Estimate the contrast in blinded reproduction success between instructional conditions.

  2. 02

    Measure which provenance fields most often prevent or enable successful reproduction.

  3. 03

    Test transfer to a new dataset and analysis environment.

  4. 04

    Quantify hidden manual decisions and unsupported analytic changes introduced during AI use.

DESIGN CANDIDATE

What Would Be Tested

Every element remains provisional until the protocol is registered. Unknowns are shown as unknowns instead of being filled with unsupported precision.

Design candidate
Clustered or individually randomized instructional comparison followed by a blinded reproduction challenge
Unit of assignment
Learner or course section, selected after contamination and delivery constraints are assessed
Conditions
Provenance-first AI workflow and conventional worked-example instruction
Reproduction unit
A complete analysis package evaluated by an analyst who did not create it
Sample size
Not set. The protocol must account for clustering, repeated tasks, and the binary or ordinal primary endpoint
Data source
Open, versioned scientific datasets with documented licenses and stable reference analyses

Learning sequence

01

Inspect

Evaluate dataset documentation, license, provenance, missingness, and analysis suitability before code is written.

02

Declare

State the question, intended estimand or analytic target, exclusions, transformations, and acceptance tolerances.

03

Execute

Run analysis while capturing code, environment, AI interactions, sources, and material human decisions.

04

Reproduce

Transfer the package to a blinded analyst and record every ambiguity, failure, and undocumented dependency.

MEASUREMENT

Outcomes Defined Before Observation

Candidate roles may change during protocol review. Any primary outcome will be fixed before data collection or access to relevant outcome data.

OutcomeRoleOperational definitionTiming
Reproduction successCandidate primaryWhether a blinded analyst obtains the target outputs within prespecified numerical and qualitative tolerances.After package submission
Provenance completenessCandidate co-primaryProportion of required provenance fields that are accurate, specific, and sufficient for reuse.At package audit
Unresolved decision countCandidate secondaryNumber of material analytic choices that cannot be reconstructed from the submitted record.During reproduction
Transfer performanceCandidate secondaryReproducibility score for a new dataset completed without the original teaching examples.Delayed task
Time and support burdenImplementationLearner time, instructor assistance, and reproducer time under prespecified logging rules.Throughout each task

ANALYSIS DISCIPLINE

A Plan That Can Report Unfavorable Results

Final estimands, models, exclusions, missing-data rules, multiplicity decisions, and stopping conditions will be specified in the registered protocol where applicable.

  1. 01

    Prespecify the reproduction tolerance and adjudication rule before packages are evaluated.

  2. 02

    Estimate condition contrasts with models appropriate to the assignment unit, repeated tasks, and course clustering.

  3. 03

    Report failure categories, not only an aggregate reproduction rate.

  4. 04

    Conduct sensitivity analyses using stricter and looser tolerances declared in advance.

  5. 05

    Separate pedagogical effectiveness, technical reliability, and time cost in interpretation.

RESEARCH INTEGRITY

Implementation, Access, and AI Disclosure

The record must be detailed enough to audit what learners experienced, what the AI system could do, and where qualified humans remained responsible.

Implementation record

  • Dataset identifiers, versions, licenses, checksums, and download dates
  • Code, package lockfiles, container definition, operating environment, and random seeds
  • AI prompts or instructions, outputs used, model details, tool calls, and human approval points
  • Instructor interventions, hints, exceptions, and technical support
  • Reference-output generation and tolerance-setting procedure

Equity and access

  • Provide equivalent computing access and document hardware or network constraints.
  • Design tasks so success does not depend on access to a private paid tool outside the study.
  • Measure baseline programming and statistical experience using interpretable categories.
  • Publish accessible teaching materials and lower-compute alternatives where feasible.

AI-system disclosure

  • Model and provider identifiers, access mode, version dates, and data-retention configuration
  • Code-execution tools, file access, network access, retrieval sources, and environment boundaries
  • Prompts or instructions that generated or revised analysis code
  • Exact AI-generated changes accepted, rejected, or modified by the learner
  • Service or model changes that could affect repeatability

VALIDITY REGISTER

Main Risks and Planned Responses

These responses reduce specific risks. They do not eliminate uncertainty or guarantee that the final design will support a causal claim.

R01

Reference-answer dependence

Define scientifically acceptable output tolerances and include analyses with more than one valid implementation.

R02

Prior technical expertise

Measure baseline skills, block or stratify assignment if justified, and report heterogeneous effects cautiously.

R03

Environment drift

Version environments, archive container definitions, and timestamp all externally hosted dependencies.

R04

Hidden instructor rescue

Log instructional support and distinguish independent completion from assisted completion.

R05

Tool-specific findings

Describe the tested configuration precisely and avoid generalizing beyond evaluated systems and tasks.

STUDY GATES

Status Changes Require an Exit Record

A stage label is a public claim. The record moves forward only when its stated exit condition is documented.

01

Concept

Complete

Learning problem, intended decision, candidate tasks, and reproducibility definition are recorded.

02

Protocol drafting

Current

Assignment unit, primary tolerance, outcomes, and analysis plan are fixed.

03

Review

Not started

Scientific, education, privacy, accessibility, and infrastructure reviews are documented.

04

Registration and execution

Not started

Protocol, tasks, environments, and scoring rules are time-stamped before use.

05

Reporting

Not started

Reproduction outcomes, failure taxonomy, costs, deviations, and artifacts are released.

EVIDENCE CONTEXT

Real Sources That Inform the Protocol

These external sources provide context for design decisions. They are not Santaros Labs outputs, endorsements, or evidence that this proposed study has been completed.

01

Scientific Data · 2016

The FAIR Guiding Principles for scientific data management and stewardship

Supports rich metadata, persistent identification, provenance, interoperability, and reuse across data, tools, and workflows.

Open source
02

National Institute of Standards and Technology · 2021

Ensuring reproducibility in the application of artificial intelligence to materials R&D

Provides a real scientific-computing context for reusable datasets, models, and executable research infrastructure.

Open source
03

Institute of Education Sciences · 2022

Sharing Study Data: A Guide for Education Researchers

Informs decisions about what education data can be shared and what documentation is needed for responsible reuse.

Open source

OPEN SCIENCE PLAN

Planned Public Open

  • Preregistered teaching study protocol
  • Open lesson materials
  • Provenance schema and validator
  • Versioned benchmark datasets
  • Reference containers
  • Reproduction scoring manual
  • Analysis code and governed data statement

Availability will depend on consent, ethics review, licensing, privacy, security, and institutional requirements. A restriction will be explained rather than presented as open access.

METHODS COLLABORATION

Improve This Study Before Registration

We welcome educators, learning scientists, domain researchers, statisticians, research software specialists, and governance reviewers who can strengthen the protocol.

Start a research inquiry
A research inquiry does not imply study enrollment, institutional approval, funding, or authorship.