experiment EX-HANDOFF-2026-6E42
Handoff rubric and protocol preflight
Experiment
Research question
Can the safe-continuation rubric, fixtures, study workspace, and safeguard protocol operate reliably enough to justify a treatment-effect study?
Non-efficacy boundary
This preflight does not estimate whether structured handoffs work. Synthetic and deliberately varied example responses may be used to test scoring. Any condition-labelled observations are for defect finding only and must not be reported as treatment effects.
Method
- Build a response set spanning correct, incomplete, unsupported, high-confidence incorrect, accessibility-failed, and noncompleted cases.
- Two independent reviewers score concealed examples using the draft rubric.
- Measure category-specific agreement, not only an overall average.
- Adjudicate disagreements and change the rubric only in a versioned log.
- Repeat with a held-out response set; do not report only the training set.
- Run the complete accessible workspace and consent/withdrawal path with participants representing required assistive-technology use.
Advancement criteria
- Primary safe-continuation and critical-harm categories meet a preregistered agreement threshold on held-out examples.
- No unresolved accessibility-critical or security-critical defect.
- Every fixture passes authorization, secret/PII, realism, and safe-action review.
- Reviewers can score without participant identity and with condition concealed as far as artifact format permits.
- The study owner, data steward, and qualified ethics reviewer approve the revised protocol.
Falsification and stopping
Stop and redesign if critical categories cannot be scored reliably, if format reveals condition in a way that biases scoring, if accommodations change the construct, or if safe/authorized fixtures are too artificial to represent the target task.
Results
Not run. It requires independent reviewers and direct accessibility participation; autonomous repository agents cannot satisfy those roles.