python3 study_lab.py run --scenario worker-contention
python3 study_lab.py evaluate
The source includes three fixtures, seven contract checks, and a model-adapter interface. Output must cite known evidence IDs and include an experiment, measurements, uncertainty, and a rollback condition.
Keyword coverage can miss synonyms and accept weak answers. It is not a correctness score. Contract validation cannot establish semantic support for a citation. Human review remains necessary.
Offline mode has no model-token cost. Elapsed time is measured by the CLI; it is not a model or production benchmark.
What this demonstrates
The harness makes evidence and review boundaries explicit. A researcher proposes hypotheses, a challenger identifies missing information, and a reviewer checks a small contract before the exercise is accepted. In the baseline these stages are represented by fixture outputs; a model adapter can replace that provider.