# STDetail evaluation bundle v1

`schema: "stdetail-evaluation-bundle-v1"` is a scores-only companion to the unchanged paired-expression `stdetail-analysis-bundle-v1`. It accepts a single prediction, multiple models and multiple biological units. It never fabricates a second arm or forces five bands on Fourier or region-ordering results.

```json
{
  "schema": "stdetail-evaluation-bundle-v1",
  "title": "My HEST evaluation",
  "workflow": "hest",
  "synthetic": false,
  "metrics": [{
    "id": "q_fine",
    "label": "Fine-scale accuracy",
    "description": "Combined B4+B5 components, not their arithmetic mean.",
    "unit": "score",
    "higher_is_better": true,
    "object": "prediction"
  }],
  "rows": [{
    "model": "prediction_A",
    "task": "LUNG",
    "unit_id": "specimen_1",
    "unit_type": "specimen",
    "metric": "q_fine",
    "value": null,
    "status": "ZERO_MODEL_VARIANCE",
    "support": "graph-retained locations",
    "details": {},
    "model_role": "prediction"
  }],
  "sources": [{"name": "run1/specimen_scores.tsv", "columns": [], "rows": []}],
  "notes": ["Export preserves the saved scientific scores."]
}
```

The illustrated null is structural, not an experimental result. The computed examples beside this file contain actual numerical outputs from generated inputs.

## Rules

- Workflow is `hest`, `her2`, `fourier`, or `cell`. One bundle uses one workflow. Different protocols are not pooled.
- Row identity is `(model, task, unit_id, metric)`. Duplicate identities fail the export instead of doubling the sample count. Model or unit names are never inferred from a directory name.
- `value` is a finite number or `null`. Undefined scores retain a status; negative values and scores above one remain unchanged. All details and source-table values are JSON-compatible; nonfinite TSV values become null.
- `metrics[].object` and `rows[].model_role` are `prediction`, `rna_reference`, or `baseline`. References must not enter prediction-model leaderboards. Comparison panels select one metric, one task and matching unit support.
- HEST metrics are `pearson_graph_support`, `q_broad`, `q_fine`, `q_B1` through `q_B5`. Broad and fine use saved combined components; they are not arithmetic means of two band scores. The five graph bands have no fixed micrometre interpretation.
- HER2 task labels identify across-region, matched-context or within-region questions. Prediction metrics are `model_accuracy` and `accuracy_relative_to_rna_repeatability`; reference metrics are `lopo_gene_majority_baseline` and `rna_split_half_accuracy`. Baseline/reference rows are emitted once per task and patient.
- Fourier metric is `physical_q`; task includes the physical wavelength band and the grid spacing. `details.nyquist_status` explicitly identifies partial high-frequency support. `ALL_NON_DC` is a combined-components score. The original table lacks specimen IDs, so `--unit` is mandatory and repeated in the same order as `--run`.
- Cell metric `cell_q` preserves track/view/band. RNA-only tables yield `rna_split_agreement`. Hotspot tables yield `E_model_normalized_log_score` and `E_RNA_normalized_log_score`, with model-specific support in the task. Specimens remain specimens; patient identity stays in details. Source tables can have different model support, which is not fixed or harmonized during export.
- `sources` stores scored table contents for inspection; expression matrices, images, coordinates and model weights are not added. This bundle supports score analysis, not spatial-expression plotting. Use the original paired HEST export when expression maps are wanted.
- `synthetic` labels generated software examples. It is not evidence of performance. The exporter itself cannot infer whether arbitrary data are synthetic.

## Commands

```bash
stdetail export-results --run outputs/hest --output outputs/hest-results.json
stdetail export-results --workflow her2 --run outputs/her2 --output outputs/ordering-results.json
stdetail export-results --run outputs/fourier_P1.tsv --unit P1 --run outputs/fourier_P2.tsv --unit P2 --output outputs/physical-results.json
stdetail export-results --workflow cell --run outputs/cell_scores --output outputs/cell-results.json
```

`--workflow auto` is the default. Explicit workflow must match the existing output. Repeat `--run` to merge compatible completed results, including multiple HEST specimens or cell output directories. Existing output files are never overwritten.

## Computed examples

Files in `synthetic-examples-ready/`:

| File | Numerical path used | Structure |
|---|---|---|
| `hest.json` | Original HEST count/graph scorer; 20 molecule splits | One prediction, one specimen, 8 metrics |
| `her2.json` | Original ordering scorer; one specified molecule split | Two predictions, two patients, 3 tasks, separate reference rows |
| `fourier.json` | Original Fourier scorer; one specified molecule split | Two predictions, two specimens, eight windows including ALL_NON_DC |
| `cell.json` | Original lazy graph band and component functions | One prediction, one specimen, five bands and RNA reference |

The cell example exercises numerical components, not the transcript preparation or full hotspot pipeline. The short HER2/Fourier split lists are explicitly synthetic software examples, not paper reproductions. Generate again with `python examples/export_results_example.py --output outputs/export-examples` after installing the Python package.
