User guide

What each part of the site takes in and gives back.

Expression tables

In Analyze data, choose one file of measured RNA and one file per model, type the expression scale they share, and press Analyze data. Analyze the example does the same with one HEST lung section; the example files can be copied as a template.

Each file has one row per location. It starts with sample_id, unit_id, followed by one column per gene. Coordinates x, y and a region label are optional and can sit in the RNA file or in a file of their own. patient_id groups the sections of one patient. Rows can be in any order. Genes or locations missing from any file are dropped, and Data details lists them.

Values are used as supplied. RNA and predictions must be on the same scale, since every map of a gene shares one colour range.

Choose a region from the list or drag a rectangle on a map. The figure then sets the correlation inside the region beside the one for the whole section. Sorting the gene list by whole minus regional correlation puts the genes that lose most inside the region at the top.

These are plain correlations of the supplied values. G and q need molecule counts and come from the Python package.

Score tables

A score table compares the same predictions under two conditions, with one row per model, task and specimen or patient. The example under Computed results is the table behind Figure 6a,b.

ColumnContent
modelModel name
taskCohort or tissue group. Figure columns are grouped by it
unit_idSpecimen or patient ID. A column named specimen or patient also works
unit_typespecimen or patient, one kind per file
fine_baseline, fine_interventionFine-scale q under the two conditions
coarse_baseline, coarse_interventionBroad-scale q under the two conditions
pcc_baseline, pcc_interventionOverall Pearson r under the two conditions
score_definitionOptional. band_mean for the mean q of two bands, combined_components for q from their pooled components
baseline_condition, intervention_conditionOptional names of the two conditions
experiment_id, source_tableOptional provenance. Keep separate experiments in separate files

All models must cover the same tasks and IDs. Scores of an ID that appears in several tasks are averaged first, then every model and every ID carries equal weight. Pearson r stays between −1 and 1. q is corrected for noise and can exceed 1.

Fine- and broad-scale q come in two forms: the mean q of the two bands, or q from their pooled components, which is what stdetail hest writes. Name the form in the file or when you open it.

Intervals in the browser come from 20,000 fixed-seed resamples of IDs within tasks and need at least five IDs. The paper computed its intervals differently; they are in its Source Data.

Analysis bundles

A finished HEST run packed with its expression data, so that maps can be drawn next to the scores.

stdetail export --manifest path/to/manifest.json --run outputs/hest-run --output outputs/hest-analysis.json

The command only packs what the run produced. It checks each saved gene correlation against the expression it packs.

Opened under Computed results, a bundle gives q in the five bands and its change, RNA variance and covariance per band, expression maps, gene correlations and prediction errors. The bands are relative to the section, not distances in micrometres.

The schema is stdetail-analysis-bundle-v1: a score_table and a list of specimens, each with genes, spot IDs, coordinates, measured expression and two prediction arms. The HEST example shows the full layout.

Python package

Splitting molecules, building the spatial graph and computing G and q need raw counts, so they run on your machine. Each workflow writes a file that opens under Computed results.

QuestionInputWorkflow
Whole-section agreementExpression tables of RNA and predictionsAnalyze data, no installation
Ordering within regionsInteger counts, predictions, coordinates and pathology regions from two or more patientsstdetail her2 · input specification
q in relative bands (HEST)Integer counts for 50 genes, coordinates, held-out predictions and the reference cohortstdetail hest · input specification
q by wavelength on a gridInteger counts, predictions, grid coordinates in µm and tissue areastdetail fourier · input specification
Single cells and pooled viewsCell counts and metadata, predictions and the cell protocolstdetail.cell_rna, stdetail.cell_score · input specification

Install

Python 3.10 or later. Install the PyTorch build for your CPU or GPU first. The package holds the code, workflow notes and small simulated examples, with no patient data or model weights. Terms of use are in the software notice.

python -m pip install stdetail-*.whl   # in the folder with the wheel
stdetail --help
stdetail example --output outputs/synthetic

stdetail example scores a small simulated section and writes outputs/synthetic/analysis-bundle.json.

Export finished runs

stdetail export-results --run outputs/her2 --output outputs/ordering-results.json
stdetail export-results --run outputs/hest --output outputs/hest-results.json
stdetail export-results --run outputs/fourier_P1.tsv --unit P1 --run outputs/fourier_P2.tsv --unit P2 --output outputs/physical-results.json
stdetail export-results --run outputs/cell_scores --output outputs/cell-results.json

Repeat --run to combine runs. The file layout is in the bundle schema and the export notes.

HEST sections

stdetail hest scores existing predictions with the paper's HEST protocol. Nothing is retrained. It needs:

  1. an NPZ file per specimen with counts (raw integers), xy, barcodes and genes, including the reference specimens of the same task;
  2. an HDF5 file per prediction with prediction, barcodes and genes, predicting log1p(raw count) at every original location (maps need both a baseline and an intervention);
  3. a manifest naming the files and the model, as in the HEST guide.

Run it from the package folder, into a new output folder:

stdetail hest --manifest path/to/manifest.json --output outputs/hest-run
stdetail export-results --run outputs/hest-run --output outputs/hest-results.json

The run writes result_comparison.tsv, which opens as a score table. The CPU version handles up to 6,000 graph locations. Cropping a section changes its graph, and with it the scores.

Simulated test files

Output of each workflow on simulated data, for checking that package files open here.

Figures

The figures follow the paper: Arial at 7 pt when printed 180 mm wide, lower-case panel letters, left and bottom axes only. Expression runs from slate blue through cream to rust. Changes run from brick (lower) through white to teal (higher).

SVG keeps the text editable. PNG is the same figure at 300 dpi. Source data is a CSV of every value in the figure, including genes or locations left out for space.

Rows of patients or specimens are drawn as in Figure 5e, with one mark per patient, a box for the 95% t interval of the mean (from three patients up) and a white diamond at the mean. Paired changes are drawn as in Figure 6.

Study results

HER2ST reopens Figure 2a–f, Figure 5e,f and the FASN case for 11 methods in eight patients. Clicking a method or a patient marks it in every panel. The cohort values stay as published.

HEST TENX118 is one lung section predicted by Ridge-Phikon-v2, with standard training and with the neighbouring-difference loss.

Limits

  • Expression tables: 32 MB per file, 64 MB in total, 2 million shared expression values.
  • Score tables: 2 MB and 500 rows.
  • Analysis bundles: 64 MB, 2 million expression values, 20,000 locations and 5,000 genes per specimen.
  • Evaluation bundles: 64 MB and 100,000 score rows. Tables show 300 rows; downloads keep all of them.
  • h5ad, h5, RDS and image files are prepared with the package and do not open in the browser.

Molecules can only be split from raw integer counts. Normalised or log values give no RNA repeatability and no q.

Definitions

RNA repeatability
Each counted molecule goes at random to one of two halves. Differences found in both halves are repeatable, and agreement between the halves is the reference a prediction is measured against.
Local recovery
Agreement on local differences: within matched tissue context, at fine spatial scales and between neighbouring cells.
Ordering gain, G
Whether a prediction puts two sites in the right order for a gene. At 0% it is no better than a held-out gene-majority baseline. At 100% it orders as well as one RNA half orders the other. G can be negative or exceed 100%.
Across regions, matched context, within regions
The three sets of site pairs in HER2ST. Pairs across regions may span pathology regions. Matched pairs share RNA-defined tissue context and depth. Pairs within regions share a pathology region, so knowing the region no longer helps.
Noise-corrected correlation, q
How closely a prediction follows the measured variation over a range of distances, corrected for the sampling noise of the counts. A prediction that follows all repeatable variation scores close to 1.
Relative bands
Broadest, Broad, Mid, Fine and Finest, from section-wide gradients to differences between neighbouring spots. Fine-scale q is the mean q of Fine and Finest, broad-scale q that of Broad and Broadest.
Overall Pearson r
For each gene, the correlation between prediction and measured RNA across all sites of a section, averaged over genes and then over the sections of a patient. A gene predicted as constant scores 0. Knowing the tissue regions alone can earn a high value.
Neighbouring-difference loss
A training term that also penalises errors in the predicted differences between neighbouring training sites. In the paper it raised fine-scale q in all eight spot models tested.