# Physical-scale scoring on a regular grid

```bash
stdetail fourier --input inputs/grid.npz --spacing-um 32 --out outputs/fourier.tsv
```

This command starts from raw integer RNA counts and fixed predictions. It
performs count splitting, exposure correction, normalization, Fourier
decomposition and RNA-referenced scoring. It does not require saved sufficient
statistics, a model checkpoint or the original private experiment directories.

The NPZ file contains:

| Array | Shape | Meaning |
|---|---|---|
| `counts` | occupied sites × genes | Nonnegative integer RNA molecule counts |
| `prediction` | occupied sites × genes, optionally × models | Fixed, finite expression predictions |
| `coords` | occupied sites × 2 | Grid-centre x and y coordinates in micrometres |
| `area` | occupied sites | Effective tissue area in square micrometres, strictly positive |
| `methods` | models | Optional unique model names, stored as strings |

All arrays must describe the same occupied sites in the same order and the same
fixed gene panel. The command does not select genes from the test outcomes or
automatically remove sites. `area` is the tissue area within each occupied bin;
it can be smaller than the squared grid spacing at tissue boundaries.

Only occupied tissue sites are supplied. They must lie on the specified regular
grid. The bounding rectangle is reconstructed and sites outside the supplied
tissue outline are zero-padded. Predictions are centred and L2-normalized for
each gene across occupied sites.

RNA counts undergo five independent binomial 50/50 splits with the original
seeds, 2026073101 through 2026073105. Each half is transformed as
`sqrt(half_counts / (0.5 * effective_area))`, centred across sites and divided
by the same gene-specific scale for both halves:
`sqrt((sum(A²) + sum(B²)) / 2)`. Constant columns become zero. The original
32-gene blocks and block-specific seed offsets are retained. For exact
reproduction, retain the original site order, gene order and gene panel.
`--seeds` can specify a different analysis, but it changes the molecule splits.

The orthonormal two-dimensional Fourier transform partitions all nonconstant
modes into seven physical wavelength bands: below 128, 128–256, 256–512,
512–1,024, 1,024–2,048, 2,048–4,096 and at least 4,096 micrometres.
Wavelength is one complete spatial cycle, not grid spacing. Bounds include the
lower limit and exclude the upper limit.

For normalized model coefficients `P` and RNA-half coefficients `A` and `B`,
each band reports the summed sufficient statistics:

```text
v_rna   = mean over splits of sum Re(conj(A) * B)
c_model = mean over splits of 0.5 * sum Re(conj(P) * (A + B))
v_model = sum |P|²
q       = c_model / sqrt(v_rna * v_model)
```

Sums cover the band's modes and supplied genes. Negative RNA signal estimates
are retained. A band with no Fourier modes, nonpositive RNA signal or zero
model variance receives a missing `q` and an explicit status. Valid `q` values
are signed and are not clipped to [0, 1]. `ALL_NON_DC` combines the sufficient
statistics of every band before calculating its score; it is not a mean of
band-specific correlations.

`nyquist_status` distinguishes fully represented physical windows from partial
high-frequency windows. A band is marked full when its lower wavelength bound
is at least twice the grid spacing. Scores from partial windows remain visible
for inspection, but the paper's physical-scale comparisons use full windows.

The implementation follows the pure transforms and accumulation loop in
`g0048_crc_physical_spectrum.py`. Model fitting, top-gene selection, virtual-spot
construction from transcript coordinates, specimen aggregation and confidence
interval estimation are outside this command. It evaluates an already prepared
regular grid; it does not claim to reproduce the full CRC analysis from raw
transcript files alone.

The Python API is `stdetail.fourier.score(counts, prediction, coords, area,
spacing_um=32, methods=None, seeds=...)`; `counts` may also be a SciPy CSR matrix.
