Skip to main content
Slide-level aggregation utilities. Small, dependency-light helpers that reduce a written tile-output file (the (N_tiles, D) outputs array produced by bioptimus.inference.writers) to a single slide-level vector. Two convenience entry points build on the same mean-pool:
  • slide_embedding — averages tile embeddings into one slide embedding.
  • pseudobulk — averages predicted spatial gene expression into a single bulk RNA-seq-like profile, paired with the output gene names.
These operate on the output of SlideInference.predict/.embed (or the Inference facade), so no model or server is needed to aggregate.

Pseudobulk

A predicted pseudobulk gene-expression profile for one slide. Pairs the aggregated expression vector with the gene names it is indexed by, so callers can align it to a reference panel without a separate lookup. eq=False keeps identity equality: a value-based __eq__ would compare the numpy expression field with == and raise on truthiness. Attributes:
  • expression - The pseudobulk expression vector of shape (G,) as float32.
  • gene_names - Ordered gene identifiers for each entry of expression, or None when the output has no gene names.

slide_reduce

Reduces per-tile outputs to one slide-level vector by the chosen pooling.
The (N_tiles, D) per-tile outputs.
The pooling strategy over the tile axis. One of "mean", "max", "sum", or "median". Defaults to "mean".
Returns: The slide-level vector of shape (D,) as float32. Raises:
  • BioptimusValueError - If pool is not a supported strategy, or outputs is not a 2-D (N_tiles, D) array or has no tiles.

slide_mean

Mean-pools per-tile outputs into one slide-level vector. Thin alias for slide_reduce(outputs, pool="mean"), kept for convenience and backward compatibility.
The (N_tiles, D) per-tile outputs.
Returns: The slide-level vector of shape (D,) as float32. Raises:
  • BioptimusValueError - If outputs is not a 2-D (N_tiles, D) array or has no tiles.

slide_embedding

Reduces a slide’s tile embeddings to one slide-level embedding. Reads the per-tile embeddings from a file written in embedding mode (the embedding is the outputs array) or in the combined PREDICTION_WITH_EMBEDDING mode (the embedding is the separate embeddings array, alongside the outputs prediction). The embeddings array is preferred when present.
Path to a tile-output file written in embedding or combined mode.
The on-disk format. Inferred from the extension when omitted.
The pooling strategy over tiles ("mean", "max", "sum", or "median"). Defaults to "mean".
When True, saves the vector next to the input as <name>.slide.npy. Defaults to False.
Returns: The slide embedding of shape (D,) as float32.

pseudobulk

Reduces a slide’s predicted spatial expression to a pseudobulk profile.
Path to a tile-output file written in predict mode.
The on-disk format. Inferred from the extension when omitted.
The pooling strategy over tiles ("mean", "max", "sum", or "median"). Defaults to "mean".
When True, saves the vector next to the input as <name>.slide.npy. Defaults to False.
Returns: A Pseudobulk pairing the (G,) expression vector with the output gene names (None if the file has none).