> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bioptimus.com/llms.txt
> Use this file to discover all available pages before exploring further.

# bioptimus.inference.aggregate

Slide-level aggregation utilities.

Small, dependency-light helpers that reduce a written tile-output file (the `(N_tiles, D)` `outputs` array produced by `bioptimus.inference.writers`) to a single slide-level vector.

Two convenience entry points build on the same mean-pool:

* [`slide_embedding`](/sdk-reference/inference/aggregate#slide_embedding) — averages tile embeddings into one slide embedding.
* [`pseudobulk`](/sdk-reference/inference/aggregate#pseudobulk) — averages predicted spatial gene expression into a single bulk RNA-seq-like profile, paired with the output gene names.

These operate on the *output of* `SlideInference.predict`/`.embed` (or the `Inference` facade), so no model or server is needed to aggregate.

## Pseudobulk

```python theme={null}
@dataclass(frozen=True, eq=False)
class Pseudobulk()
```

A predicted pseudobulk gene-expression profile for one slide.

Pairs the aggregated expression vector with the gene names it is indexed by, so callers can align it to a reference panel without a separate lookup.

`eq=False` keeps identity equality: a value-based `__eq__` would compare the numpy `expression` field with `==` and raise on truthiness.

**Attributes**:

* `expression` - The pseudobulk expression vector of shape `(G,)` as `float32`.
* `gene_names` - Ordered gene identifiers for each entry of `expression`, or `None` when the output has no gene names.

***

#### slide\_reduce

```python theme={null}
def slide_reduce(outputs: np.ndarray, pool: Pool = "mean") -> np.ndarray
```

Reduces per-tile outputs to one slide-level vector by the chosen pooling.

<ParamField body="outputs">
  The `(N_tiles, D)` per-tile outputs.
</ParamField>

<ParamField body="pool">
  The pooling strategy over the tile axis. One of `"mean"`, `"max"`, `"sum"`, or `"median"`. Defaults to `"mean"`.
</ParamField>

**Returns**:

The slide-level vector of shape `(D,)` as `float32`.

**Raises**:

* `BioptimusValueError` - If `pool` is not a supported strategy, or `outputs` is not a 2-D `(N_tiles, D)` array or has no tiles.

***

#### slide\_mean

```python theme={null}
def slide_mean(outputs: np.ndarray) -> np.ndarray
```

Mean-pools per-tile outputs into one slide-level vector.

Thin alias for `slide_reduce(outputs, pool="mean")`, kept for convenience and backward compatibility.

<ParamField body="outputs">
  The `(N_tiles, D)` per-tile outputs.
</ParamField>

**Returns**:

The slide-level vector of shape `(D,)` as `float32`.

**Raises**:

* `BioptimusValueError` - If `outputs` is not a 2-D `(N_tiles, D)` array or has no tiles.

***

#### slide\_embedding

```python theme={null}
def slide_embedding(path: str | Path,
                    output_format: OutputFormat | None = None,
                    *,
                    pool: Pool = "mean",
                    save: bool = False) -> np.ndarray
```

Reduces a slide's tile embeddings to one slide-level embedding.

Reads the per-tile embeddings from a file written in embedding mode (the embedding is the `outputs` array) or in the combined `PREDICTION_WITH_EMBEDDING` mode (the embedding is the separate `embeddings` array, alongside the `outputs` prediction). The `embeddings` array is preferred when present.

<ParamField body="path">
  Path to a tile-output file written in embedding or combined mode.
</ParamField>

<ParamField body="output_format">
  The on-disk format. Inferred from the extension when omitted.
</ParamField>

<ParamField body="pool">
  The pooling strategy over tiles (`"mean"`, `"max"`, `"sum"`, or `"median"`). Defaults to `"mean"`.
</ParamField>

<ParamField body="save">
  When `True`, saves the vector next to the input as `<name>.slide.npy`. Defaults to `False`.
</ParamField>

**Returns**:

The slide embedding of shape `(D,)` as `float32`.

***

#### pseudobulk

```python theme={null}
def pseudobulk(path: str | Path,
               output_format: OutputFormat | None = None,
               *,
               pool: Pool = "mean",
               save: bool = False) -> Pseudobulk
```

Reduces a slide's predicted spatial expression to a pseudobulk profile.

<ParamField body="path">
  Path to a tile-output file written in `predict` mode.
</ParamField>

<ParamField body="output_format">
  The on-disk format. Inferred from the extension when omitted.
</ParamField>

<ParamField body="pool">
  The pooling strategy over tiles (`"mean"`, `"max"`, `"sum"`, or `"median"`). Defaults to `"mean"`.
</ParamField>

<ParamField body="save">
  When `True`, saves the vector next to the input as `<name>.slide.npy`. Defaults to `False`.
</ParamField>

**Returns**:

A [`Pseudobulk`](/sdk-reference/inference/aggregate#pseudobulk) pairing the `(G,)` expression vector with the
output gene names (`None` if the file has none).
