> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bioptimus.com/llms.txt
> Use this file to discover all available pages before exploring further.

# bioptimus.data.multi_modal.multi_modal_dataset

Multi-modal dataset for joint histology and omics inference.

Wraps a single WSI together with optional bulk RNA data, providing lazy patch access and a PyTorch-compatible `__getitem__` interface.

## MultiModalData

```python theme={null}
class MultiModalData(wsi_path: Union[str, Path],
                     extractor: TileExtractor,
                     bulk_rna_path: Union[str, Path] | None = None,
                     bulk_rna_gene_names: List[str] | None = None,
                     image_transform: Callable[[np.ndarray], Any] | None = None,
                     bulk_rna_transform: str | None = None,
                     bulk_rna_separator: str | None = None,
                     bulk_rna_gene_column: str | None = None,
                     bulk_rna_value_column: str | None = None,
                     bulk_rna_strip_version: bool = False,
                     wsi_backend: str | None = None)
```

Wraps a single Whole Slide Image with its extraction plan.

On construction the slide is opened via [`WSI`](/sdk-reference/io/wsi/factory#wsi) and the extractor is fitted + executed to produce a list of [`RegionSpec`](/sdk-reference/io/wsi/types#regionspec).  Each spec describes one tile/patch location.  Patches are read lazily via [`get_patch`](/sdk-reference/data/multi_modal/multi_modal_dataset#get_patch).

<ParamField body="wsi_path">
  Path to the WSI file (any format supported by [`WSI`](/sdk-reference/io/wsi/factory#wsi)).
</ParamField>

<ParamField body="bulk_rna_path">
  Optional path to bulk RNA CSV.
</ParamField>

<ParamField body="extractor">
  A *configured* [`TileExtractor`](/sdk-reference/extraction/wsi/tile_extraction#tileextractor).  A fresh `fit_extract` is called for every slide so the same extractor object can be reused across slides.
</ParamField>

<ParamField body="transform">
  Optional callable applied to the raw `np.ndarray` patch **before** it is returned.  Receives `(H, W, C)` uint8 and should return a transformed array (or tensor).
</ParamField>

**Attributes**:

* `path` *Path* - Resolved slide path.
* `bulk_rna_path` *Path | None* - Resolved path to bulk RNA CSV.
* `reader` *WSIReader* - Open reader for the slide.
* `specs` *List\[RegionSpec]* - Extraction plan (one entry per patch).
* `slide_name` *str* - Stem of the slide filename.

**Example**:

```python theme={null}
wsi = MultiModalData("tissue.svs", extractor)
patch, meta = wsi.get_patch(0)
```

***

#### \_\_len\_\_

```python theme={null}
def __len__() -> int
```

Returns the number of patches in the extraction plan.

**Returns**:

* `int` - Total patch count for this slide.

**Example**:

```python theme={null}
>>> wsi = MultiModalData("slide.svs", extractor)
>>> len(wsi)
150
```

***

#### get\_patch

```python theme={null}
def get_patch(idx: int) -> Tuple[Any, Dict[str, Any]]
```

Reads a single patch from the slide.

<ParamField body="idx">
  Index into `specs`.
</ParamField>

**Returns**:

A tuple `(patch, metadata)` where *patch* is an
`np.ndarray` of shape `(H, W, C)` (or whatever the
transform returns), and *metadata* is a dict with at
minimum `source`, `x`, `y`, `width`, `height`,
`slide_name`, and `tissue_ratio`.

**Raises**:

* `IndexError` - If *idx* is out of range.

***

#### close

```python theme={null}
def close() -> None
```

Closes the underlying WSI reader and releases resources.

**Example**:

```python theme={null}
>>> wsi.close()
```
