Skip to main content
The Bioptimus Python SDK runs whole-slide inference against either an on-premise server or a SageMaker endpoint. It handles WSI reading, tiling, tissue masking, bulk-RNA alignment, and concurrent dispatch. The Bioptimus SDK talks to whichever model is deployed at the endpoint you connect to. Backbone.available_backbones() lists the models the Bioptimus SDK knows how to build (h1, m-optimus, tissue-seg) — to see what a server actually has loaded, check its /ping response.

Installation

pip install bioptimus-sdk

Two ways to use it

Inference pipeline (recommended)

One object configures the whole pipeline. Caches tissue masks, organizes outputs into a workspace, and is reproducible. Best for most users and for cohorts.

Core API (advanced)

Backbone + SlideInference give explicit, per-slide control over the model client, mask provider, and writer.

Connecting to a model

The Backbone factory is the low-level client used by both layers.
from bioptimus.models.backbones import Backbone
from bioptimus.models.types import Models

print(Backbone.available_backbones())   # ['h1', 'm-optimus', 'tissue-seg']
model = Backbone(Models.H1, backend="remote", base_url="http://localhost:8080")
from bioptimus.models.backbones import Backbone
from bioptimus.models.types import Models

print(Backbone.available_backbones())   # ['h1', 'm-optimus', 'tissue-seg']
model = Backbone(Models.M_OPTIMUS, backend="remote", base_url="http://localhost:8080")
For M-Optimus, gene sets are fetched from the server automatically (model.input_gene_names, model.output_gene_names).

Guides

Inference pipeline

One-object pipeline, workspaces, reproducible config.

Cohorts

Multi-slide cohorts and late-binding bulk RNA.

Spatial transcriptomics

M-Optimus gene-expression prediction end to end.

Tile embeddings & PCA

Extract embeddings and visualize morphology.

WSI processing

Read slides: levels, MPP, regions, thumbnails.

Visualizing results

Load Zarr/HDF5/NPZ and overlay genes and masks.

Output formats

FormatExtensionNotes
OutputFormat.ZARR.zarrDefault. Directory store, memory-efficient
OutputFormat.HDF5.h5Single file, memory-efficient
OutputFormat.NPZ.npzAccumulates in memory, compressed on close
Every output file contains the same contents:
  • Datasets: outputs, coords, tissue_ratios, thumbnail, tissue_mask — plus input_gene_names and output_gene_names for M-Optimus.
  • Metadata attributes: slide_name, tile_size, stride, mpp, slide_dimensions, slide_dimensions_at_mpp, num_tiles.
See Visualizing results to load and plot them.