> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bioptimus.com/llms.txt
> Use this file to discover all available pages before exploring further.

# bioptimus.models.clients.local

In-process model inference client.

Satisfies the `Client` protocol by running model inference directly in the current process, without requiring the separate FastAPI microservice. This enables local GPU inference for notebooks, CI/CD pipelines, and researcher workflows with zero infrastructure overhead.

## Serialization contract

The [`Client`](/sdk-reference/models/clients#client) protocol exchanges JSON strings (`str`). The `InProcessClient` deserializes each request body into a `ModelRequest`, runs the server-model pipeline (preprocess → GPU forward → postprocess), and returns a `ModelResponse` serialized to JSON.

This double-serialization preserves drop-in compatibility with the remote backends. The overhead is negligible compared to eliminating network I/O.

## Batching

Two modes are supported:

* **Dynamic batching** (`dynamic_batching=True`, default): An internal `BatchScheduler` accumulates concurrent `predict_async` / `embed_async` calls into GPU batches, exactly matching the behavior of the FastAPI server. One scheduler is created per asyncio event loop (to support multi-WSI processing across threads).

* **Explicit batching** (`dynamic_batching=False`): Callers use `predict_batch` / `embed_batch` to submit pre-formed batches. Async calls run at batch-size-1 (no accumulation).

## InProcessClient

```python theme={null}
class InProcessClient(model: BaseServerModel,
                      max_batch_size: int = 48,
                      max_wait_ms: float = 30.0,
                      dynamic_batching: bool = True)
```

Runs model inference in-process using a loaded .pt2 checkpoint.

Satisfies the `Client` protocol.

<ParamField body="model">
  A loaded server-model instance (or subclass).
</ParamField>

<ParamField body="max_batch_size">
  Maximum tiles per GPU forward pass for dynamic batching.
</ParamField>

<ParamField body="max_wait_ms">
  Maximum milliseconds to wait for additional requests before dispatching a partial batch.
</ParamField>

<ParamField body="dynamic_batching">
  When *True* (default), async calls are accumulated into GPU batches via an internal scheduler. When *False*, each async call runs immediately at batch size 1.
</ParamField>

**Example**:

```python theme={null}
from bioptimus.runtime import build_server_model
from bioptimus.models.clients.local import InProcessClient

model = build_server_model("h1", "/path/to/h1.pt2", device="cuda")
client = InProcessClient(model)
response_json = client.embed(request_json)
```

***

#### predict

```python theme={null}
@log_failures
def predict(body: str) -> str
```

Runs a prediction request synchronously.

Deserializes the JSON body, runs the server-model pipeline (prepare\_sample → forward\_batch), and returns the response as a JSON string.

<ParamField body="body">
  Serialized JSON request payload.
</ParamField>

**Returns**:

Response body as a JSON string.

***

#### embed

```python theme={null}
@log_failures
def embed(body: str) -> str
```

Runs an embedding request synchronously.

<ParamField body="body">
  Serialized JSON request payload.
</ParamField>

**Returns**:

Response body as a JSON string.

***

#### predict\_with\_embedding

```python theme={null}
@log_failures
def predict_with_embedding(body: str) -> str
```

Runs a combined prediction+embedding request synchronously.

Returns both the gene prediction and the tile embedding from a single forward pass (the embedding is a free byproduct of the prediction backbone), so callers needing both avoid a second inference.

<ParamField body="body">
  Serialized JSON request payload.
</ParamField>

**Returns**:

Response body as a JSON string, with `embedding` populated
alongside `output`.

***

#### metadata

```python theme={null}
@log_failures
def metadata() -> str
```

Returns model metadata as a JSON string.

**Returns**:

JSON-serialized model metadata dictionary.

***

#### predict\_async

```python theme={null}
@log_failures
async def predict_async(body: str, session: Any = None) -> str
```

Runs a prediction request asynchronously.

When dynamic batching is enabled, the request enters a per-loop batch scheduler. Otherwise it runs immediately.

<ParamField body="body">
  Serialized JSON request payload.
</ParamField>

<ParamField body="session">
  Unused. Accepted for interface compatibility.
</ParamField>

**Returns**:

Response body as a JSON string.

***

#### embed\_async

```python theme={null}
@log_failures
async def embed_async(body: str, session: Any = None) -> str
```

Runs an embedding request asynchronously.

<ParamField body="body">
  Serialized JSON request payload.
</ParamField>

<ParamField body="session">
  Unused. Accepted for interface compatibility.
</ParamField>

**Returns**:

Response body as a JSON string.

***

#### predict\_with\_embedding\_async

```python theme={null}
@log_failures
async def predict_with_embedding_async(body: str, session: Any = None) -> str
```

Runs a combined prediction+embedding request asynchronously.

When dynamic batching is enabled, the request enters the combined-mode per-loop batch scheduler. Otherwise it runs immediately.

<ParamField body="body">
  Serialized JSON request payload.
</ParamField>

<ParamField body="session">
  Unused. Accepted for interface compatibility.
</ParamField>

**Returns**:

Response body as a JSON string, with `embedding` populated
alongside `output`.

***

#### predict\_batch

```python theme={null}
@log_failures
def predict_batch(bodies: list[str]) -> list[str]
```

Runs a batch of prediction requests synchronously.

Bypasses dynamic batching — the caller controls the batch composition directly.

<ParamField body="bodies">
  List of serialized JSON request payloads.
</ParamField>

**Returns**:

List of response JSON strings, positionally aligned.

***

#### embed\_batch

```python theme={null}
@log_failures
def embed_batch(bodies: list[str]) -> list[str]
```

Runs a batch of embedding requests synchronously.

<ParamField body="bodies">
  List of serialized JSON request payloads.
</ParamField>

**Returns**:

List of response JSON strings, positionally aligned.

***

#### predict\_with\_embedding\_batch

```python theme={null}
@log_failures
def predict_with_embedding_batch(bodies: list[str]) -> list[str]
```

Runs a batch of combined prediction+embedding requests synchronously.

Bypasses dynamic batching — the caller controls the batch composition directly.

<ParamField body="bodies">
  List of serialized JSON request payloads.
</ParamField>

**Returns**:

List of response JSON strings, positionally aligned, each with
`embedding` populated alongside `output`.

***

#### from\_checkpoint

```python theme={null}
@classmethod
@log_failures
def from_checkpoint(cls,
                    model_name: str,
                    checkpoint: str | Path,
                    device: str = "cuda",
                    *,
                    assets_root: str | Path | None = None,
                    max_batch_size: int = 48,
                    max_wait_ms: float = 30.0,
                    dynamic_batching: bool = True) -> InProcessClient
```

Creates an InProcessClient from a checkpoint path.

Convenience factory that combines [`build_server_model`](/sdk-reference/runtime/loader#build_server_model) with client construction.

<ParamField body="model_name">
  Model registry key (e.g. `"h1"`).
</ParamField>

<ParamField body="checkpoint">
  Path to the `.pt2` file.
</ParamField>

<ParamField body="device">
  Torch device string.
</ParamField>

<ParamField body="assets_root">
  Assets directory for models that need it.
</ParamField>

<ParamField body="max_batch_size">
  Maximum batch size for dynamic batching.
</ParamField>

<ParamField body="max_wait_ms">
  Maximum wait time for batch accumulation.
</ParamField>

<ParamField body="dynamic_batching">
  Whether to enable dynamic batching.
</ParamField>

**Returns**:

A ready-to-use [`InProcessClient`](/sdk-reference/models/clients/local#inprocessclient).
