Skip to main content
In-process model inference client. Satisfies the Client protocol by running model inference directly in the current process, without requiring the separate FastAPI microservice. This enables local GPU inference for notebooks, CI/CD pipelines, and researcher workflows with zero infrastructure overhead.

Serialization contract

The Client protocol exchanges JSON strings (str). The InProcessClient deserializes each request body into a ModelRequest, runs the server-model pipeline (preprocess → GPU forward → postprocess), and returns a ModelResponse serialized to JSON. This double-serialization preserves drop-in compatibility with the remote backends. The overhead is negligible compared to eliminating network I/O.

Batching

Two modes are supported:
  • Dynamic batching (dynamic_batching=True, default): An internal BatchScheduler accumulates concurrent predict_async / embed_async calls into GPU batches, exactly matching the behavior of the FastAPI server. One scheduler is created per asyncio event loop (to support multi-WSI processing across threads).
  • Explicit batching (dynamic_batching=False): Callers use predict_batch / embed_batch to submit pre-formed batches. Async calls run at batch-size-1 (no accumulation).

InProcessClient

Runs model inference in-process using a loaded .pt2 checkpoint. Satisfies the Client protocol.
A loaded server-model instance (or subclass).
Maximum tiles per GPU forward pass for dynamic batching.
Maximum milliseconds to wait for additional requests before dispatching a partial batch.
When True (default), async calls are accumulated into GPU batches via an internal scheduler. When False, each async call runs immediately at batch size 1.
Example:

predict

Runs a prediction request synchronously. Deserializes the JSON body, runs the server-model pipeline (prepare_sample → forward_batch), and returns the response as a JSON string.
Serialized JSON request payload.
Returns: Response body as a JSON string.

embed

Runs an embedding request synchronously.
Serialized JSON request payload.
Returns: Response body as a JSON string.

predict_with_embedding

Runs a combined prediction+embedding request synchronously. Returns both the gene prediction and the tile embedding from a single forward pass (the embedding is a free byproduct of the prediction backbone), so callers needing both avoid a second inference.
Serialized JSON request payload.
Returns: Response body as a JSON string, with embedding populated alongside output.

metadata

Returns model metadata as a JSON string. Returns: JSON-serialized model metadata dictionary.

predict_async

Runs a prediction request asynchronously. When dynamic batching is enabled, the request enters a per-loop batch scheduler. Otherwise it runs immediately.
Serialized JSON request payload.
Unused. Accepted for interface compatibility.
Returns: Response body as a JSON string.

embed_async

Runs an embedding request asynchronously.
Serialized JSON request payload.
Unused. Accepted for interface compatibility.
Returns: Response body as a JSON string.

predict_with_embedding_async

Runs a combined prediction+embedding request asynchronously. When dynamic batching is enabled, the request enters the combined-mode per-loop batch scheduler. Otherwise it runs immediately.
Serialized JSON request payload.
Unused. Accepted for interface compatibility.
Returns: Response body as a JSON string, with embedding populated alongside output.

predict_batch

Runs a batch of prediction requests synchronously. Bypasses dynamic batching — the caller controls the batch composition directly.
List of serialized JSON request payloads.
Returns: List of response JSON strings, positionally aligned.

embed_batch

Runs a batch of embedding requests synchronously.
List of serialized JSON request payloads.
Returns: List of response JSON strings, positionally aligned.

predict_with_embedding_batch

Runs a batch of combined prediction+embedding requests synchronously. Bypasses dynamic batching — the caller controls the batch composition directly.
List of serialized JSON request payloads.
Returns: List of response JSON strings, positionally aligned, each with embedding populated alongside output.

from_checkpoint

Creates an InProcessClient from a checkpoint path. Convenience factory that combines build_server_model with client construction.
Model registry key (e.g. "h1").
Path to the .pt2 file.
Torch device string.
Assets directory for models that need it.
Maximum batch size for dynamic batching.
Maximum wait time for batch accumulation.
Whether to enable dynamic batching.
Returns: A ready-to-use InProcessClient.