Client protocol by running model inference directly in the current process, without requiring the separate FastAPI microservice. This enables local GPU inference for notebooks, CI/CD pipelines, and researcher workflows with zero infrastructure overhead.
Serialization contract
TheClient protocol exchanges JSON strings (str). The InProcessClient deserializes each request body into a ModelRequest, runs the server-model pipeline (preprocess → GPU forward → postprocess), and returns a ModelResponse serialized to JSON.
This double-serialization preserves drop-in compatibility with the remote backends. The overhead is negligible compared to eliminating network I/O.
Batching
Two modes are supported:-
Dynamic batching (
dynamic_batching=True, default): An internalBatchScheduleraccumulates concurrentpredict_async/embed_asynccalls into GPU batches, exactly matching the behavior of the FastAPI server. One scheduler is created per asyncio event loop (to support multi-WSI processing across threads). -
Explicit batching (
dynamic_batching=False): Callers usepredict_batch/embed_batchto submit pre-formed batches. Async calls run at batch-size-1 (no accumulation).
InProcessClient
Client protocol.
A loaded server-model instance (or subclass).
Maximum tiles per GPU forward pass for dynamic batching.
Maximum milliseconds to wait for additional requests before dispatching a partial batch.
When True (default), async calls are accumulated into GPU batches via an internal scheduler. When False, each async call runs immediately at batch size 1.
predict
Serialized JSON request payload.
embed
Serialized JSON request payload.
predict_with_embedding
Serialized JSON request payload.
embedding populated
alongside output.
metadata
predict_async
Serialized JSON request payload.
Unused. Accepted for interface compatibility.
embed_async
Serialized JSON request payload.
Unused. Accepted for interface compatibility.
predict_with_embedding_async
Serialized JSON request payload.
Unused. Accepted for interface compatibility.
embedding populated
alongside output.
predict_batch
List of serialized JSON request payloads.
embed_batch
List of serialized JSON request payloads.
predict_with_embedding_batch
List of serialized JSON request payloads.
embedding populated alongside output.
from_checkpoint
build_server_model with client construction.
Model registry key (e.g.
"h1").Path to the
.pt2 file.Torch device string.
Assets directory for models that need it.
Maximum batch size for dynamic batching.
Maximum wait time for batch accumulation.
Whether to enable dynamic batching.
InProcessClient.
