- Deserialization (parallel, per-request): Each incoming request is decoded, transformed, and validated in a thread pool. Failures are returned immediately as per-request errors (HTTP 400) without entering the batch queue.
-
Batching (GPU, shared): Successfully prepared samples are accumulated into batches and dispatched to the GPU. A GPU-level failure (OOM, CUDA error) is shared by all requests in the batch and mapped to HTTP 503 via
_ForwardError.
BatchScheduler
predict call submits the request to a thread pool for image decoding, transform, shape validation, and (for M-Optimus) bulk RNA preprocessing. If any step fails, the request’s future is immediately resolved with a ValueError (→ HTTP 400) and the request never enters the batch queue.
Stage 2 — Batching (shared GPU): Valid PreparedSample objects are collected into batches of up to max_batch_size and dispatched to the GPU. A forward failure raises _ForwardError (a RuntimeError subclass → HTTP 503) to all requests in that batch.
The server-side model to batch requests for.
Maximum tiles per GPU forward pass.
Maximum milliseconds to wait before dispatching a partial batch.
Thread pool size for parallel deserialization.
"prediction", "embedding", or "prediction_with_embedding".shutdown
predict
The tile request to process.

