Skip to main content
Hugging Face client for in-process timm models. Loads an embedding backbone (h0-mini, h0, h1) from the Hugging Face Hub via timm and exposes the same JSON-string transport contract as the HTTP, AWS, and local clients — so it plugs into EndpointModel unchanged. No running server, Docker image, or AWS account is required — only network access to the Hub and acceptance of each model’s gated license. The build recipe (Hub repository and timm keyword arguments) lives in the SDK model configs under the huggingface block (see bioptimus.models.config_loader); only the plain embedding backbones declare such a block. Pooling is derived from the model output itself: backbones that already return a pooled (N, embedding_dim) feature (h0, h1) are used as-is, while token-sequence outputs (h0-mini) are pooled according to the config output_type (e.g. cls+patch_mean).

hf_supported_models

Returns the sorted model names that can be built from the Hub. A model is supported when its SDK config declares a huggingface block. Returns: Sorted list of supported model names.

HuggingFaceClient

Runs a Hugging Face Hub embedding backbone in-process. The build recipe (Hub repository and timm keyword arguments) is taken wholesale from the model’s SDK config, so the config stays the single source of truth and inconsistent spec/recipe combinations cannot be constructed. Prefer build_hf_client (or Backbone(..., backend="huggingface")) as the supported construction path.
The model’s SDK config. Must declare a huggingface build recipe (config.huggingface).
Torch device string.
Inference dtype. On CUDA, reduced-precision dtypes (float16/bfloat16) enable autocast; float32 runs in full precision.
Optional preloaded model, primarily for tests.
Raises:
  • ValueError - If config has no huggingface build recipe.

predict

Runs prediction synchronously. Embedding-only Hugging Face backbones share the predict and embed output contract so generic callers remain compatible.
Serialized ModelRequest JSON string.
Returns: Serialized ModelResponse JSON string.

predict_with_embedding

Not supported for embedding-only Hugging Face backbones. The combined mode returns a spatial-transcriptomics prediction alongside the embedding, but Hugging Face backbones are pure embedders with no prediction head, so there is no second output to pair.
Serialized ModelRequest JSON string.
Returns: Never returns; this method always raises. Raises:
  • NotImplementedError - Always, since these backbones produce only an embedding.

embed

Runs embedding synchronously.
Serialized ModelRequest JSON string.
Returns: Serialized ModelResponse JSON string.

metadata

Returns model metadata as a JSON string. Returns: JSON string with the model name, version, weights source, output configuration, and the effective runtime precision (derived from the dtype the client runs with, not the config default).

predict_async

Runs prediction asynchronously via the event loop executor.
Serialized ModelRequest JSON string.
Unused; accepted for Client-protocol parity.
Returns: Serialized ModelResponse JSON string.

embed_async

Runs embedding asynchronously via the event loop executor.
Serialized ModelRequest JSON string.
Unused; accepted for Client-protocol parity.
Returns: Serialized ModelResponse JSON string.

predict_with_embedding_async

Not supported for embedding-only Hugging Face backbones.
Serialized ModelRequest JSON string.
Unused; accepted for Client-protocol parity.
Returns: Never returns; this method always raises. Raises:

build_hf_client

Builds an in-process client backed by a Hugging Face backbone. Reads the model’s SDK config (spec + huggingface build recipe) and constructs a ready-to-use HuggingFaceClient.
Registry key ("h0-mini", "h0" or "h1").
Torch device string ("cuda", "cuda:0", "cpu").
Inference precision — "fp32" (default) or "fp16".
Optional Hugging Face access token for gated models. When omitted, a cached login or the HF_TOKEN environment variable is used.
Returns: A ready-to-use HuggingFaceClient. Raises:
  • KeyError - If the model has no SDK config.
  • ValueError - If the model is not a supported embedding backbone or if precision is invalid.