timm models.
Loads an embedding backbone (h0-mini, h0, h1) from the Hugging Face Hub via timm and exposes the same JSON-string transport contract as the HTTP, AWS, and local clients — so it plugs into EndpointModel unchanged.
No running server, Docker image, or AWS account is required — only network access to the Hub and acceptance of each model’s gated license. The build recipe (Hub repository and timm keyword arguments) lives in the SDK model configs under the huggingface block (see bioptimus.models.config_loader); only the plain embedding backbones declare such a block.
Pooling is derived from the model output itself: backbones that already return a pooled (N, embedding_dim) feature (h0, h1) are used as-is, while token-sequence outputs (h0-mini) are pooled according to the config output_type (e.g. cls+patch_mean).
hf_supported_models
huggingface block.
Returns:
Sorted list of supported model names.
HuggingFaceClient
timm keyword arguments) is taken wholesale from the model’s SDK config, so the config stays the single source of truth and inconsistent spec/recipe combinations cannot be constructed. Prefer build_hf_client (or Backbone(..., backend="huggingface")) as the supported construction path.
The model’s SDK config. Must declare a
huggingface build recipe (config.huggingface).Torch device string.
Inference dtype. On CUDA, reduced-precision dtypes (
float16/bfloat16) enable autocast; float32 runs in full precision.Optional preloaded model, primarily for tests.
ValueError- Ifconfighas nohuggingfacebuild recipe.
predict
predict and embed output contract so generic callers remain compatible.
Serialized
ModelRequest JSON string.ModelResponse JSON string.
predict_with_embedding
Serialized
ModelRequest JSON string.NotImplementedError- Always, since these backbones produce only an embedding.
embed
Serialized
ModelRequest JSON string.ModelResponse JSON string.
metadata
predict_async
Serialized
ModelRequest JSON string.Unused; accepted for Client-protocol parity.
ModelResponse JSON string.
embed_async
Serialized
ModelRequest JSON string.Unused; accepted for Client-protocol parity.
ModelResponse JSON string.
predict_with_embedding_async
Serialized
ModelRequest JSON string.Unused; accepted for Client-protocol parity.
NotImplementedError- Always; seepredict_with_embedding.
build_hf_client
huggingface build recipe) and constructs a ready-to-use HuggingFaceClient.
Registry key (
"h0-mini", "h0" or "h1").Torch device string (
"cuda", "cuda:0", "cpu").Inference precision —
"fp32" (default) or "fp16".Optional Hugging Face access token for gated models. When omitted, a cached login or the
HF_TOKEN environment variable is used.HuggingFaceClient.
Raises:
KeyError- If the model has no SDK config.ValueError- If the model is not a supported embedding backbone or ifprecisionis invalid.

