> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bioptimus.com/llms.txt
> Use this file to discover all available pages before exploring further.

# bioptimus.models.clients.hf

Hugging Face client for in-process `timm` models.

Loads an embedding backbone (`h0-mini`, `h0`, `h1`) from the Hugging Face Hub via `timm` and exposes the same JSON-string transport contract as the HTTP, AWS, and local clients — so it plugs into `EndpointModel` unchanged.

No running server, Docker image, or AWS account is required — only network access to the Hub and acceptance of each model's gated license. The build recipe (Hub repository and `timm` keyword arguments) lives in the SDK model configs under the `huggingface` block (see `bioptimus.models.config_loader`); only the plain embedding backbones declare such a block.

Pooling is derived from the model output itself: backbones that already return a pooled `(N, embedding_dim)` feature (`h0`, `h1`) are used as-is, while token-sequence outputs (`h0-mini`) are pooled according to the config `output_type` (e.g. `cls+patch_mean`).

***

#### hf\_supported\_models

```python theme={null}
def hf_supported_models() -> list[str]
```

Returns the sorted model names that can be built from the Hub.

A model is supported when its SDK config declares a `huggingface` block.

**Returns**:

Sorted list of supported model names.

## HuggingFaceClient

```python theme={null}
class HuggingFaceClient(config: ModelConfig,
                        *,
                        device: str = 'cuda',
                        dtype: torch.dtype = torch.float16,
                        model: torch.nn.Module | None = None)
```

Runs a Hugging Face Hub embedding backbone in-process.

The build recipe (Hub repository and `timm` keyword arguments) is taken wholesale from the model's SDK config, so the config stays the single source of truth and inconsistent spec/recipe combinations cannot be constructed. Prefer [`build_hf_client`](/sdk-reference/models/clients/hf#build_hf_client) (or `Backbone(..., backend="huggingface")`) as the supported construction path.

<ParamField body="config">
  The model's SDK config. Must declare a `huggingface` build recipe (`config.huggingface`).
</ParamField>

<ParamField body="device">
  Torch device string.
</ParamField>

<ParamField body="dtype">
  Inference dtype. On CUDA, reduced-precision dtypes (`float16`/`bfloat16`) enable autocast; `float32` runs in full precision.
</ParamField>

<ParamField body="model">
  Optional preloaded model, primarily for tests.
</ParamField>

**Raises**:

* `ValueError` - If `config` has no `huggingface` build recipe.

***

#### predict

```python theme={null}
def predict(body: str) -> str
```

Runs prediction synchronously.

Embedding-only Hugging Face backbones share the `predict` and `embed` output contract so generic callers remain compatible.

<ParamField body="body">
  Serialized [`ModelRequest`](/sdk-reference/inference/schemas#modelrequest) JSON string.
</ParamField>

**Returns**:

Serialized [`ModelResponse`](/sdk-reference/inference/schemas#modelresponse) JSON string.

***

#### predict\_with\_embedding

```python theme={null}
def predict_with_embedding(body: str) -> str
```

Not supported for embedding-only Hugging Face backbones.

The combined mode returns a spatial-transcriptomics prediction alongside the embedding, but Hugging Face backbones are pure embedders with no prediction head, so there is no second output to pair.

<ParamField body="body">
  Serialized [`ModelRequest`](/sdk-reference/inference/schemas#modelrequest) JSON string.
</ParamField>

**Returns**:

Never returns; this method always raises.

**Raises**:

* `NotImplementedError` - Always, since these backbones produce only an embedding.

***

#### embed

```python theme={null}
def embed(body: str) -> str
```

Runs embedding synchronously.

<ParamField body="body">
  Serialized [`ModelRequest`](/sdk-reference/inference/schemas#modelrequest) JSON string.
</ParamField>

**Returns**:

Serialized [`ModelResponse`](/sdk-reference/inference/schemas#modelresponse) JSON string.

***

#### metadata

```python theme={null}
def metadata() -> str
```

Returns model metadata as a JSON string.

**Returns**:

JSON string with the model name, version, weights source, output
configuration, and the *effective* runtime precision (derived
from the dtype the client runs with, not the config default).

***

#### predict\_async

```python theme={null}
async def predict_async(body: str, session: Any = None) -> str
```

Runs prediction asynchronously via the event loop executor.

<ParamField body="body">
  Serialized [`ModelRequest`](/sdk-reference/inference/schemas#modelrequest) JSON string.
</ParamField>

<ParamField body="session">
  Unused; accepted for Client-protocol parity.
</ParamField>

**Returns**:

Serialized [`ModelResponse`](/sdk-reference/inference/schemas#modelresponse) JSON string.

***

#### embed\_async

```python theme={null}
async def embed_async(body: str, session: Any = None) -> str
```

Runs embedding asynchronously via the event loop executor.

<ParamField body="body">
  Serialized [`ModelRequest`](/sdk-reference/inference/schemas#modelrequest) JSON string.
</ParamField>

<ParamField body="session">
  Unused; accepted for Client-protocol parity.
</ParamField>

**Returns**:

Serialized [`ModelResponse`](/sdk-reference/inference/schemas#modelresponse) JSON string.

***

#### predict\_with\_embedding\_async

```python theme={null}
async def predict_with_embedding_async(body: str, session: Any = None) -> str
```

Not supported for embedding-only Hugging Face backbones.

<ParamField body="body">
  Serialized [`ModelRequest`](/sdk-reference/inference/schemas#modelrequest) JSON string.
</ParamField>

<ParamField body="session">
  Unused; accepted for Client-protocol parity.
</ParamField>

**Returns**:

Never returns; this method always raises.

**Raises**:

* `NotImplementedError` - Always; see [`predict_with_embedding`](/sdk-reference/models/clients/hf#predict_with_embedding).

***

#### build\_hf\_client

```python theme={null}
def build_hf_client(model_name: str,
                    device: str = "cuda",
                    *,
                    precision: str = "fp32",
                    hf_token: str | None = None) -> HuggingFaceClient
```

Builds an in-process client backed by a Hugging Face backbone.

Reads the model's SDK config (spec + `huggingface` build recipe) and constructs a ready-to-use [`HuggingFaceClient`](/sdk-reference/models/clients/hf#huggingfaceclient).

<ParamField body="model_name">
  Registry key (`"h0-mini"`, `"h0"` or `"h1"`).
</ParamField>

<ParamField body="device">
  Torch device string (`"cuda"`, `"cuda:0"`, `"cpu"`).
</ParamField>

<ParamField body="precision">
  Inference precision — `"fp32"` (default) or `"fp16"`.
</ParamField>

<ParamField body="hf_token">
  Optional Hugging Face access token for gated models. When omitted, a cached login or the `HF_TOKEN` environment variable is used.
</ParamField>

**Returns**:

A ready-to-use [`HuggingFaceClient`](/sdk-reference/models/clients/hf#huggingfaceclient).

**Raises**:

* `KeyError` - If the model has no SDK config.
* `ValueError` - If the model is not a supported embedding backbone or if `precision` is invalid.
