Skip to content

Developing a Service with a Servable Model ​

To make a model deployable through ALIDA (see Model Deployment and Usage), the training Service must produce two mandatory elements in the --output-model folder:

  1. The model artifacts in the correct physical format for the chosen inference runtime
  2. The model-settings.json file that tells MLServer how to load the model and which input/output schema to expose

The model-settings.json File ​

The model-settings.json file must be created by the developer and saved in the --output-model folder. It contains:

  • name — model name, must follow the ALIDA convention model-{output_model_id}
  • implementation — Python path to the MLServer runtime to use (e.g. mlserver_sklearn.SKLearnModel)
  • parameters — runtime configuration:
    • uri — relative path to the model artifact
    • content_type — (optional) default content type for payload decoding (see Content type)
    • extra — (optional) runtime-specific extra parameters (e.g. task for HuggingFace)
  • inputs — input schema: name, V2 type and shape of each feature
  • outputs — output schema: name, V2 type and shape of each prediction

The inputs and outputs fields are what ALIDA displays as Model Metadata on the model detail page and what the user uses to build inference requests. It is the developer's responsibility to fill them in correctly.

General Structure ​

json
{
  "name": "model-<output_model_id>",
  "implementation": "<runtime-mlserver>",
  "parameters": {
    "uri": "<percorso-artefatto-relativo>"
  },
  "inputs": [
    {
      "name": "<nome-feature>",
      "datatype": "<tipo-v2>",
      "shape": [-1]
    }
  ],
  "outputs": [
    {
      "name": "predict",
      "datatype": "<tipo-v2>",
      "shape": [-1, 1]
    }
  ]
}

The shape uses -1 as a placeholder for the batch dimension (variable number of samples). For a scalar feature per sample use [-1]; for a vector feature use [-1, N].

outputs as the Exposed Contract ​

The outputs field of model-settings.json defines the contract exposed by the model: everything declared there appears in the Model Metadata on the ALIDA model detail page, and is what the user sees and can request.

The developer has full freedom over which outputs to declare. For example, for a sklearn classifier one can choose to expose only the predicted class (predict), or also the per-class probabilities (predict_proba), or both:

json
"outputs": [
  { "name": "predict",       "datatype": "INT64", "shape": [-1, 1] },
  { "name": "predict_proba", "datatype": "FP64",  "shape": [-1, 3] }
]

The user can then explicitly request one of the declared outputs in the inference payload:

json
{
  "inputs": [...],
  "outputs": [{ "name": "predict_proba" }]
}

"Note — available outputs by runtime"

The output names that MLServer recognises depend on the runtime:
RuntimeAvailable outputsNotes
mlserver_sklearn.SKLearnModelpredict, predict_proba, transformpredict_proba for classifiers only; transform for pipelines only; predict is the default
mlserver_xgboost.XGBoostModelpredictSingle output
mlserver_lightgbm.LightGBMModelpredictSingle output
mlserver_catboost.CatBoostModelpredictSingle output
mlserver_mlflow.MLflowRuntimepredictCalls MLflow pyfunc predict(); does not expose predict_proba separately
mlserver_huggingface.HuggingFaceRuntimedepends on taskOutput returned in the configured HuggingFace task format
custom (mlserver.MLModel)defined by developerDeveloper builds the InferenceResponse in the predict() method

For regression never declare predict_proba in outputs regardless of the runtime.

Content Type ​

The content_type determines how MLServer decodes the V2 payload before passing it to the model. The available content types are:

Content typeResulting Python typeRequest levelInput level
npnumpy.ndarray✅✅
pdpandas.DataFrame✅❌
strstr (UTF-8)✅✅
base64byte decodificati da base64❌✅
datetimedatetime.datetime❌✅

With np, the data field is flat and shape is used to reshape the tensor. Suitable for models that accept a pure numeric array with a fixed column order.

With pd, inputs are aggregated into a pandas.DataFrame, where each input name becomes a column name. Suitable for tabular sklearn pipelines that use column names to apply different transformations. pd operates only at request level and can be combined with input-level content types (e.g. np for numeric columns, str for string columns).

The content type can be declared in model-settings.json (in the parameters at request level, or in each input's parameters section). This becomes the default: requests that do not specify an explicit content type will use the one from the metadata. Content types explicitly specified in the request always take precedence over those in the metadata.

"Note — sklearn default content type"

If no content type is declared in the metadata or in the request, the sklearn runtime decodes the payload as a NumPy array. To use Pandas DataFrame, the developer must explicitly declare `"content_type": "pd"` in the `parameters` of `model-settings.json`.

Supported Data Types (V2 Inference Protocol) ​

TypeDescriptionCorresponding pandas type
BOOLBooleanbool
INT8Signed 8-bit integerint8
INT16Signed 16-bit integerint16
INT32Signed 32-bit integerint32
INT64Signed 64-bit integerint64
UINT8Unsigned 8-bit integeruint8
UINT16Unsigned 16-bit integeruint16
UINT32Unsigned 32-bit integeruint32
UINT64Unsigned 64-bit integeruint64
FP1616-bit floating pointfloat16
FP3232-bit floating pointfloat32
FP6464-bit floating pointfloat64
BYTESBytes / string / categorical / objectobject, category
STRINGUTF-8 stringstring

Generating model-settings.json ​

It is recommended to generate the file programmatically from the training data, deriving V2 types from the pandas dtypes of X and y. Below is a utility function:

python
import json
from pathlib import Path
import pandas as pd


def to_v2_dtype(pd_dtype) -> str:
    """Maps a pandas dtype to the corresponding V2 type."""
    kind = getattr(pd_dtype, "kind", "O")
    if kind == "b":
        return "BOOL"
    if kind in ("i", "u"):
        return "INT64"
    if kind == "f":
        return "FP64" if str(pd_dtype) in ("float64", "Float64") else "FP32"
    return "BYTES"


def create_model_settings(
    model_dir: str,
    X_sample: pd.DataFrame,
    y_sample: pd.Series,
    model_id: str,
    implementation: str,
    uri: str = ".",
    is_regression: bool = False,
    content_type: str = None,
    extra: dict = None,
):
    """
    Generates model-settings.json in the model folder.

    Args:
        model_dir:      ALIDA output-model folder
        X_sample:       DataFrame with features (even a single sample)
        y_sample:       Series with target (even a single sample)
        model_id:       value of args.output_model_id injected by ALIDA
        implementation: MLServer runtime (e.g. "mlserver_sklearn.SKLearnModel")
        uri:            relative path to the model artifact
        is_regression:  True for regression, False for classification
        content_type:   content type for payload decoding (e.g. "pd", "np")
        extra:          extra parameters for the runtime (e.g. {"task": "question-answering"})
    """
    inputs = [
        {
            "name": col,
            "datatype": to_v2_dtype(dtype),
            "shape": [-1]
        }
        for col, dtype in zip(X_sample.columns, X_sample.dtypes)
    ]

    if is_regression:
        out_dtype = to_v2_dtype(y_sample.dtype)
        outputs = [{"name": "predict", "datatype": out_dtype, "shape": [-1, 1]}]
    else:
        y_kind = getattr(y_sample.dtype, "kind", "O")
        if y_kind == "b":
            out_dtype = "BOOL"
        elif y_kind in ("i", "u"):
            out_dtype = "INT64"
        else:
            out_dtype = "BYTES"
        outputs = [{"name": "predict", "datatype": out_dtype, "shape": [-1, 1]}]

    parameters = {"uri": uri}
    if content_type is not None:
        parameters["content_type"] = content_type
    if extra is not None:
        parameters["extra"] = extra

    settings = {
        "name": f"model-{model_id}",
        "implementation": implementation,
        "parameters": parameters,
        "inputs": inputs,
        "outputs": outputs
    }

    settings_path = Path(model_dir) / "model-settings.json"
    with open(settings_path, "w", encoding="utf-8") as f:
        json.dump(settings, f, indent=2, ensure_ascii=False)

    return settings_path

"Note"

The function generates a single `predict` output. To expose additional outputs (e.g. `predict_proba`), the developer can manually add further elements to the `outputs` list after the call, or extend the function as needed.

Expected Output Folder Structure ​

<output-model>/
├── model-settings.json      ← mandatory
└── <artefatto-modello>      ← dipende dal runtime scelto

Supported Runtimes ​

MLServer Runtimeformat in metamodelArtifact format
mlserver_sklearn.SKLearnModelsklearn.skops (recommended), .joblib, .pkl
mlserver_xgboost.XGBoostModelxgboost.json, .bst
mlserver_lightgbm.LightGBMModellightgbm.txt, .bst
mlserver_catboost.CatBoostModelcatboost.cbm
mlserver_mlflow.MLflowRuntimemlflowMLflow folder (with MLmodel)
mlserver_huggingface.HuggingFaceRuntimehuggingfacelocal model or from HuggingFace Hub
custom class extending mlserver.MLModelpythondirectory with .py file + artifact

The format value in the metamodel is used by ALIDA for internal routing to the correct serving runtime and is independent of the implementation field in model-settings.json.


Scikit-learn ​

Runtime: mlserver_sklearn.SKLearnModel

Recommended artifact format: .skops file, produced with the Skops library. Unlike pickle/joblib, Skops does not allow arbitrary code execution during deserialization.

The .joblib and .pkl formats are also supported.

python
import skops.io as sio

os.makedirs(args.output_model, exist_ok=True)
model_path = os.path.join(args.output_model, "model.skops")
sio.dump(pipeline, model_path)

create_model_settings(
    model_dir=args.output_model,
    X_sample=X,
    y_sample=y,
    model_id=args.output_model_id,
    implementation="mlserver_sklearn.SKLearnModel",
    uri="model.skops",
    content_type="pd",
    is_regression=False
)

XGBoost ​

Runtime: mlserver_xgboost.XGBoostModel

Artifact format: .json or .bst file — native XGBoost format (save_model()).

python
model_path = os.path.join(args.output_model, "model.json")
model.save_model(model_path)

create_model_settings(
    model_dir=args.output_model,
    X_sample=X,
    y_sample=y,
    model_id=args.output_model_id,
    implementation="mlserver_xgboost.XGBoostModel",
    uri="model.json",
    is_regression=False
)

LightGBM ​

Runtime: mlserver_lightgbm.LightGBMModel

Artifact format: .txt or .bst file — native LightGBM format (booster_.save_model()).

python
model_path = os.path.join(args.output_model, "model.txt")
model.booster_.save_model(model_path)

create_model_settings(
    model_dir=args.output_model,
    X_sample=X,
    y_sample=y,
    model_id=args.output_model_id,
    implementation="mlserver_lightgbm.LightGBMModel",
    uri="model.txt",
    is_regression=False
)

CatBoost ​

Runtime: mlserver_catboost.CatBoostModel

Artifact format: .cbm file — native CatBoost format (save_model()).

python
model_path = os.path.join(args.output_model, "model.cbm")
model.save_model(model_path)

create_model_settings(
    model_dir=args.output_model,
    X_sample=X,
    y_sample=y,
    model_id=args.output_model_id,
    implementation="mlserver_catboost.CatBoostModel",
    uri="model.cbm",
    is_regression=False
)

MLflow (native format) ​

Runtime: mlserver_mlflow.MLflowRuntime

Artifact format: complete MLflow folder containing the MLmodel file and the flavour artifacts. When the MLflow model defines a signature, MLServer automatically converts it to a V2 metadata schema. The model-settings.json must still be created to specify the model name and implementation.

python
import mlflow

run_id = mlflow.active_run().info.run_id
mlflow.artifacts.download_artifacts(
    run_id=run_id,
    artifact_path="model",
    dst_path=args.output_model
)

create_model_settings(
    model_dir=args.output_model,
    X_sample=X,
    y_sample=y,
    model_id=args.output_model_id,
    implementation="mlserver_mlflow.MLflowRuntime",
    uri=".",
    content_type="pd",
    is_regression=False
)

HuggingFace ​

Runtime: mlserver_huggingface.HuggingFaceRuntime

The HuggingFace runtime supports models from the HuggingFace Transformers library. Configuration is done through the parameters.extra field of model-settings.json:

  • task — (mandatory) the HuggingFace task type (e.g. "question-answering", "text-generation", "text-classification", etc.)
  • pretrained_model — (optional) model name in the HuggingFace Hub; takes precedence over parameters.uri
  • optimum_model — (optional) true to use models optimized with Optimum
  • device — (optional) device to load the model on (0 for GPU, -1 for CPU)

To load a local model, specify the path in parameters.uri. To load a model from the HuggingFace Hub, specify the name in parameters.extra.pretrained_model.

!!! warning "Warning — resources" HuggingFace Transformer models can require significant amounts of RAM and/or GPU memory. Verify that the ALIDA instance has sufficient resources for the chosen model. When registering the Service, declare the required resources through Resource type Service Properties (see Service Registration).

When using `pretrained_model`, the model is downloaded from the HuggingFace Hub at first load — this requires network connectivity and disk space in the container. To avoid delays at first deploy, it is recommended to include the model artifacts directly in the `--output-model` folder.
json
{
  "name": "model-<output_model_id>",
  "implementation": "mlserver_huggingface.HuggingFaceRuntime",
  "parameters": {
    "extra": {
      "task": "text-classification",
      "pretrained_model": "distilbert-base-uncased-finetuned-sst-2-english"
    }
  },
  "inputs": [
    {
      "name": "text_inputs",
      "datatype": "BYTES",
      "shape": [-1]
    }
  ],
  "outputs": [
    {
      "name": "output",
      "datatype": "BYTES",
      "shape": [-1, 1]
    }
  ]
}

Python Custom ​

Runtime: classe custom che estende mlserver.MLModel

For models not covered by standard runtimes, a custom runtime can be implemented. The developer must create a Python file with a class extending mlserver.MLModel that implements the load() and predict() methods:

python
from mlserver import MLModel
from mlserver.types import InferenceRequest, InferenceResponse
from mlserver.codecs import NumpyCodec
import numpy as np


class MyCustomModel(MLModel):

    async def load(self) -> bool:
        # Load the model from self.settings.parameters.uri
        model_uri = self.settings.parameters.uri
        # ... loading logic ...
        self.ready = True
        return self.ready

    async def predict(self, payload: InferenceRequest) -> InferenceResponse:
        # Decode input
        input_data = self.decode(payload.inputs[0])
        # ... inference logic ...
        result = np.array([0])  # example
        return InferenceResponse(
            model_name=self.settings.name,
            outputs=[NumpyCodec.encode_output("predict", result)]
        )

The model-settings.json specifies the runtime as <file_name>.<ClassName>:

json
{
  "name": "model-<output_model_id>",
  "implementation": "my_model.MyCustomModel",
  "parameters": {
    "uri": "."
  },
  "inputs": [...],
  "outputs": [...]
}

The .py file with the class must be in the same folder as model-settings.json (--output-model folder).

"Note"

The `load()` and `predict()` methods are **asynchronous** (`async`). All Python dependencies used by the class must be installed in the *Service* container (declared in `requirements.txt`). The `.py` file must be self-contained or import only modules present in the Docker image.

!!! warning "Warning — preferred formats for serving" Formats based on Python serialization (pickle, joblib) are supported but discouraged for models intended for serving. Prefer the native formats of each framework (Skops for sklearn, .json for XGBoost, etc.) which guarantee safer deserialization.

Metamodel Declaration ​

The output-model port in the Service metamodel must declare the model field with the appropriate format value. This value is used by ALIDA for internal routing to the correct serving runtime and is independent of the implementation field in model-settings.json:

json
{
    "key": "output-model",
    "type": "application",
    "mandatory": true,
    "invisible": true,
    "valueType": "STRING",
    "defaultValue": null,
    "model": { "format": "sklearn" }
}

Complete Example — sklearn Training with Skops and Serving ​

python
import argparse

def str2bool(v):
    if isinstance(v, bool):
        return v
    if v.lower() in ('yes', 'true', 't', 'y', '1'):
        return True
    elif v.lower() in ('no', 'false', 'f', 'n', '0', ''):
        return False

parser = argparse.ArgumentParser()

parser.add_argument('--input-dataset', dest='input_dataset', type=str, required=True)
parser.add_argument('--input-dataset.s3_bucket', dest='input_s3_bucket', type=str, required=True)
parser.add_argument('--input-dataset.s3_URL', dest='input_s3_url', type=str, required=True)
parser.add_argument('--input-dataset.s3_ACCESS_KEY', dest='input_access_key', type=str, required=True)
parser.add_argument('--input-dataset.s3_SECRET_KEY', dest='input_secret_key', type=str, required=True)
parser.add_argument('--input-dataset.s3_REGION', dest='input_dataset_region_name', type=str, required=False)
parser.add_argument('--input-dataset.use_ssl', dest='input_use_ssl', type=str2bool, required=True)
parser.add_argument('--input-columns', dest='input_columns', type=str, required=False)

parser.add_argument('--output-model', dest='output_model', type=str, required=True)
parser.add_argument('--output-model.s3_bucket', dest='output_s3_bucket', type=str, required=True)
parser.add_argument('--output-model.s3_URL', dest='output_s3_url', type=str, required=True)
parser.add_argument('--output-model.s3_ACCESS_KEY', dest='output_access_key', type=str, required=True)
parser.add_argument('--output-model.s3_SECRET_KEY', dest='output_secret_key', type=str, required=True)
parser.add_argument('--output-model.s3_REGION', dest='output_model_region_name', type=str, required=False)
parser.add_argument('--output-model.use_ssl', dest='output_use_ssl', type=str2bool, required=True)
parser.add_argument('--output-model.id', dest='output_model_id', type=str, required=True)

parser.add_argument('--label-column', dest='label_column', type=str, required=True)

args, unknown = parser.parse_known_args()
python
import json
import os
from pathlib import Path

import pandas as pd
import skops.io as sio
from minio import Minio
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

from arguments import args


def to_v2_dtype(pd_dtype) -> str:
    kind = getattr(pd_dtype, "kind", "O")
    if kind == "b":
        return "BOOL"
    if kind in ("i", "u"):
        return "INT64"
    if kind == "f":
        return "FP64" if str(pd_dtype) in ("float64", "Float64") else "FP32"
    return "BYTES"


def create_model_settings(model_dir, X_sample, y_sample, model_id, implementation,
                          uri=".", is_regression=False, content_type=None, extra=None):
    inputs = [
        {"name": col, "datatype": to_v2_dtype(dtype), "shape": [-1]}
        for col, dtype in zip(X_sample.columns, X_sample.dtypes)
    ]
    if is_regression:
        out_dtype = to_v2_dtype(y_sample.dtype)
    else:
        y_kind = getattr(y_sample.dtype, "kind", "O")
        out_dtype = "BOOL" if y_kind == "b" else "INT64" if y_kind in ("i", "u") else "BYTES"
    outputs = [{"name": "predict", "datatype": out_dtype, "shape": [-1, 1]}]

    parameters = {"uri": uri}
    if content_type is not None:
        parameters["content_type"] = content_type
    if extra is not None:
        parameters["extra"] = extra

    settings = {
        "name": f"model-{model_id}",
        "implementation": implementation,
        "parameters": parameters,
        "inputs": inputs,
        "outputs": outputs
    }
    settings_path = Path(model_dir) / "model-settings.json"
    with open(settings_path, "w", encoding="utf-8") as f:
        json.dump(settings, f, indent=2, ensure_ascii=False)
    return settings_path


def s3_ls(address, access_key, secret_key, region, bucket_name, folder, extension, use_ssl=False):
    if not folder.endswith("/"):
        folder = folder + "/"
    cleaned = address.replace("http://", "").replace("https://", "")
    client = Minio(cleaned, access_key=access_key, secret_key=secret_key, secure=use_ssl, region=region)
    objects = client.list_objects(bucket_name=bucket_name, prefix=folder)
    files_list = [x._object_name for x in objects if x._object_name.endswith(extension)]
    if len(files_list) == 0:
        raise Exception("Empty dataset!")
    return "s3://" + bucket_name + "/" + files_list[0]


# Load dataset
storage_options = {
    'key': args.input_access_key,
    'secret': args.input_secret_key,
    'region': args.input_dataset_region_name,
    'client_kwargs': {'endpoint_url': args.input_s3_url}
}
file_path = s3_ls(
    args.input_s3_url, args.input_access_key, args.input_secret_key,
    args.input_dataset_region_name, args.input_s3_bucket, args.input_dataset, ".csv"
)
dataset = pd.read_csv(file_path, storage_options=storage_options, sep=None, engine='python')

if args.input_columns is not None and args.input_columns.strip() != '*':
    dataset = dataset[[c.strip() for c in args.input_columns.split(",")]].copy()

X = dataset.drop(columns=[args.label_column])
y = dataset[args.label_column]

# Training
pipeline = Pipeline([
    ("scaler", StandardScaler()),
    ("clf", LogisticRegression())
])
pipeline.fit(X, y)

# Save model in Skops format
os.makedirs(args.output_model, exist_ok=True)
model_path = os.path.join(args.output_model, "model.skops")
sio.dump(pipeline, model_path)

# Generate model-settings.json
create_model_settings(
    model_dir=args.output_model,
    X_sample=X,
    y_sample=y,
    model_id=args.output_model_id,
    implementation="mlserver_sklearn.SKLearnModel",
    uri="model.skops",
    content_type="pd",
    is_regression=False
)

print(f"Model saved to {args.output_model}", flush=True)

The corresponding metamodel declares the output port with format: "sklearn":

json
{
    "key": "output-model",
    "type": "application",
    "mandatory": true,
    "invisible": true,
    "valueType": "STRING",
    "defaultValue": null,
    "model": { "format": "sklearn" }
}