Developing a Service with a Servable Model
To make a model deployable through ALIDA (see Model Deployment and Usage), the training Service must produce two mandatory elements in the --output-model folder:
- The model artifacts in the correct physical format for the chosen inference runtime
- The
model-settings.jsonfile that tells MLServer how to load the model and which input/output schema to expose
The model-settings.json File
The model-settings.json file must be created by the developer and saved in the --output-model folder. It contains:
name— model name, must follow the ALIDA conventionmodel-{output_model_id}implementation— Python path to the MLServer runtime to use (e.g.mlserver_sklearn.SKLearnModel)parameters— runtime configuration:uri— relative path to the model artifactcontent_type— (optional) default content type for payload decoding (see Content type)extra— (optional) runtime-specific extra parameters (e.g.taskfor HuggingFace)
inputs— input schema: name, V2 type and shape of each featureoutputs— output schema: name, V2 type and shape of each prediction
The inputs and outputs fields are what ALIDA displays as Model Metadata on the model detail page and what the user uses to build inference requests. It is the developer's responsibility to fill them in correctly.
General Structure
{
"name": "model-<output_model_id>",
"implementation": "<runtime-mlserver>",
"parameters": {
"uri": "<percorso-artefatto-relativo>"
},
"inputs": [
{
"name": "<nome-feature>",
"datatype": "<tipo-v2>",
"shape": [-1]
}
],
"outputs": [
{
"name": "predict",
"datatype": "<tipo-v2>",
"shape": [-1, 1]
}
]
}The shape uses -1 as a placeholder for the batch dimension (variable number of samples). For a scalar feature per sample use [-1]; for a vector feature use [-1, N].
outputs as the Exposed Contract
The outputs field of model-settings.json defines the contract exposed by the model: everything declared there appears in the Model Metadata on the ALIDA model detail page, and is what the user sees and can request.
The developer has full freedom over which outputs to declare. For example, for a sklearn classifier one can choose to expose only the predicted class (predict), or also the per-class probabilities (predict_proba), or both:
"outputs": [
{ "name": "predict", "datatype": "INT64", "shape": [-1, 1] },
{ "name": "predict_proba", "datatype": "FP64", "shape": [-1, 3] }
]The user can then explicitly request one of the declared outputs in the inference payload:
{
"inputs": [...],
"outputs": [{ "name": "predict_proba" }]
}"Note — available outputs by runtime"
The output names that MLServer recognises depend on the runtime:
| Runtime | Available outputs | Notes |
|---|---|---|
mlserver_sklearn.SKLearnModel | predict, predict_proba, transform | predict_proba for classifiers only; transform for pipelines only; predict is the default |
mlserver_xgboost.XGBoostModel | predict | Single output |
mlserver_lightgbm.LightGBMModel | predict | Single output |
mlserver_catboost.CatBoostModel | predict | Single output |
mlserver_mlflow.MLflowRuntime | predict | Calls MLflow pyfunc predict(); does not expose predict_proba separately |
mlserver_huggingface.HuggingFaceRuntime | depends on task | Output returned in the configured HuggingFace task format |
custom (mlserver.MLModel) | defined by developer | Developer builds the InferenceResponse in the predict() method |
For regression never declare predict_proba in outputs regardless of the runtime.
Content Type
The content_type determines how MLServer decodes the V2 payload before passing it to the model. The available content types are:
| Content type | Resulting Python type | Request level | Input level |
|---|---|---|---|
np | numpy.ndarray | ✅ | ✅ |
pd | pandas.DataFrame | ✅ | ❌ |
str | str (UTF-8) | ✅ | ✅ |
base64 | byte decodificati da base64 | ❌ | ✅ |
datetime | datetime.datetime | ❌ | ✅ |
With np, the data field is flat and shape is used to reshape the tensor. Suitable for models that accept a pure numeric array with a fixed column order.
With pd, inputs are aggregated into a pandas.DataFrame, where each input name becomes a column name. Suitable for tabular sklearn pipelines that use column names to apply different transformations. pd operates only at request level and can be combined with input-level content types (e.g. np for numeric columns, str for string columns).
The content type can be declared in model-settings.json (in the parameters at request level, or in each input's parameters section). This becomes the default: requests that do not specify an explicit content type will use the one from the metadata. Content types explicitly specified in the request always take precedence over those in the metadata.
"Note — sklearn default content type"
If no content type is declared in the metadata or in the request, the sklearn runtime decodes the payload as a NumPy array. To use Pandas DataFrame, the developer must explicitly declare `"content_type": "pd"` in the `parameters` of `model-settings.json`.
Supported Data Types (V2 Inference Protocol)
| Type | Description | Corresponding pandas type |
|---|---|---|
BOOL | Boolean | bool |
INT8 | Signed 8-bit integer | int8 |
INT16 | Signed 16-bit integer | int16 |
INT32 | Signed 32-bit integer | int32 |
INT64 | Signed 64-bit integer | int64 |
UINT8 | Unsigned 8-bit integer | uint8 |
UINT16 | Unsigned 16-bit integer | uint16 |
UINT32 | Unsigned 32-bit integer | uint32 |
UINT64 | Unsigned 64-bit integer | uint64 |
FP16 | 16-bit floating point | float16 |
FP32 | 32-bit floating point | float32 |
FP64 | 64-bit floating point | float64 |
BYTES | Bytes / string / categorical / object | object, category |
STRING | UTF-8 string | string |
Generating model-settings.json
It is recommended to generate the file programmatically from the training data, deriving V2 types from the pandas dtypes of X and y. Below is a utility function:
import json
from pathlib import Path
import pandas as pd
def to_v2_dtype(pd_dtype) -> str:
"""Maps a pandas dtype to the corresponding V2 type."""
kind = getattr(pd_dtype, "kind", "O")
if kind == "b":
return "BOOL"
if kind in ("i", "u"):
return "INT64"
if kind == "f":
return "FP64" if str(pd_dtype) in ("float64", "Float64") else "FP32"
return "BYTES"
def create_model_settings(
model_dir: str,
X_sample: pd.DataFrame,
y_sample: pd.Series,
model_id: str,
implementation: str,
uri: str = ".",
is_regression: bool = False,
content_type: str = None,
extra: dict = None,
):
"""
Generates model-settings.json in the model folder.
Args:
model_dir: ALIDA output-model folder
X_sample: DataFrame with features (even a single sample)
y_sample: Series with target (even a single sample)
model_id: value of args.output_model_id injected by ALIDA
implementation: MLServer runtime (e.g. "mlserver_sklearn.SKLearnModel")
uri: relative path to the model artifact
is_regression: True for regression, False for classification
content_type: content type for payload decoding (e.g. "pd", "np")
extra: extra parameters for the runtime (e.g. {"task": "question-answering"})
"""
inputs = [
{
"name": col,
"datatype": to_v2_dtype(dtype),
"shape": [-1]
}
for col, dtype in zip(X_sample.columns, X_sample.dtypes)
]
if is_regression:
out_dtype = to_v2_dtype(y_sample.dtype)
outputs = [{"name": "predict", "datatype": out_dtype, "shape": [-1, 1]}]
else:
y_kind = getattr(y_sample.dtype, "kind", "O")
if y_kind == "b":
out_dtype = "BOOL"
elif y_kind in ("i", "u"):
out_dtype = "INT64"
else:
out_dtype = "BYTES"
outputs = [{"name": "predict", "datatype": out_dtype, "shape": [-1, 1]}]
parameters = {"uri": uri}
if content_type is not None:
parameters["content_type"] = content_type
if extra is not None:
parameters["extra"] = extra
settings = {
"name": f"model-{model_id}",
"implementation": implementation,
"parameters": parameters,
"inputs": inputs,
"outputs": outputs
}
settings_path = Path(model_dir) / "model-settings.json"
with open(settings_path, "w", encoding="utf-8") as f:
json.dump(settings, f, indent=2, ensure_ascii=False)
return settings_path"Note"
The function generates a single `predict` output. To expose additional outputs (e.g. `predict_proba`), the developer can manually add further elements to the `outputs` list after the call, or extend the function as needed.
Expected Output Folder Structure
<output-model>/
├── model-settings.json ← mandatory
└── <artefatto-modello> ← dipende dal runtime sceltoSupported Runtimes
| MLServer Runtime | format in metamodel | Artifact format |
|---|---|---|
mlserver_sklearn.SKLearnModel | sklearn | .skops (recommended), .joblib, .pkl |
mlserver_xgboost.XGBoostModel | xgboost | .json, .bst |
mlserver_lightgbm.LightGBMModel | lightgbm | .txt, .bst |
mlserver_catboost.CatBoostModel | catboost | .cbm |
mlserver_mlflow.MLflowRuntime | mlflow | MLflow folder (with MLmodel) |
mlserver_huggingface.HuggingFaceRuntime | huggingface | local model or from HuggingFace Hub |
custom class extending mlserver.MLModel | python | directory with .py file + artifact |
The format value in the metamodel is used by ALIDA for internal routing to the correct serving runtime and is independent of the implementation field in model-settings.json.
Scikit-learn
Runtime: mlserver_sklearn.SKLearnModel
Recommended artifact format: .skops file, produced with the Skops library. Unlike pickle/joblib, Skops does not allow arbitrary code execution during deserialization.
The .joblib and .pkl formats are also supported.
import skops.io as sio
os.makedirs(args.output_model, exist_ok=True)
model_path = os.path.join(args.output_model, "model.skops")
sio.dump(pipeline, model_path)
create_model_settings(
model_dir=args.output_model,
X_sample=X,
y_sample=y,
model_id=args.output_model_id,
implementation="mlserver_sklearn.SKLearnModel",
uri="model.skops",
content_type="pd",
is_regression=False
)XGBoost
Runtime: mlserver_xgboost.XGBoostModel
Artifact format: .json or .bst file — native XGBoost format (save_model()).
model_path = os.path.join(args.output_model, "model.json")
model.save_model(model_path)
create_model_settings(
model_dir=args.output_model,
X_sample=X,
y_sample=y,
model_id=args.output_model_id,
implementation="mlserver_xgboost.XGBoostModel",
uri="model.json",
is_regression=False
)LightGBM
Runtime: mlserver_lightgbm.LightGBMModel
Artifact format: .txt or .bst file — native LightGBM format (booster_.save_model()).
model_path = os.path.join(args.output_model, "model.txt")
model.booster_.save_model(model_path)
create_model_settings(
model_dir=args.output_model,
X_sample=X,
y_sample=y,
model_id=args.output_model_id,
implementation="mlserver_lightgbm.LightGBMModel",
uri="model.txt",
is_regression=False
)CatBoost
Runtime: mlserver_catboost.CatBoostModel
Artifact format: .cbm file — native CatBoost format (save_model()).
model_path = os.path.join(args.output_model, "model.cbm")
model.save_model(model_path)
create_model_settings(
model_dir=args.output_model,
X_sample=X,
y_sample=y,
model_id=args.output_model_id,
implementation="mlserver_catboost.CatBoostModel",
uri="model.cbm",
is_regression=False
)MLflow (native format)
Runtime: mlserver_mlflow.MLflowRuntime
Artifact format: complete MLflow folder containing the MLmodel file and the flavour artifacts. When the MLflow model defines a signature, MLServer automatically converts it to a V2 metadata schema. The model-settings.json must still be created to specify the model name and implementation.
import mlflow
run_id = mlflow.active_run().info.run_id
mlflow.artifacts.download_artifacts(
run_id=run_id,
artifact_path="model",
dst_path=args.output_model
)
create_model_settings(
model_dir=args.output_model,
X_sample=X,
y_sample=y,
model_id=args.output_model_id,
implementation="mlserver_mlflow.MLflowRuntime",
uri=".",
content_type="pd",
is_regression=False
)HuggingFace
Runtime: mlserver_huggingface.HuggingFaceRuntime
The HuggingFace runtime supports models from the HuggingFace Transformers library. Configuration is done through the parameters.extra field of model-settings.json:
task— (mandatory) the HuggingFace task type (e.g."question-answering","text-generation","text-classification", etc.)pretrained_model— (optional) model name in the HuggingFace Hub; takes precedence overparameters.urioptimum_model— (optional)trueto use models optimized with Optimumdevice— (optional) device to load the model on (0for GPU,-1for CPU)
To load a local model, specify the path in parameters.uri. To load a model from the HuggingFace Hub, specify the name in parameters.extra.pretrained_model.
!!! warning "Warning — resources" HuggingFace Transformer models can require significant amounts of RAM and/or GPU memory. Verify that the ALIDA instance has sufficient resources for the chosen model. When registering the Service, declare the required resources through Resource type Service Properties (see Service Registration).
When using `pretrained_model`, the model is downloaded from the HuggingFace Hub at first load — this requires network connectivity and disk space in the container. To avoid delays at first deploy, it is recommended to include the model artifacts directly in the `--output-model` folder.
{
"name": "model-<output_model_id>",
"implementation": "mlserver_huggingface.HuggingFaceRuntime",
"parameters": {
"extra": {
"task": "text-classification",
"pretrained_model": "distilbert-base-uncased-finetuned-sst-2-english"
}
},
"inputs": [
{
"name": "text_inputs",
"datatype": "BYTES",
"shape": [-1]
}
],
"outputs": [
{
"name": "output",
"datatype": "BYTES",
"shape": [-1, 1]
}
]
}Python Custom
Runtime: classe custom che estende mlserver.MLModel
For models not covered by standard runtimes, a custom runtime can be implemented. The developer must create a Python file with a class extending mlserver.MLModel that implements the load() and predict() methods:
from mlserver import MLModel
from mlserver.types import InferenceRequest, InferenceResponse
from mlserver.codecs import NumpyCodec
import numpy as np
class MyCustomModel(MLModel):
async def load(self) -> bool:
# Load the model from self.settings.parameters.uri
model_uri = self.settings.parameters.uri
# ... loading logic ...
self.ready = True
return self.ready
async def predict(self, payload: InferenceRequest) -> InferenceResponse:
# Decode input
input_data = self.decode(payload.inputs[0])
# ... inference logic ...
result = np.array([0]) # example
return InferenceResponse(
model_name=self.settings.name,
outputs=[NumpyCodec.encode_output("predict", result)]
)The model-settings.json specifies the runtime as <file_name>.<ClassName>:
{
"name": "model-<output_model_id>",
"implementation": "my_model.MyCustomModel",
"parameters": {
"uri": "."
},
"inputs": [...],
"outputs": [...]
}The .py file with the class must be in the same folder as model-settings.json (--output-model folder).
"Note"
The `load()` and `predict()` methods are **asynchronous** (`async`). All Python dependencies used by the class must be installed in the *Service* container (declared in `requirements.txt`). The `.py` file must be self-contained or import only modules present in the Docker image.
!!! warning "Warning — preferred formats for serving" Formats based on Python serialization (pickle, joblib) are supported but discouraged for models intended for serving. Prefer the native formats of each framework (Skops for sklearn, .json for XGBoost, etc.) which guarantee safer deserialization.
Metamodel Declaration
The output-model port in the Service metamodel must declare the model field with the appropriate format value. This value is used by ALIDA for internal routing to the correct serving runtime and is independent of the implementation field in model-settings.json:
{
"key": "output-model",
"type": "application",
"mandatory": true,
"invisible": true,
"valueType": "STRING",
"defaultValue": null,
"model": { "format": "sklearn" }
}Complete Example — sklearn Training with Skops and Serving
import argparse
def str2bool(v):
if isinstance(v, bool):
return v
if v.lower() in ('yes', 'true', 't', 'y', '1'):
return True
elif v.lower() in ('no', 'false', 'f', 'n', '0', ''):
return False
parser = argparse.ArgumentParser()
parser.add_argument('--input-dataset', dest='input_dataset', type=str, required=True)
parser.add_argument('--input-dataset.s3_bucket', dest='input_s3_bucket', type=str, required=True)
parser.add_argument('--input-dataset.s3_URL', dest='input_s3_url', type=str, required=True)
parser.add_argument('--input-dataset.s3_ACCESS_KEY', dest='input_access_key', type=str, required=True)
parser.add_argument('--input-dataset.s3_SECRET_KEY', dest='input_secret_key', type=str, required=True)
parser.add_argument('--input-dataset.s3_REGION', dest='input_dataset_region_name', type=str, required=False)
parser.add_argument('--input-dataset.use_ssl', dest='input_use_ssl', type=str2bool, required=True)
parser.add_argument('--input-columns', dest='input_columns', type=str, required=False)
parser.add_argument('--output-model', dest='output_model', type=str, required=True)
parser.add_argument('--output-model.s3_bucket', dest='output_s3_bucket', type=str, required=True)
parser.add_argument('--output-model.s3_URL', dest='output_s3_url', type=str, required=True)
parser.add_argument('--output-model.s3_ACCESS_KEY', dest='output_access_key', type=str, required=True)
parser.add_argument('--output-model.s3_SECRET_KEY', dest='output_secret_key', type=str, required=True)
parser.add_argument('--output-model.s3_REGION', dest='output_model_region_name', type=str, required=False)
parser.add_argument('--output-model.use_ssl', dest='output_use_ssl', type=str2bool, required=True)
parser.add_argument('--output-model.id', dest='output_model_id', type=str, required=True)
parser.add_argument('--label-column', dest='label_column', type=str, required=True)
args, unknown = parser.parse_known_args()import json
import os
from pathlib import Path
import pandas as pd
import skops.io as sio
from minio import Minio
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from arguments import args
def to_v2_dtype(pd_dtype) -> str:
kind = getattr(pd_dtype, "kind", "O")
if kind == "b":
return "BOOL"
if kind in ("i", "u"):
return "INT64"
if kind == "f":
return "FP64" if str(pd_dtype) in ("float64", "Float64") else "FP32"
return "BYTES"
def create_model_settings(model_dir, X_sample, y_sample, model_id, implementation,
uri=".", is_regression=False, content_type=None, extra=None):
inputs = [
{"name": col, "datatype": to_v2_dtype(dtype), "shape": [-1]}
for col, dtype in zip(X_sample.columns, X_sample.dtypes)
]
if is_regression:
out_dtype = to_v2_dtype(y_sample.dtype)
else:
y_kind = getattr(y_sample.dtype, "kind", "O")
out_dtype = "BOOL" if y_kind == "b" else "INT64" if y_kind in ("i", "u") else "BYTES"
outputs = [{"name": "predict", "datatype": out_dtype, "shape": [-1, 1]}]
parameters = {"uri": uri}
if content_type is not None:
parameters["content_type"] = content_type
if extra is not None:
parameters["extra"] = extra
settings = {
"name": f"model-{model_id}",
"implementation": implementation,
"parameters": parameters,
"inputs": inputs,
"outputs": outputs
}
settings_path = Path(model_dir) / "model-settings.json"
with open(settings_path, "w", encoding="utf-8") as f:
json.dump(settings, f, indent=2, ensure_ascii=False)
return settings_path
def s3_ls(address, access_key, secret_key, region, bucket_name, folder, extension, use_ssl=False):
if not folder.endswith("/"):
folder = folder + "/"
cleaned = address.replace("http://", "").replace("https://", "")
client = Minio(cleaned, access_key=access_key, secret_key=secret_key, secure=use_ssl, region=region)
objects = client.list_objects(bucket_name=bucket_name, prefix=folder)
files_list = [x._object_name for x in objects if x._object_name.endswith(extension)]
if len(files_list) == 0:
raise Exception("Empty dataset!")
return "s3://" + bucket_name + "/" + files_list[0]
# Load dataset
storage_options = {
'key': args.input_access_key,
'secret': args.input_secret_key,
'region': args.input_dataset_region_name,
'client_kwargs': {'endpoint_url': args.input_s3_url}
}
file_path = s3_ls(
args.input_s3_url, args.input_access_key, args.input_secret_key,
args.input_dataset_region_name, args.input_s3_bucket, args.input_dataset, ".csv"
)
dataset = pd.read_csv(file_path, storage_options=storage_options, sep=None, engine='python')
if args.input_columns is not None and args.input_columns.strip() != '*':
dataset = dataset[[c.strip() for c in args.input_columns.split(",")]].copy()
X = dataset.drop(columns=[args.label_column])
y = dataset[args.label_column]
# Training
pipeline = Pipeline([
("scaler", StandardScaler()),
("clf", LogisticRegression())
])
pipeline.fit(X, y)
# Save model in Skops format
os.makedirs(args.output_model, exist_ok=True)
model_path = os.path.join(args.output_model, "model.skops")
sio.dump(pipeline, model_path)
# Generate model-settings.json
create_model_settings(
model_dir=args.output_model,
X_sample=X,
y_sample=y,
model_id=args.output_model_id,
implementation="mlserver_sklearn.SKLearnModel",
uri="model.skops",
content_type="pd",
is_regression=False
)
print(f"Model saved to {args.output_model}", flush=True)The corresponding metamodel declares the output port with format: "sklearn":
{
"key": "output-model",
"type": "application",
"mandatory": true,
"invisible": true,
"valueType": "STRING",
"defaultValue": null,
"model": { "format": "sklearn" }
}