Use the VMAF Python library to run VMAF from the command line on .yuv files,
train and validate a new VMAF model on a video dataset, and experiment with
other video quality metrics. It also holds the analysis tools (BD-rate, LIME)
and the format converters.
This page covers setup first, then the command-line tools, the analysis and format tools, and the core classes for extending the library.
Make sure you have python3 (Python 3.10 or higher for the harness package; the
repository's own tooling asks for 3.14). Check the version with
python3 --version.
On Debian-based systems such as Debian 13 or Ubuntu 26.04 (the project's CI base images):
sudo apt install nasm doxygen python3-devOn Fedora 22 or higher, or CentOS 8 or higher:
sudo dnf install nasm doxygen python3-develOn older CentOS and RHEL:
sudo yum install nasm doxygen python3-develMake sure nasm is 2.13.02 or higher (check by nasm --version).
First, install Homebrew.
If you don't already have a python3 installation on your Mac, run the
following to install Python 3 via Homebrew:
brew install python3Install the remaining dependencies:
brew install nasm doxygen llvm libompNote that brew requires no sudo.
Follow these steps to set up a clean virtual environment and build the binary.
-
Create and activate the virtual environment:
python3 -m pip install virtualenv python3 -m virtualenv .venv source .venv/bin/activateFrom here on,
python3andpip3are relative to the virtualenv and isolated from the system Python. In later shell sessions, re-activate it withsource .venv/bin/activate. -
Install the tools required to build VMAF into the virtualenv:
pip3 install cython numpy meson ninja
Make sure
ninjais 1.7.1 or higher (ninja --version). -
Clean-build the binary (the default
maketarget builds intocore/build):make clean; make -
Check that the build succeeded:
./core/build/tools/vmaf --version
-
Install the remaining Python packages:
pip3 install -r python/requirements.txt
!!! note "macOS needs the Homebrew LLVM" The macOS clang has no OpenMP support, which libsvm-official needs. Install the requirements with the Homebrew LLVM:
```bash
CC=$HOMEBREW_PREFIX/opt/llvm/bin/clang CXX=$HOMEBREW_PREFIX/opt/llvm/bin/clang++ pip3 install -r python/requirements.txt
```
Run unittests and make sure they all pass:
./scripts/run_unittests.shrun_vmaf runs VMAF from the command line on .yuv inputs. The package ships a
__main__ module, so the shortest invocation is:
python -m vmaf fmt width height ref_path dis_path [options]The same command through the full module path:
python -m vmaf.script.run_vmaf \
format width height \
reference_path \
distorted_path \
[--out-fmt output_format]formatcan be one of:yuv420p,yuv422p,yuv444p(8-Bit YUV)yuv420p10le,yuv422p10le,yuv444p10le(10-Bit little-endian YUV)yuv420p12le,yuv422p12le,yuv444p12le(12-Bit little-endian YUV)yuv420p16le,yuv422p16le,yuv444p16le(16-Bit little-endian YUV)
widthandheightare the width and height of the videos, in pixelsreference_pathanddistorted_pathare the paths to the reference and distorted video filesoutput_formatcan be one of:text(default)xmljson
| Option | Effect |
|---|---|
--model model_path |
Use a model file instead of the default. |
--out-fmt output_format |
Output format: text (default), xml or json. |
--pool method |
Pooling of the per-frame scores: mean, harmonic_mean, min, median, perc5, perc10 or perc20. |
--phone-model |
Enable the score transform (enable_transform_score) of the phone-viewing model. |
--ci |
Use the bootstrap runner to report a confidence interval. Cannot be combined with --local-explain. |
--local-explain |
Add a LIME explanation of the score. |
--save-plot plot_dir |
With --local-explain, write the explanation plots to plot_dir instead of showing them. |
The following command runs VMAF on a pair of .yuv inputs
(src01_hrc00_576x324.yuv,
src01_hrc01_576x324.yuv):
python -m vmaf.script.run_vmaf \
yuv420p 576 324 \
src01_hrc00_576x324.yuv \
src01_hrc01_576x324.yuv \
--out-fmt jsonIt prints JSON with one entry per frame and an aggregate block like this (the
aggregate of a CPU run on this pair):
{
...
"aggregate": {
"VMAF_integer_feature_adm2_score": 0.9345057916666667,
"VMAF_integer_feature_motion2_score": 3.8943608125,
"VMAF_integer_feature_vif_scale0_score": 0.36366208333333333,
"VMAF_integer_feature_vif_scale1_score": 0.7674953125,
"VMAF_integer_feature_vif_scale2_score": 0.8631078125,
"VMAF_integer_feature_vif_scale3_score": 0.9157200416666665,
"VMAF_score": 76.66783253176855,
"method": "mean"
}
}where VMAF_score is the final score and the others are the scores for VMAF's
elementary metrics:
adm2,vif_scalexscores range from 0 (worst) to 1 (best)motion2score typically ranges from 0 (static) to 20 (high-motion)
VMAF first extracts quality-relevant features (elementary metrics) from a distorted video and its reference. It then fuses them into a final score with a non-linear regressor (for example an SVM), hence the name "Video Multi-method Assessment Fusion".
The package also lets you train your own perceptual quality model. The directory
model contains a number of pre-trained models, which the
commands above can load:
python -m vmaf.script.run_vmaf \
format width height \
reference_path \
distorted_path \
[--model model_path]For example:
python -m vmaf.script.run_vmaf \
yuv420p 576 324 \
python/test/resource/yuv/src01_hrc00_576x324.yuv \
python/test/resource/yuv/src01_hrc01_576x324.yuv \
--model model/other_models/nflxtrain_vmafv3.pklA user can customize the model based on:
- The video dataset it is trained on
- The list of features used
- The regressor used (and its hyper-parameters)
Once a model is trained, the VMAF package also provides tools to cross validate it on a different dataset and visualization.
To begin with, create a dataset file following the format in
example_dataset.py.
A dataset is a collection of distorted videos. Each has a unique asset ID and a
corresponding reference video, identified by a unique content ID. Each distorted
video is also associated with subjective quality score, typically a MOS (mean
opinion score), obtained through subjective study. An example code snippet that
defines a dataset is as follows:
dataset_name = 'example'
yuv_fmt = 'yuv420p'
width = 1920
height = 1080
ref_videos = [
{'content_id':0, 'path':'checkerboard.yuv'},
{'content_id':1, 'path':'flat.yuv'},
]
dis_videos = [
{'content_id':0, 'asset_id': 0, 'dmos':100, 'path':'checkerboard.yuv'},
{'content_id':0, 'asset_id': 1, 'dmos':50, 'path':'checkerboard_dis.yuv'},
{'content_id':1, 'asset_id': 2, 'dmos':100, 'path':'flat.yuv'},
{'content_id':1, 'asset_id': 3, 'dmos':80, 'path':'flat_dis.yuv'},
]See the directory
compat/python-vmaf/resource/dataset
for more examples. Also refer to the Datasets document
regarding publicly available datasets.
Dataset-level settings (width, height, yuv_fmt, quality_width,
quality_height, resampling_type, crop_cmd, pad_cmd, fps_cmd,
workfile_yuv_type, ...) apply to every video. Each video can also carry its
own:
widthandheighton a reference or a distorted video set that side's size, and both must be given together. When the two sides end up with different sizes, the dataset needsquality_widthandquality_heightto scale them to; reading it fails otherwise.resampling_typeon a reference or a distorted video sets how that side is scaled. A side without one uses the dataset's, else the distorted video's.crop_cmd,pad_cmdandfps_cmdon a video apply to that side only; a dataset-level command applies to both.workfile_yuv_typeon a distorted video sets the work-file format of that asset when the dataset sets none.
vmaf.routine.read_dataset(dataset, **kwargs) turns a dataset into Asset
objects; it is SubjectiveDatasetReader(dataset, **kwargs).read(), where
kwargs can select content_ids or asset_ids, name the groundtruth_key,
or skip assets without groundtruth (skip_asset_with_none_groundtruth=True).
SubjectiveDatasetTester(reader, quality_runner_class, ...) runs a quality
runner on the assets of a reader and, after run(), holds test_assets,
results and the correlation stats (SRCC, PCC, RMSE, ...);
run_test_on_dataset(), which run_testing below calls, is built on it.
Once a dataset is created, first validate the dataset using existing VMAF or other (PSNR, SSIM or MS-SSIM) metrics. Run:
python -m vmaf.script.run_testing \
quality_type \
test_dataset_file \
[--vmaf-model optional_VMAF_model_path] \
[--cache-result] \
[--parallelize]where quality_type can be VMAF, PSNR, SSIM, MS_SSIM, etc.
For example:
python -m vmaf.script.run_testing \
VMAF \
compat/python-vmaf/resource/example/example_dataset.py \
--cache-result \
--parallelizeEnabling --cache-result allows storing/retrieving extracted features (or
elementary quality metrics) in a data store (under
compat/python-vmaf/workspace/result_store_dir/file_result_store by default;
overridable via the VMAF_WORKSPACE environment variable — see
docs/architecture/workspace.md), since feature
extraction is the most expensive operations here.
Enabling --parallelize allows execution on multiple reference-distorted video
pairs in parallel. Sometimes it is desirable to disable parallelization for
debugging purpose (e.g. some error messages can only be displayed when parallel
execution is disabled).
Make sure matplotlib is installed to visualize the MOS-prediction scatter plot
and inspect the statistics:
- PCC – Pearson correlation coefficient
- SRCC – Spearman rank order correlation coefficient
- RMSE – root mean squared error
When creating a dataset file, one may make errors (for example, having a typo in
a file path) that could go unnoticed but make the execution of run_testing
fail. For debugging purposes, it is recommended to disable --parallelize.
If the problem persists, one may need to run the script:
python -m vmaf.script.run_cleaning_cache \
quality_type \
test_dataset_fileto clean up corrupted results in the store before retrying. For example:
python -m vmaf.script.run_cleaning_cache \
VMAF \
compat/python-vmaf/resource/example/example_dataset.pyNow that we are confident that the dataset is created correctly and we have some benchmark result on existing metrics, we proceed to train a new quality assessment model. Run:
python -m vmaf.script.run_vmaf_training \
train_dataset_filepath \
feature_param_file \
model_param_file \
output_model_file \
[--cache-result] \
[--parallelize]For example:
python -m vmaf.script.run_vmaf_training \
compat/python-vmaf/resource/example/example_dataset.py \
compat/python-vmaf/resource/feature_param/vmaf_feature_v2.py \
compat/python-vmaf/resource/model_param/libsvmnusvr_v2.py \
compat/python-vmaf/workspace/model/test_model.pkl \
--cache-result \
--parallelizefeature_param_file defines the set of features used. For example, both
dictionaries below:
feature_dict = {'VMAF_feature': 'all', }and
feature_dict = {'VMAF_feature': ['vif', 'adm'], }are valid specifications of selected features. Here VMAF_feature is an
"aggregate" feature type, and vif, adm are the "atomic" feature types within
the aggregate type. In the first case, all specifies that all atomic features
of VMAF_feature are selected. A feature_dict dictionary can also contain
more than one aggregate feature types.
model_param_file defines the type and hyper-parameters of the regressor to be
used. For details, refer to the self-explanatory examples in directory
resource/model_param. One example is:
model_type = "LIBSVMNUSVR"
model_param_dict = {
# ==== preprocess: normalize each feature ==== #
'norm_type':'clip_0to1', # rescale to within [0, 1]
# ==== postprocess: clip final quality score ==== #
'score_clip':[0.0, 100.0], # clip to within [0, 100]
# ==== libsvmnusvr parameters ==== #
'gamma':0.85, # selected
'C':1.0, # default
'nu':0.5, # default
'cache_size':200, # default
}The trained model is output to output_model_file. Once it is obtained, it can
be used by the run_vmaf, or by run_testing to validate another dataset.
Above are two example scatter plots obtained from running the
run_vmaf_training and run_testing commands on a training and a testing
dataset, respectively.
The commands run_vmaf_training and run_testing also support custom
subjective models (e.g. MLE_CO_AP2 (default), MOS, DMOS, SR_MOS (i.e. ITU-R
BT.500), BR_SR_MOS (i.e. ITU-T P.913) and more), through the
sureal package.
The subjective model option can be specified with option
--subj-model subjective_model, for example:
python -m vmaf.script.run_vmaf_training \
compat/python-vmaf/resource/example/example_raw_dataset.py \
compat/python-vmaf/resource/feature_param/vmaf_feature_v2.py \
compat/python-vmaf/resource/model_param/libsvmnusvr_v2.py \
compat/python-vmaf/workspace/model/test_model.pkl \
--subj-model MLE_CO_AP2 \
--cache-result \
--parallelize
python -m vmaf.script.run_testing \
VMAF \
compat/python-vmaf/resource/example/example_raw_dataset.py \
--subj-model MLE_CO_AP2 \
--cache-result \
--parallelizeNote that for the --subj-model option to have effect, the input dataset file
must follow a format similar to
example_raw_dataset.py.
Specifically, for each dictionary element in dis_videos, instead of having a
key named dmos or groundtruth as in
example_dataset.py,
it must have a key named os (stands for opinion score), and the value must be
a list of numbers. This is the "raw opinion score" collected from subjective
experiments, which is used as the input to the custom subjective models.
run_vmaf_cross_validation.py provides tools for cross-validation of hyper-parameters and models. run_vmaf_cv runs training on a training dataset using hyper-parameters specified in a parameter file, output a trained model file, and then test the trained model on another test dataset and report testing correlation scores. run_vmaf_kfold_cv takes in a dataset file, a parameter file, and a data structure (list of lists) that specifies the folds based on video content's IDs, and run k-fold cross validation on the video dataset. This can be useful for manually tuning the model parameters.
You can also customize VMAF by plugging in third-party features or inventing new
features, and specify them in a feature_param_file. Essentially, the
"aggregate" feature type (for example: VMAF_feature) specified in the
feature_dict corresponds to the TYPE field of a FeatureExtractor subclass
(for example: VmafFeatureExtractor). All you need to do is to create a new
class extending the FeatureExtractor base class.
Similarly, you can plug in a third-party regressor or invent a new regressor and
specify them in a model_param_file. The model_type (for example:
LIBSVMNUSVR) corresponds to the TYPE field of a TrainTestModel subclass
(for example: LibsvmnusvrTrainTestModel). All needed is to create a new class
extending the TrainTestModel base class.
For instructions on how to extending the FeatureExtractor and TrainTestModel
base classes, refer to CONTRIBUTING.md.
Overtime, a number of helper tools have been incorporated into the package, to facilitate training and validating VMAF models. An overview of the tools available can be found in this slide deck.
A Bjøntegaard-Delta (BD) rate
implementation is added.
An example usage is available in
bd_rate_calculator_test.py.
The implementation is validated against MPEG JCTVC-E137; see
bd-rate.md for usage.
An implementation of LIME is also added as part of the repository. For more information, refer to our analysis tools presentation. The main idea is to perform a local linear approximation to any regressor or classifier and then use the coefficients of the linearized model as indicators of feature importance. LIME can be used as part of the VMAF regression framework, for example:
python -m vmaf.script.run_vmaf \
yuv420p 576 324 \
src01_hrc00_576x324.yuv \
src01_hrc00_576x324.yuv \
--local-explainLIME can also be applied to any other regression scheme as long as a pre-trained model exists. For example, with the VMAF v1 SVM model:
python -m vmaf.script.run_vmaf yuv420p 576 324 \
src01_hrc00_576x324.yuv \
src01_hrc00_576x324.yuv \
--local-explain \
--model model/other_models/nflxall_vmafv1.pklA tool to convert a model file (currently support libsvm model) from pickle to
json is added at compat/python-vmaf/script/convert_model_from_pkl_to_json.py.
Usage:
usage: convert_model_from_pkl_to_json.py [-h] --input-pkl-filepath
INPUT_PKL_FILEPATH
--output-json-filepath
OUTPUT_JSON_FILEPATH
optional arguments:
-h, --help show this help message and exit
--input-pkl-filepath INPUT_PKL_FILEPATH
path to the input pkl file, example:
model/vmaf_float_v0.6.1.pkl or
model/vmaf_float_b_v0.6.3/vmaf_float_b_v0.6.3.pkl
--output-json-filepath OUTPUT_JSON_FILEPATH
path to the output json file, example:
model/vmaf_float_v0.6.1.json or model/vmaf_float_b_v0.6.3.json
Examples:
compat/python-vmaf/script/convert_model_from_pkl_to_json.py \
--input-pkl-filepath model/vmaf_float_b_v0.6.3/vmaf_float_b_v0.6.3.pkl \
--output-json-filepath ./vmaf_float_b_v0.6.3.json
compat/python-vmaf/script/convert_model_from_pkl_to_json.py \
--input-pkl-filepath model/vmaf_float_v0.6.1.pkl \
--output-json-filepath ./vmaf_float_v0.6.1.jsonThe table summarises the core classes; the diagram and per-class sections follow.
| Class | Role | Extends or uses |
|---|---|---|
Asset |
One reference/distorted pair with its preprocessing settings. | WorkdirEnabled |
Executor |
Runs a computation on a list of Assets and returns Results. |
TypeVersionEnabled |
Result |
Read-only key-value store of an Executor's scores for one Asset. |
|
ResultStore |
Saves and loads Results (FileSystemResultStore). |
|
FeatureExtractor |
Extracts features (elementary metrics). | Executor |
FeatureAssembler |
Assembles feature vectors from several extractors. | FeatureExtractors |
TrainTestModel |
Base class of every regressor. | TypeVersionEnabled |
CrossValidation |
Static helpers to validate and tune a TrainTestModel. |
|
QualityRunner |
Computes the final quality score. | Executor |
The core classes can be depicted in the diagram below:
An Asset is the most basic unit with enough information to perform a task on a
media. It includes basic information about a distorted video and its undistorted
reference counterpart, as well as the auxiliary preprocessing information that
can be understood by the Executor and its subclasses. For example:
- The frame range on which to perform a task (i.e.
dis_start_end_frameandref_start_end_frame) - At what resolution to perform a task (e.g. a video frame is upscaled with a
resampling_typemethod to the resolution specified byquality_width_heightbefore feature extraction) - Optional FFmpeg preprocessing filters in
asset_dict. Shared keys such ascrop_cmd,pad_cmd,gblur_cmd,eq_cmd,lutyuv_cmd,yadif_cmd,format_cmd,fps_cmdandselect_cmdapply to both reference and distorted inputs; target-specific keys such asref_fps_cmdordis_fps_cmdoverride the shared value for one side. The filters run in that order (Asset.ORDERED_FILTER_LIST, Netflix's order), so aformat_cmdacts on the output ofgblur_cmd, andfps_cmdandselect_cmdcome last. Each value is the argument of the FFmpeg filter of the same name (select_cmdbecomesselect=<value>).
Asset extends the WorkdirEnabled mixin, which comes with a thread-safe working
directory to facilitate parallel execution.
An Executor takes a list of Assets, runs computations on them and returns
the corresponding Results.
An Executor extends the TypeVersionEnabled mixin and must specify a unique
type and version combination (the TYPE and VERSION attributes), so that its
Results can be identified uniquely. This enables shared housekeeping: storing
and reusing Results (result_store), creating FIFO pipes (fifo_mode) and so
on.
An Executor understands the preprocessing steps in its input Assets and
relies on FFmpeg to apply them (FFmpeg 5.1 or newer must be pre-installed and
its path set as FFMPEG_PATH in an optional compat/python-vmaf/externals.py
file, which config.py reads). The decode step passes -fps_mode passthrough,
which FFmpeg 5.0 and older reject; FFmpeg 9 removed the -vsync 0 spelling used
before.
An Executor and its subclasses can take optional parameters during
initialization. There are two fields to put the optional parameters:
optional_dict: a dictionary field to specify parameters that will impact numerical result (e.g. which wavelet transform to use).VmafexecQualityRunneralso readscolor_refandcolor_distfrom it, each a dict with exactly the keysrange,primaries,trcandmatrix(for example{'range': 'limited', 'primaries': 'bt2020', 'trc': 'smpte2084', 'matrix': 'bt2020nc'}), and passes them tovmafexecas the--color_*_ref/--color_*_distflags of the CLI; an incomplete dict is anAssertionError.optional_dict2: a dictionary field to specify parameters that will NOT impact numerical result (e.g. outputting optional results).
Executor is the base class for FeatureExtractor and QualityRunner (and the
sub-subclass VmafQualityRunner).
A Result is a key-value store of read-only execution results generated by an
Executor on an Asset. A key corresponds to an "atom" feature type or a type
of a quality score, and a value is a list of score values, each corresponding to
a computation unit (i.e. in the current implementation, a frame).
When the vmaf executable reports an atom feature under a name with option
suffixes (for example integer_vif_scale0_egl_1), the key carries the suffix
(VMAF_integer_feature_vif_scale0_egl_1_score). If it reports the same atom
feature under several suffixes, because the feature ran with several option
sets, the Result has one key per suffix. A suffixed name belongs to the
longest atom feature that is a prefix of it, so integer_vif_scale0_egl_1 is a
vif_scale0 key and never a vif key. VmafexecQualityRunner reports its
*_feature_* keys the same way.
The Result class also provides a number of tools for aggregating the per-unit
scores into an average score. The default aggregatijon method is the mean, but
Result.set_score_aggregate_method() allows customizing other methods (see
test_to_score_str() in test/result_test.py for examples).
ResultStore provides capability to save and load a Result. Current
implementation FileSystemResultStore persists results by a simple file system
that save/load result in a directory. The directory has multiple subdirectories,
each corresponding to an Executor. Each subdirectory contains multiple files,
each file storing the dataframe for an Asset.
FeatureExtractor subclasses Executor, and is specifically for extracting
features (aka elementary quality metrics) from Assets. Any concrete feature
extraction implementation should extend the FeatureExtractor base class (e.g.
VmafFeatureExtractor). The TYPE field corresponds to the "aggregate" feature
name, and the ATOM_FEATURES/DERIVED_ATOM_FEATURES field corresponds to the
"atom" feature names.
FeatureAssembler assembles features for an input list of Assets on a input
list of FeatureExtractor subclasses. The constructor argument feature_dict
specifies the list of FeatureExtractor subclasses (i.e. the "aggregate"
feature) and selected "atom" features. For each asset on a FeatureExtractor,
it outputs a BasicResult object. FeatureAssembler is used by a
QualityRunner to assemble the vector of features to be used by a
TrainTestModel.
TrainTestModel is the base class for any concrete implementation of regressor,
which must provide a train() method to perform training on a set of data and
their groud-truth labels, and a predict() method to predict the labels on a
set of data, and a to_file() and a from_file() method to save and load
trained models.
A TrainTestModel constructor must supply a dictionary of parameters (i.e.
param_dict) that contains the regressor's hyper-parameters. The base class
also provides shared functionalities such as input data normalization/output
data denormalization, evaluating prediction performance, etc.
A model can carry a chroma_correction_parameter (a float between 0 and 200,
set through the property of that name). Before predicting, every
speed_chroma score that is exactly 0 is then replaced by
p * (1 - adm3) with p the parameter, from the same frame's adm3 score:
the Python form of the chroma_from_luma correction that the C predictor
applies (TrainTestModel.postprocess_feature_from_another()). A model without
the parameter predicts as before.
Like an Executor, a TrainTestModel extends TypeVersionEnabled and must
specify a unique type and version combination (by the TYPE and VERSION
attribute).
CrossValidation provides a collection of static methods to facilitate
validation of a TrainTestModel object. As such, it also provides means to
search the optimal hyper-parameter set for a TrainTestModel object.
QualityRunner subclasses Executor, and is specifically for evaluating the
quality score for Assets. Any concrete implementation to generate the final
quality score should extend the QualityRunner base class (e.g.
VmafQualityRunner, PsnrQualityRunner).
There are two ways to extend a QualityRunner base class -- either by directly
implementing the quality calculation (e.g. by calling a C executable, as in
PsnrQualityRunner), or by calling a FeatureAssembler (with indirectly calls
a FeatureExtractor) and a TrainTestModel subclass (as in
VmafQualityRunner).


