VMAF is a ML (machine learning) model[^0] that computes several traditional metrics across frames, aggregates results, and computes a single score between 0 and 100. Developed by Netflix, it explicitly targets human perception of video quality degradation when viewed at a nominal distance on TV, mobiles or 4K displays.
A encoded video and a reference video is required for comparison. In all practical scenarios, the reference video is the source video that was encoded. The reference video is often the highest quality (e.g. archival grade/master) version of the video in question.
- Choosing a right model:[^5] Using the wrong model will inflate or deflate the true results. Such as a Mobile/4K model used for HD comparison inflates results. By default, HD model is chosen.
- Frame alignment and resolution sensitivity: Encode and the reference must have the same resolution, frame rate[^1], pixel format, and the frames must h