VIDEO TEMPORAL COVERAGE IN DEEPFAKE DETECTION: ACCURACY AND COMPUTATIONAL COST

Authors

DOI:

https://doi.org/10.31673/2412-4338.2026.035602

Abstract

Video deepfake detectors commonly analyse one or several short clips instead of the entire recording. When forgery evidence is unevenly distributed over time, one clip may not represent the video as a whole, while processing more clips increases computational cost almost proportionally. This study determined how many uniformly placed clips are needed for reliable evaluation of short videos and where additional temporal coverage ceases to provide a practically meaningful gain. The detector was fixed throughout the experiment and combined spatial and frequency-domain frame features with Transformer-based sequence analysis. Predictions from three independently trained models were averaged and evaluated on a separate external sample of 200 matched real-fake video pairs. Nine uniformly spaced 16-frame clips were processed once for every video. The one-, three-, and five-clip modes were reconstructed as nested subsets of the same nine-position grid, so observed differences were not caused by resampling or another model run. ROC AUC was 0.8349 for one central clip, 0.8688 for three clips, 0.8724 for five clips, and 0.8732 for nine clips. Nine clips increased AUC by 0.0383 over one clip, with p = 0.001. Three clips retained 88.5% of this gain at one third of the maximum computational budget; five clips retained 97.7% at five ninths of that budget. No confirmed improvement was observed between five and nine clips. Prediction variability across clips was approximately 1.9 times higher for fake videos than for real videos. The median, maximum, and mean of the two largest clip scores did not outperform the arithmetic mean. A single central clip therefore under-represents short videos under cross-dataset evaluation. Three clips are suitable under a strict resource limit, whereas five clips are a conservative choice near the observed plateau. The finding should not be transferred directly to multi-minute recordings, which require a separate evaluation of duration-aware and adaptive temporal sampling.

Keywords: digital forensics, temporal coverage, ROC AUC, cross-dataset evaluation, temporal heterogeneity, prediction aggregation, inference efficiency

Published

2026-10-01

Issue

Section

Articles