Multimodal Benchmark Dataset
Creators:
Wang, X. (Shane), Bendle, N., & Pan, Y. (2024)
Publication Date:
Published by the authors in 2024
Data Category:
Dataset Description:
To benchmark the machine learning capabilities of AWS, Google Cloud Platform, and Microsoft Azure for marketing tasks, the authors ran several publicly available unstructured datasets (image, video, and audio) through each platform. The supporting data and code are publicly available on GitHub. The data comprise three datasets: an image dataset, a video dataset, and an audio dataset.
The image dataset consists of three labelled datasets used to evaluate image classification and object detection performance: (1) a vehicle dataset containing 17,760 images labelled as vehicle or no vehicle; (2) an apparel dataset containing 11,385 images with 24 labels formed by crossing six colours (black, blue, brown, green, red, and white) with four clothing types (dress, pants, shirt, and shoes); and (3) a face authenticity dataset containing 2,041 images labelled as real or fake. The image datasets were curated from third-party datasets available on Kaggle and assembled and re-shared for this study by Wang, Bendle, and Pan.
The video dataset consists of three publicly available datasets used to evaluate video content classification on a custom model: (1) the Fight Scenes dataset, containing fighting and non-fighting actions; and two subsets of a public video advertisement dataset labelled according to persuasive strategies, including (2) exciting versus not exciting and (3) funny versus not funny (excluding private or unavailable videos). The Fight Scenes dataset is publicly available, while the advertising video subsets were derived from Hussain et al. (2017), Automatic Understanding of Image and Video Advertisements (IEEE CVPR). These datasets were assembled for this study by Wang, Bendle, and Pan.
The audio dataset consists of 3,021 podcast summaries used to evaluate speech-to-text transcription performance. Each platform generates a transcript together with confidence scores and alternative word predictions. The audio data were collected online and released by the authors through their GitHub repository.
Variables:
Details:

