Resources by weimiao0112

Inside Airbnb — Listing Data

Creators: Murray Cox
Publication Date: Ongoing since February 2015
Creators: Murray Cox

Inside Airbnb provides listing-level data scraped from publicly available information on Airbnb.com for cities worldwide. Each snapshot contains the listing title and full-text description, host attributes (e.g., Superhost status, verifications, experience), property characteristics (room/property type, capacity, amenities, price and fees, cancellation policy), quality ratings, and review counts. The data is well suited to text/emoji analysis and to modelling how listing content relates to demand and electronic word-of-mouth.

Multimodal Benchmark Dataset

Creators: Wang, X. (Shane), Bendle, N., & Pan, Y. (2024)
Publication Date: Published by the authors in 2024
Creators: Wang, X. (Shane), Bendle, N., & Pan, Y. (2024)

To benchmark the machine learning capabilities of AWS, Google Cloud Platform, and Microsoft Azure for marketing tasks, the authors ran several publicly available unstructured datasets (image, video, and audio) through each platform. The supporting data and code are publicly available on GitHub. The data comprise three datasets: an image dataset, a video dataset, and an audio dataset.

The image dataset consists of three labelled datasets used to evaluate image classification and object detection performance: (1) a vehicle dataset containing 17,760 images labelled as vehicle or no vehicle; (2) an apparel dataset containing 11,385 images with 24 labels formed by crossing six colours (black, blue, brown, green, red, and white) with four clothing types (dress, pants, shirt, and shoes); and (3) a face authenticity dataset containing 2,041 images labelled as real or fake. The image datasets were curated from third-party datasets available on Kaggle and assembled and re-shared for this study by Wang, Bendle, and Pan.

The video dataset consists of three publicly available datasets used to evaluate video content classification on a custom model: (1) the Fight Scenes dataset, containing fighting and non-fighting actions; and two subsets of a public video advertisement dataset labelled according to persuasive strategies, including (2) exciting versus not exciting and (3) funny versus not funny (excluding private or unavailable videos). The Fight Scenes dataset is publicly available, while the advertising video subsets were derived from Hussain et al. (2017), Automatic Understanding of Image and Video Advertisements (IEEE CVPR). These datasets were assembled for this study by Wang, Bendle, and Pan.

The audio dataset consists of 3,021 podcast summaries used to evaluate speech-to-text transcription performance. Each platform generates a transcript together with confidence scores and alternative word predictions. The audio data were collected online and released by the authors through their GitHub repository.

Kickstarter Campaign Data

Creators: Kickstarter, PBC (primary source). Public re-distributions: Web Robots (monthly scrapes since 2014); a Zenodo mirror also exists.
Publication Date: Kickstarter data is generated continuously from campaigns; public Web Robots snapshots run monthly (since March 2016).
Creators: Kickstarter, PBC (primary source). Public re-distributions: Web Robots (monthly scrapes since 2014); a Zenodo mirror also exists.

Reward-based crowdfunding campaign data from Kickstarter. The core unstructured content is each campaign’s project description (pitch text), which the paper mines for rhetorical signals — emotional tone, cognitive tone, communal language style, and linguistic style match — alongside a substantive signal (backer support), and relates these to the trajectory of funds pledged over the campaign. Campaign metadata typically includes the funding goal, amount pledged, number of backers, category, launch/deadline dates, and final outcome (success/failure). The data supports research on persuasion, entrepreneurship, and consumer/backer behaviour.

Text-Based Measures of Firm Position and Differentiation in Technology Space

Creators: Sam Arts; Bruno Cassiman; Jianan Hou (built on the DISCERN database by Ashish Arora, Sharon Belenzon & Lia Sheer)
Publication Date: 2023
Creators: Sam Arts; Bruno Cassiman; Jianan Hou (built on the DISCERN database by Ashish Arora, Sharon Belenzon & Lia Sheer)

Data and code that represent each firm–year technology portfolio as a vector built from the stemmed technical keywords in the firm’s patents (title, abstract, claims). Cosine similarity between vectors gives a text-based measure of how similar two firms are technologically, and of how differentiated a firm is. Built on DISCERN, which links roughly 1.35 million U.S. patents (1980–2015) to Compustat firms.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.