Showing 1-8 of 573 results
Creators: Thomas, Michael
Important features / special characteristics: Large replication archive studying whether promotional puffery in Airbnb listing language affects outcomes. It includes original multi-city Inside Airbnb downloads, ChatGPT-generated assignment/classification outputs, and preparation/analysis code.

Data size: 69.722 GB (69,721,605,344 bytes).

Number of observations: Not reported in repository metadata.

Temporal coverage: Not reported in repository metadata; coverage depends on the included Inside Airbnb city snapshots.

Structure: 5 restricted files: a very large compressed city-data archive, two compressed ChatGPT-output archives, a compressed code archive, and a README.

Classification note: ChatGPT-derived labels are synthetic/derived, but the underlying listing data are scraped empirical Airbnb records; the requested category is retained.

Consumer–Spokesperson Vocal Similarity: Audio, Extracted Voice Features, Experiment Data, and Models

Creators: Hyun, Kimberly; Lowe, Michael; Krishna, Aradhna
Publication Date: 2026-04-21
Creators: Hyun, Kimberly; Lowe, Michael; Krishna, Aradhna
Important features / special characteristics: Large multimodal replication collection supporting research on vocal similarity/timbre, trust, and persuasion. It includes audio files, structured files, feature-extraction/preprocessing resources, and modeling materials. The repository description distinguishes restricted Studies 1–2 from publicly shared Studies 3–6 resources hosted on OSF.

Data size: 44.081 GB (44,081,476,587 bytes).

Number of observations: The Dataverse record contains 4,230 files; the number of study participants/analytical observations is not reported in repository metadata.

Temporal coverage: Not reported in repository metadata.

Structure: 4,230 Dataverse files, predominantly WAV audio plus spreadsheets and analysis resources. The Dataverse metadata marks all files restricted; additional public code/data for Studies 3–6 are referenced at https://osf.io/gsx32/.

Classification note: The archive includes real audio/experimental material and derived features, not solely synthetic data; the requested category is retained.

Multimodal Benchmark Dataset

Creators: Wang, X. (Shane), Bendle, N., & Pan, Y. (2024)
Publication Date: Published by the authors in 2024
Creators: Wang, X. (Shane), Bendle, N., & Pan, Y. (2024)

To benchmark the machine learning capabilities of AWS, Google Cloud Platform, and Microsoft Azure for marketing tasks, the authors ran several publicly available unstructured datasets (image, video, and audio) through each platform. The supporting data and code are publicly available on GitHub. The data comprise three datasets: an image dataset, a video dataset, and an audio dataset.

The image dataset consists of three labelled datasets used to evaluate image classification and object detection performance: (1) a vehicle dataset containing 17,760 images labelled as vehicle or no vehicle; (2) an apparel dataset containing 11,385 images with 24 labels formed by crossing six colours (black, blue, brown, green, red, and white) with four clothing types (dress, pants, shirt, and shoes); and (3) a face authenticity dataset containing 2,041 images labelled as real or fake. The image datasets were curated from third-party datasets available on Kaggle and assembled and re-shared for this study by Wang, Bendle, and Pan.

The video dataset consists of three publicly available datasets used to evaluate video content classification on a custom model: (1) the Fight Scenes dataset, containing fighting and non-fighting actions; and two subsets of a public video advertisement dataset labelled according to persuasive strategies, including (2) exciting versus not exciting and (3) funny versus not funny (excluding private or unavailable videos). The Fight Scenes dataset is publicly available, while the advertising video subsets were derived from Hussain et al. (2017), Automatic Understanding of Image and Video Advertisements (IEEE CVPR). These datasets were assembled for this study by Wang, Bendle, and Pan.

The audio dataset consists of 3,021 podcast summaries used to evaluate speech-to-text transcription performance. Each platform generates a transcript together with confidence scores and alternative word predictions. The audio data were collected online and released by the authors through their GitHub repository.

Inside Airbnb — Listing Data

Creators: Murray Cox
Publication Date: Ongoing since February 2015
Creators: Murray Cox

Inside Airbnb provides listing-level data scraped from publicly available information on Airbnb.com for cities worldwide. Each snapshot contains the listing title and full-text description, host attributes (e.g., Superhost status, verifications, experience), property characteristics (room/property type, capacity, amenities, price and fees, cancellation policy), quality ratings, and review counts. The data is well suited to text/emoji analysis and to modelling how listing content relates to demand and electronic word-of-mouth.

Kickstarter Campaign Data

Creators: Kickstarter, PBC (primary source). Public re-distributions: Web Robots (monthly scrapes since 2014); a Zenodo mirror also exists.
Publication Date: Kickstarter data is generated continuously from campaigns; public Web Robots snapshots run monthly (since March 2016).
Creators: Kickstarter, PBC (primary source). Public re-distributions: Web Robots (monthly scrapes since 2014); a Zenodo mirror also exists.

Reward-based crowdfunding campaign data from Kickstarter. The core unstructured content is each campaign’s project description (pitch text), which the paper mines for rhetorical signals — emotional tone, cognitive tone, communal language style, and linguistic style match — alongside a substantive signal (backer support), and relates these to the trajectory of funds pledged over the campaign. Campaign metadata typically includes the funding goal, amount pledged, number of backers, category, launch/deadline dates, and final outcome (success/failure). The data supports research on persuasion, entrepreneurship, and consumer/backer behaviour.

Text-Based Measures of Firm Position and Differentiation in Technology Space

Creators: Sam Arts; Bruno Cassiman; Jianan Hou (built on the DISCERN database by Ashish Arora, Sharon Belenzon & Lia Sheer)
Publication Date: 2023
Creators: Sam Arts; Bruno Cassiman; Jianan Hou (built on the DISCERN database by Ashish Arora, Sharon Belenzon & Lia Sheer)

Data and code that represent each firm–year technology portfolio as a vector built from the stemmed technical keywords in the firm’s patents (title, abstract, claims). Cosine similarity between vectors gives a text-based measure of how similar two firms are technologically, and of how differentiated a firm is. Built on DISCERN, which links roughly 1.35 million U.S. patents (1980–2015) to Compustat firms.

Generative Agents: Interviews, Surveys, Behavioral Tasks, and Digital Representations of 1,052 U.S. Adults

Creators: Joon Sung Park; Carolyn Q. Zou; Jonne Kamphorst; Niles Egan; Aaron Shaw; Benjamin Mako Hill; Carrie Cai; Meredith Ringel Morris; Percy Liang; Robb Willer; Michael S. Bernstein
Publication Date: 2024-11-15
Creators: Joon Sung Park; Carolyn Q. Zou; Jonne Kamphorst; Niles Egan; Aaron Shaw; Benjamin Mako Hill; Carrie Cai; Meredith Ringel Morris; Percy Liang; Robb Willer; Michael S. Bernstein
Important features / special characteristics: Digital representations of a diverse national sample of 1,052 U.S. adults, grounded in sensitive real-person self-reports. Each participant completed an approximately two-hour AI-conducted semi-structured interview (about 6,500 words per person), structured surveys, and behavioral tasks. Agents were evaluated using General Social Survey items, the 44-item Big Five inventory, five behavioral economic games, and five social-science experiments; participants repeated tasks two weeks later for a test–retest benchmark.

Data size: Not reported in GB.

Number of observations: 1,052 participants / agent representations. The exact number of row-level records is not reported.

Temporal coverage: Study collection dates are not reported in the cited public materials; evaluation includes a two-week repeat-measurement interval.

Structure: Three principal input/configuration groupings are described: (1) two-hour interview transcripts, (2) structured survey responses, and (3) combined interview-plus-survey inputs. Evaluation data cover GSS responses, Big Five personality items, economic games, and social-science experiments.

Important classification note: The digital agents and their generated predictions are synthetic, but the grounding interviews, surveys, and task responses are real individual-level self-report data.

SYNTHIA — Synthetic Urban Driving Images with Segmentation, Depth, and Instance Labels

Creators: German Ros; Laura Sellart; Joanna Materzynska; David Vazquez; Antonio M. Lopez
Publication Date: 2016
Creators: German Ros; Laura Sellart; Joanna Materzynska; David Vazquez; Antonio M. Lopez
Synthetic urban driving scenes with RGB images, semantic and instance segmentation, depth, camera information, and varied seasons, weather, and viewpoints. Organized into multiple rendered sequences and image subsets.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.