Showing 505-512 of 573 results

Twitch Livestreaming Interactions

Creators: Rappaz, Jérémie; McAuley, Julian; Aberer, Karl
Publication Date: 2021
Creators: Rappaz, Jérémie; McAuley, Julian; Aberer, Karl

This is a dataset of users consuming streaming content on Twitch. We retrieved all streamers, and all users connected in their respective chats, every 10 minutes during 43 days. The dataset is unique because it captures real-time interactions between users and streamers at a high temporal resolution, allowing for detailed analysis of how audiences engage with live content. Below are the key features that make this dataset particularly valuable. The dataset is 6,47 GB in size and covers a 43-day period in July 2019, with data collected every 10 minutes, resulting in 6,148 time steps.

Overall, it includes:

  • Users: 100k
  • Streamers (items): 162.6k
  • Interactions: 3M
  • Time steps: 6148

Structurally, the dataset encompasses the following information:

  • User ID: Anonymized identifier for each user.

  • Stream ID: Identifier for each streaming session.

  • Streamer Username: Name of the channel or streamer.

  • Time Start: The initial time step (in 10-minute intervals) when the user was observed in the chat.

  • Time Stop: The final time step (in 10-minute intervals) when the user was observed in the chat.

Social Recommendation Data

Creators: Cai, Chenwei; He, Ruining; McAuley, Julian; Zhao, Tong; King, Irwin
Publication Date: 2017
Creators: Cai, Chenwei; He, Ruining; McAuley, Julian; Zhao, Tong; King, Irwin

These datasets include ratings as well as social (or trust) relationships between users. Data are from LibraryThing (a book review website) and epinions (general consumer reviews). Those specific user ratings allow for detailed analysis of user preferences. By capturing the social (or trust) relationships between users, this dataset enables the study of how social connections influence user behavior and recommendations. The dataset is approximately 660 MB in size and includes:

Number of Observations:

  • LibraryThing:

    • Users: 73,882
    • Items: 337,561
    • Ratings: 979,053
    • Social Relations: 120,536
  • Epinions:

    • Users: 116,260
    • Items: 41,269
    • Ratings/Feedback: 181,394
    • Social Relations: 181,304

The dataset is structured into:

  • User Information: Anonymized user identifiers.

  • Item Information: Identifiers for items such as books or products.

  • Ratings/Feedback: User-provided ratings or feedback scores for items.

  • Social Relations: Mappings of social or trust relationships between users.


Steam Video Game and Bundle Data

Creators: Kang, Wang-Cheng; McAuley, Julian; Pathak, Apurva; Gupta, Kshitiz
Publication Date: 2018
Creators: Kang, Wang-Cheng; McAuley, Julian; Pathak, Apurva; Gupta, Kshitiz
This datasets collect user interactions and metadata from the Steam platform, aimed at facilitating research in recommendation systems and user behavior analysis. They encompass a vast number of user reviews, capturing detailed feedback and engagement levels. Also, it provides insights into game bundles, detailing which games are frequently purchased together, aiding in the analysis of bundling strategies and their effectiveness. The The dataset encompasses user reviews and interactions from October 2010 to January 2018 and is approximately 1.4 GB in size and includes:
  • Reviews: 7,793,069
  • Users: 2,567,538
  • Items: 15,474
  • Bundles: 615

Structurally, the dataset comprises several key components:

  • User Reviews: Each review entry includes the user ID, game ID, review text, and associated metadata such as timestamps and ratings.

  • Game Metadata: Information about each game, including game ID, title, genre, developer, and pricing details.

  • Bundle Details: Descriptions of game bundles, specifying the bundle ID, included games, and pricing information.

 

EndoMondo Fitness Tracking Data

Creators: Ni, Jianmo; Muhlstein, Larry; McAuley, Julian
Publication Date: 2019
Creators: Ni, Jianmo; Muhlstein, Larry; McAuley, Julian

This is a collection of workout logs from users of EndoMondo. It contains sequential sensor data such as GPS coordinates (latitude, longitude, altitude), heart rate measurements, speed, and distance, making it valuable for studying workout patterns, performance tracking, and personalized fitness recommendations. Additionally, it includes user metadata such as anonymized user IDs, gender, and sport type, along with contextual factors like weather conditions. The dataset has a size of approximately 2.9 GB and consists of 1,104 users with 253,020 recorded workouts.

The dataset covers multiple components:

  • User Information: Anonymized user identifiers and gender.

  • Workout Details: Each workout log includes sport type, sequential data for GPS coordinates (latitude, longitude, altitude) with timestamps, heart rate measurements, and derived metrics such as speed and distance.

Pinterest Fashion Compatibility

Creators: Kang, Wang-Cheng; Kim, Eric; Leskovec, Jure; Rosenberg, Charles; McAuley, Julian
Publication Date: 2019
Creators: Kang, Wang-Cheng; Kim, Eric; Leskovec, Jure; Rosenberg, Charles; McAuley, Julian

This dataset is a structured collection of images and metadata designed to study the compatibility of fashion products within real-world scenes. It enables detailed analysis of how fashion items appear in different settings and supports applications in machine learning, recommendation systems, and virtual styling tools. One of its key features is the scene-product pairing, where fashion items in real-world images are annotated with bounding boxes and linked to corresponding product images. In total, the dataset includes 47,739 scene images, 38,111 product images, and 93,274 scene-product pairs, making it a comprehensive resource for fashion compatibility research.

The dataset is about 29 MB large and includes:

  • Scenes: 47,739
  • Products: 38,111
  • Scene-Product Pairs: 93,274

Behance Community Art Data

Creators: He, Ruining; Fang, Chen; Wang, Zhaowen; McAuley, Julian
Publication Date: 2016
Creators: He, Ruining; Fang, Chen; Wang, Zhaowen; McAuley, Julian

Being a small, anonymized, version of a larger proprietary dataset, this dataset covers likes and image data from the community art website Behance. It provides valuable insights into user engagement with digital art, making it a significant resource for research in recommender systems, social network analysis, and the study of artistic preferences. Also, the dataset captures user interactions in the form of “appreciations” (akin to likes) on various art items. Each appreciation reflects a user’s positive acknowledgment of an artwork, offering a measurable indicator of engagement. Additionally, the dataset includes image features extracted from the artworks, facilitating analyses that combine user behavior with visual content characteristics.

In total, the dataset is about 3.5 GB large and encompasses:

  • Users: 63,497
  • Items: 178,788
  • Appreciates (“likes”): 1,000,000

The dataset is structured to include:

  • User Data: Anonymized identifiers representing individual users.

  • Item Data: Identifiers for each artwork, accompanied by associated image features.

  • Appreciation Data: Records of user-item interactions, indicating which user appreciated which artwork.

Google Restaurants

Creators: Zhankui He, Yan; Li, Jiacheng; Zhang, Tianyang; McAuley, Julian
Publication Date: 2022
Creators: Zhankui He, Yan; Li, Jiacheng; Zhang, Tianyang; McAuley, Julian

This is a mutli-modal dataset of restaurants from Google Local (Google Maps). Data includes images and reviews posted by users, as well as other metadata for each restaurant. The rich combination of textual reviews, numerical ratings, and visual content helps to provide a holistic view of user experiences and restaurant characteristics. Such a multi-faceted dataset is particularly valuable for developing and testing recommendation systems, conducting sentiment analysis, and exploring the relationships between visual content and user perceptions in the context of dining establishment. The total size of the dataset is approximately 120 GB and structured into:

  • Restaurant Metadata: Information such as restaurant names, locations, contact details, and operational hours.

  • User Reviews: Textual feedback and numerical ratings provided by users.

  • Images: Photographs uploaded by users, showcasing various aspects of the restaurants.

Multi-aspect Reviews

Creators: Julian McAuley; Jure Leskovec; Dan Jurafsky
Publication Date: 2013
Creators: Julian McAuley; Jure Leskovec; Dan Jurafsky
These datasets include reviews with multiple rated dimensions.It is particularly valuable for research in sentiment analysis, recommender systems, and user modeling, as it allows for a nuanced understanding of user opinions beyond overall ratings.​The most comprehensive of these are beer review datasets from Ratebeer and Beeradvocate, which include sensory aspects such as taste, look, feel, and smell. The data set is about 1 GB large.
Ratebeer:

  • Number of users: 40,213
  • Number of items: 110,419
  • Number of ratings/reviews: 2,855,232
  • Timespan: April, 2000 – November, 2011

BeerAdvocate:

  • Number of users: 33,387
  • Number of items: 66,051
  • Number of ratings/reviews: 1,586,259
  • Timespan: January, 1998 – November, 2011

The datasets are structured in a JSON format, with each entry representing a single review that includes:

  • Product Information: Details about the beer being reviewed.

  • User Information: Anonymized identifiers of the reviewers.

  • Review Content: Textual feedback provided by the user.

  • Ratings: Numerical scores for overall satisfaction and specific aspects (appearance, aroma, palate, taste).

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.