books

Showing 1-2 of 2 results

goodbooks-10k

Creators: Zając, Zygmunt
Publication Date: 2017
Creators: Zając, Zygmunt

The dataset contains six million ratings for ten thousand most popular books (with most ratings). It offers a rich resource for analyzing reading habits, book popularity, and user engagement within the literary community. There are also books marked to read by the users, book metadata (author, year, etc.) and tags/shelves/genres.

ratings contains ratings sorted by time. Ratings go from one to five. Both book IDs and user IDs are contiguous. For books, they are 1-10000, for users, 1-53424.

to_read  provides IDs of the books marked “to read” by each user, as user_id,book_id pairs, sorted by time. There are close to a million pairs.

books has metadata for each book (goodreads IDs, authors, title, average rating, etc.). The metadata have been extracted from goodreads XML files.

book_tags contains tags/shelves/genres assigned by users to books. Tags in this file are represented by their IDs. They are sorted by goodreads_book_id  ascending and count descending.

The date set is 68.8 MB large.

Goodreads Datasets

Creators: Wan, Mengting; McAuley, Julian
Publication Date: 2017
Creators: Wan, Mengting; McAuley, Julian

The Goodreads Datasets provide a large-scale collection of book-related data, making them valuable for analyzing reading behavior, book popularity, and recommendation systems. They contain rich metadata on over 2.3 million books, including titles, authors, publication years, genres, and average ratings. Additionally, they feature nearly 229 million user-book interactions, capturing explicit preferences such as bookshelf assignments (“read,” “to-read”) and user ratings. The dataset also covers a vast collection of user-generated textual reviews, offering insights into reader sentiments and opinions. The Goodreads Datasets contain three primary components: book metadata, user-book interactions, and book reviews. The book metadata includes details on 2,360,655 books, such as title, author, publication date, and genre. The user-book interactions dataset comprises 228,648,342 interactions from 876,145 users, capturing activities like adding books to shelves (“read,” “to-read”) and providing ratings. The book reviews dataset contains detailed textual reviews written by users, providing insights into reader sentiments and book popularity.The total size of the combined datasets is approximately 11 GB, available in JSON and CSV formats.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.