machine learning

Showing 1-5 of 5 results

GitHub – Missing Data Handling Code for Machine Learning Portfolios

Creators: Andrew Y. Chen, Jack McCoy
Publication Date: 2024
Creators: Andrew Y. Chen, Jack McCoy

This GitHub repository contains the code implementation for the study “Missing Values Handling for Machine Learning Portfolios” (Chen & McCoy, 2024, Journal of Financial Economics). It provides replication code for the various missing data handling methods evaluated in the paper, enabling researchers to apply and compare these techniques in their own machine learning portfolio research.

Replication Data: Missing Values Handling for Machine Learning Portfolios

Creators: Andrew Chen , Jack McCoy
Publication Date: 20 February 2024
Creators: Andrew Chen , Jack McCoy

This dataset, hosted on Mendeley Data, contains the replication data for the study “Missing Values Handling for Machine Learning Portfolios” (Chen & McCoy, 2024, Journal of Financial Economics). It includes stock return and firm characteristic data used to examine how different methods for handling missing values in financial datasets affect the performance of machine learning-based portfolio construction strategies, providing guidance on best practices for missing data imputation in empirical asset pricing research.

BERTopic – Neural Topic Modelling Python Library

Creators: Maarten Grootendorst
Publication Date: 2024
Creators: Maarten Grootendorst

BERTopic is an open-source Python library for topic modelling that leverages transformer-based language models (such as BERT) and clustering techniques to discover coherent topics in large text corpora. It enables researchers to extract and analyse thematic structures from unstructured text data.

Replication Data: From Man vs. Machine to Man + Machine – AI and Human Stock Analysis

Creators: Sean Cao, Wei Jiang, Junbo Wang, and Baozhong Yang
Publication Date: 2024-05-31
Creators: Sean Cao, Wei Jiang, Junbo Wang, and Baozhong Yang

This dataset, hosted on Mendeley Data, contains the replication data for the study “From Man vs. Machine to Man + Machine: The Art and AI of Stock Analyses” (Cao, Jiang, Wang & Yang, 2024, Journal of Financial Economics). It includes analyst forecast data and AI-generated stock assessments, used to compare the performance of human analysts versus machine learning models in predicting stock returns, and to examine the complementary effects when human and AI analysis are combined.

Multi-aspect Reviews

Creators: Julian McAuley; Jure Leskovec; Dan Jurafsky
Publication Date: 2013
Creators: Julian McAuley; Jure Leskovec; Dan Jurafsky
These datasets include reviews with multiple rated dimensions.It is particularly valuable for research in sentiment analysis, recommender systems, and user modeling, as it allows for a nuanced understanding of user opinions beyond overall ratings.​The most comprehensive of these are beer review datasets from Ratebeer and Beeradvocate, which include sensory aspects such as taste, look, feel, and smell. The data set is about 1 GB large.
Ratebeer:

  • Number of users: 40,213
  • Number of items: 110,419
  • Number of ratings/reviews: 2,855,232
  • Timespan: April, 2000 – November, 2011

BeerAdvocate:

  • Number of users: 33,387
  • Number of items: 66,051
  • Number of ratings/reviews: 1,586,259
  • Timespan: January, 1998 – November, 2011

The datasets are structured in a JSON format, with each entry representing a single review that includes:

  • Product Information: Details about the beer being reviewed.

  • User Information: Anonymized identifiers of the reviewers.

  • Review Content: Textual feedback provided by the user.

  • Ratings: Numerical scores for overall satisfaction and specific aspects (appearance, aroma, palate, taste).

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.