Showing 473-480 of 573 results

IRS E-File Bucket

Creators: Internal Revenue Service
Publication Date: 2016
Creators: Internal Revenue Service

This bucket contains a mirror of the IRS e-file release as of December 31, 2016, which are annual information returns submitted by tax-exempt organizations in the United States. The data helps to understand the financial and operational aspects of nonprofit organizations. Each Form 990 provides insights into an organization’s mission, programs, and governance structures.​ The forms include detailed financial data, such as revenues, expenses, assets, and liabilities, offering a clear view of an organization’s financial health. As mandated by law, these forms are publicly accessible, promoting transparency and allowing stakeholders to make informed decision. In total, the dataset has a size of 5,3 kB and is divided into individual Form 990 filings, each corresponding to a specific tax-exempt organization. Each filing includes:

  • Organizational Details: Name, Employer Identification Number (EIN), address, and mission statement.

  • Financial Information: Detailed breakdowns of revenues (e.g., contributions, grants, program service revenue), expenses (e.g., salaries, grants, operational costs), assets, and liabilities.

  • Governance and Compliance: Information on board members, key employees, governance policies, and compliance with tax regulations.

Large-scale CelebFaces Attributes (CelebA) Dataset

Creators: Liu, Ziwei; Luo, Ping; Wang, Xiaogang; Tang, Xiaoou
Publication Date: 2015
Creators: Liu, Ziwei; Luo, Ping; Wang, Xiaogang; Tang, Xiaoou

CelebFaces Attributes Dataset (CelebA) is a large-scale face attributes dataset with more than 200K celebrity images, each with 40attribute annotations. The images in this dataset cover large pose variations and background clutter. CelebA has large diversities, large quantities, and rich annotations, including: 10,177 number of identities, 202,599 number of face images, and 5 landmark locations, 40 binary attributes annotations per image. Each image in the dataset captures various facial features and accessories, such as eyeglasses, smiling, or bangs. Additionally, five landmark points (e.g., eyes, nose, mouth corners) are provided per image, facilitating tasks like facial alignment. Also, a wide range of poses, expressions, and occlusions are included, reflecting real-world conditions and enhancing the robustness of models trained on this data. The dataset has a size of 25,3 kB and is organized in theree main components:

  • Images:

    • In-the-Wild Images: Original images depicting celebrities in various environments and conditions.

    • Aligned and Cropped Images: Faces have been aligned and cropped to a consistent size, facilitating standardized analysis.

  • Annotations:

    • Landmark Locations: Coordinates for five key facial points (left eye, right eye, nose, left mouth corner, right mouth corner) per image.

    • Attribute Labels: Binary labels indicating the presence or absence of 40 distinct facial attributes for each image.

    • Identity Labels: Each image is associated with an identity label, linking it to one of the 10,177 unique individuals.

  • Evaluation Partitions:

    • The dataset is divided into training, validation, and test sets, enabling standardized evaluation of algorithms.

Steam Video Game Database

Creators: Beliaev, Volodymyr
Publication Date: 2023
Creators: Beliaev, Volodymyr

This dataset aggregates information on all games available on the Steam platform, enriched with additional data from sources like Steam Spy, GameFAQs, Metacritic, IGDB, and HowLongToBeat (HLTB). It is particularly valuable for researchers, developers, and enthusiasts interested in analyzing various aspects of video games, such as pricing, ratings, and gameplay duration. Each entry provides detailed data, including game identifiers, store URLs, promotional content, user scores, release dates, descriptions, pricing, supported platforms, developers, publishers, available languages, genres, tags, and achievements. The dataset reflects the state of the Steam catalog as of 2023 and has a size of 7,1 kB.

Amazon Brand and Exclusives

Creators: Jeffries, Adrianne; Yin, Leon
Publication Date: 2021
Creators: Jeffries, Adrianne; Yin, Leon
My co-author Adrianne Jeffries and I found Amazon gave its own branded products an advantage over better-rated competitors in search results. This repository contains code to reproduce the findings featured in our story “Amazon Puts Its Own ‘Brands’ First Above Better-Rated Products” and “When Amazon Takes the Buy Box, it Doesn’t Give it up” from our series Amazon’s Advantage. Each product in this dataset is identified by its unique Amazon Standard Identification Number (ASIN), facilitating precise tracking and analysis. They are categorized based on their association with Amazon, distinguishing between Amazon-owned brands, exclusive partnerships, and proprietary electronics. Data is derived from extensive web scraping of Amazon’s product listings, ensuring a comprehensive and up-to-date collection. The data collection occurred primarily in early 2021, with search results gathered in January 2021 and product pages in February 2021. In total, the dataset comprises 137,428 products, each represented by a unique ASIN and has a size of 221,0 kB. It is organized into several sub-datasets, each serving a specific analytical purpose:

  1. Amazon Private Label Products: Contains detailed information on 137,428 products identified as Amazon brands, exclusives, or proprietary electronics.

  2. Search Results: Includes parsed search result pages from top and generic searches, totaling 187,534 product positions.

  3. Product Pages: Comprises parsed product pages corresponding to the search results, encompassing 157,405 product pages.

  4. Training Set: Provides metadata used to train and evaluate machine learning models, with feature engineering conducted in associated Jupyter notebooks.

  5. Trademarks: Contains a dataset of trademarked brands registered by Amazon, collected from USPTO.gov and Amazon.

Fact Check Data

Creators: Data Commons
Publication Date: 2019
Creators: Data Commons

This is a data feed of ClaimReview markups created via the Google Fact Check Markup Tool and the new ClaimReview Read/Write API. The data in the feed also follows the schema.org ClaimReview standard, namely the same schema as the data in the historical research dataset. It compiles structured metadata from fact-checking articles and serves as a valuable resource for researchers and developers aiming to analyze and combat misinformation. Each entry follows the ClaimReview schema, providing standardized fields such as the claim reviewed, the author, the date of publication, and the URL to the original fact-checking article. It aggregates fact-checks from multiple reputable organizations, offering a comprehensive view of fact-checking efforts across various domains​. In total, the dataset has a size of 73,4 kB.

Facebook Social Connectedness Index

Creators: Meta
Publication Date: 2021
Creators: Meta

We use an anonymized snapshot of all active Facebook users and their friendship networks to measure the intensity of connectedness between locations. The Social Connectedness Index (SCI) is a measure of the social connectedness between different geographies. Specifically, it measures the relative probability that two individuals across two locations are friends with each other on Facebook. Each entry represents a pair of locations, detailing the strength of social connectedness between them. By doing so, the SCI provides a measure of the relative probability that two individuals from different locations are Facebook friends, offering insights into social ties across regions. The dataset has a a size of 3,9 kB and reflects a specific snapshot in time, with the latest available data from October 2021. The dataset is organized into multiple sub-datasets, each detailing social connectedness at different geographic levels:

  1. Country-Country Pairs:

    • user_loc: ISO2 code of the first country.

    • fr_loc: ISO2 code of the second country.

    • scaled_sci: Scaled Social Connectedness Index between the two countries.

  2. US County-Country Pairs:

    • user_loc: 5-digit FIPS code of the U.S. county.

    • fr_loc: ISO2 code of the country.

    • scaled_sci: Scaled Social Connectedness Index between the U.S. county and the country.

Creators: ProPublica

This free download is a database of more than 12,000 civilian complaints filed against New York City police officers. After New York state repealed the statute that kept police disciplinary records secret, known as 50-a, ProPublica filed a records request with New York City’s Civilian Complaint Review Board, which investigates complaints by the public about NYPD officers. The board provided us with records about closed cases for every police officer still on the force as of late June 2020 who had at least one substantiated allegation against them. The records span decades, from September 1985 to January 2020. Each entry includes specifics such as the nature of the allegation (e.g., use of force, abuse of authority), the outcome of the investigation, and any disciplinary actions taken. The dataset provides information on the officers involved, including their rank and assignment at the time of the complaint. Entries contain timestamps and locations of the alleged incidents, facilitating analyses of patterns over time and across different areas.​​

Structurally, the dataset contains the following information:

  • Complaint ID: A unique identifier for each complaint.

  • Date and Time: When the incident allegedly occurred.

  • Location: Where the incident took place.

  • Officer Details: Information about the officer(s) involved, such as badge number, rank, and assignment.

  • Allegation Details: Type of misconduct reported (e.g., excessive force, discourtesy).

  • Investigation Outcome: Findings of the investigation, including whether the allegation was substantiated, unsubstantiated, exonerated, or unfounded.

  • Disciplinary Action: Any penalties or corrective actions imposed following the investigation.

The greatest hip-hop songs off all time

Creators: BBC
Publication Date: 2019
Creators: BBC

BBC Music polled over 1001 critics in 15 countries to find the best hip-hop song ever. This repo contains poll data, originally published by BBC Music, as well as code for transforming the data, adding cover artwork, and publishing charts via Datawrapper. The poll data was extracted from this article on bbc.com: The greatest hip-hop songs of all time – who voted. It covered over 300 hip-hop songs, each representing an individual observation. The songs listed span from the late 1970s to 2019, covering the evolution of hip-hop over approximately four decades. In total, the dataset has a size of 60,8 kB. Structurally, it is built upon the following key variables:

  • Rank: The position of the song in the overall ranking.

  • Title: The name of the song.

  • Artist: The performing artist(s) of the song.

  • Year: The release year of the song.

  • Total Points: The cumulative points the song received based on the poll’s scoring system.

  • Number of Votes: The count of critics who included the song in their top five lists.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.