Showing 465-472 of 573 results

AudioSet dataset

Creators: Google
Publication Date: 2017
Creators: Google

AudioSet consists of an expanding ontology of 632 audio event classes and a collection of 2,084,320 human-labeled 10-second sound clips drawn from YouTube videos. The ontology is specified as a hierarchical graph of event categories, covering a wide range of human and animal sounds, musical instruments and genres, and common everyday environmental sounds. The dataset has a size of 19,0 kB and is divided into three primary subsets:

  • Evaluation Set: Contains 20,383 segments from distinct videos, ensuring at least 59 examples for each of the 527 sound classes used.

  • Balanced Training Set: Consists of 22,176 segments from distinct videos, selected to provide a balanced representation with at least 59 examples per class.

  • Unbalanced Training Set: Includes 2,042,985 segments from distinct videos, representing the remainder of the dataset

U.S.-Mexico Border Surveillance Data

Creators: Electronic Frontier Foundaton (EFF)
Publication Date: 2024
Creators: Electronic Frontier Foundaton (EFF)

This dataset includes the locations of Customs & Border Patrol surveillance towers, proposed tower locations, and automated license plate readers. There is an accompanying blog post and map. The dataset includes precise locations of Customs & Border Protection (CBP) surveillance towers, proposed tower sites, automated license plate readers, aerostats (tethered surveillance balloons), and facial recognition systems at land ports of entry. This extensive mapping offers valuable insights into the deployment and reach of surveillance infrastructure along the border. By making this data publicly available, the database facilitates research into the implications of surveillance practices on civil liberties and border communities. The dataset was last updated on November 21, 2024. It reflects the state of surveillance infrastructure up to that date and has a total size of 187,3 kB. Structurally, the dataset is organized into several categories, each detailing a specific type of surveillance technology:

  • Surveillance Towers: Locations and specifications of existing and proposed CBP surveillance towers.

  • Automated License Plate Readers (ALPRs): Positions of ALPR systems used to monitor vehicle movements across the border.

  • Aerostats: Details on tethered surveillance balloons employed for aerial monitoring.

  • Facial Recognition Systems: Information on the deployment of facial recognition technology at land ports of entry.

The Upworthy Research Archive

Creators: The Upworthy Research Archive
Publication Date: 2019
Creators: The Upworthy Research Archive

The Upworthy Research Archive is an open dataset of thousands of A/B tests of headlines conducted by Upworthy from January 2013 to April 2015. This repository includes the full data from the archive. The dataset’s size is approximately 149,7 MB. It includes 32,488 records of headline experiments, providing insights into how different headline variations impacted user engagement. The dataset is structured as a time series of experiments, with each record detailing the performance metrics of different headline variations. This structure enables researchers to analyze the effectiveness of various headlines and understand user engagement patterns over time.

News Homepage Archive

Creators: Jones, Nick
Publication Date: 2019
Creators: Jones, Nick

This project aims to provide a visual representation of how different media organizations cover various topics. Screenshots of the homepages of five different news organizations are taken once per hour, and made public thereafter. For each website, this amounts to 24 screenshots per day. Over a year, this results in approximately 8,760 screenshots per website. Screenshots are available at every hour starting from January 1, 2019. The size of the dataset is 1,8 MB. Currently, the only websites being tracked are:
nytimes.com;
washingtonpost.com;
cnn.com;
wsj.com;
foxnews.com;
By capturing hourly screenshots, this dataset offers a unique visual chronicle of news presentation, allowing for analysis of editorial choices, headline prominence, and the evolution of news stories across different media outlets. The dataset is organized hierarchically based on the website name and timestamp of each screenshot. Each sub-dataset corresponds to a specific news website, containing a chronological collection of its homepage screenshots. This structure facilitates targeted analysis of individual news outlets over time.

Creators: Shutterstock

With millions of images in our library and billions of user-submitted keywords, we work hard at Shutterstock to make sure that bad words don’t show up in places they shouldn’t. This repo, published in 2019, contains a list of words that we use to filter results from our autocomplete server and recommendation engine. The dataset encompasses offensive terms in multiple languages. It is open for contributions, allowing users to add or refine entries, particularly in non-English languages, enhancing its comprehensiveness and applicability across diverse cultural contexts. The exact number of entries varies by language. For instance, the English list contains 403 entries. In total, the dataset has a size of 25,7 kB. The data is organized into separate files for each language, with each file containing a list of offensive words in that particular language. For example, the English words are listed in the ‘en’ file, German words in the ‘de’ file, and so on. This allows the targeted application of language-specific content filtering systems. Each sub-dataset (language file) consists of a plain text file with one offensive term per line, facilitating easy integration into various text processing pipelines.

Africapolis data

Creators: OECD; SWAC
Publication Date: 2022
Creators: OECD; SWAC

Africapolis has been designed to provide a much needed standardised and geospatial database on urbanisation dynamics in Africa, with the aim of making urban data in Africa comparable across countries and across time. This version of Africapolis is the first time that the data for the 54 countries currently covered are available for the same base year — 2015. In addition, Africapolis closes one major data gap by integrating 7,496 small towns and intermediary cities between 10,000 and 300,000 inhabitants. Africapolis data is based on a large inventory of housing and population censuses, electoral registers and other official population sources, in some cases dating back to the beginning of the 20th century. The dataset has a size of 34,5 kB and contains the following information:

  • Spatial Data: Urban agglomerations are represented as polygon vector data, delineating the spatial extent of settlements. Each polygon’s size corresponds to the settled area, providing insights into urban sprawl and density.

  • Attribute Data: Each spatial unit is linked to demographic attributes, such as total population figures for specific years (e.g., 2015, 2020). This linkage enables analyses of population distribution and urban growth patterns.

  • Data Layers: The dataset includes multiple layers corresponding to different years, supporting temporal analyses of urbanization trends.

3 Million Russian troll tweets

Creators: FiveThirtyEight; Warren, Patrick ;Linvill, Darren
Publication Date: 2018
Creators: FiveThirtyEight; Warren, Patrick ;Linvill, Darren

This directory contains data on nearly 3 million tweets sent from Twitter handles connected to the Internet Research Agency, a Russian “troll factory” and a defendant in an indictment filed by the Justice Department in February 2018, as part of special counsel Robert Mueller’s Russia investigation. The tweets in this database were sent between February 2012 and May 2018, with the vast majority posted from 2015 through 2017. Each entry includes detailed information such as the tweet’s content, author handle, language, publication date, and engagement metrics (e.g., number of followers, following count). The dataset provides classifications for each account, indicating the thematic focus (e.g., Right Troll, Left Troll, News Feed), as coded by researchers Darren Linvill and Patrick Warren.​ It has a total size of 507,2 kB.

Computer generated building footprints for the United States

Creators: Microsoft Bing Maps Team
Publication Date: 2018
Creators: Microsoft Bing Maps Team

Microsoft Maps is releasing country wide open building footprints datasets in United States. This dataset contains 129,591,852 computer generated building footprints derived using our computer vision algorithms on satellite imagery. Building footprints were extracted using deep neural networks for semantic segmentation, followed by polygonization to convert detected building pixels into vector shapes. This data is freely available for download and use. The dataset is organized by U.S. state and provided in GeoJSON format. Each GeoJSON file contains polygon geometries representing building footprints, accompanied by metadata such as the capture date of the underlying imagery. Notably, footprints within specific regions are based on imagery from 2019-2020, accounting for approximately 73,250,745 buildings.

Sign In

Register

Reset Password

Please enter your username or email address, you will receive a link to create a new password via email.