This repository contains a selection of real-world benchmark datasets for multi-label classification. They are provided in the Mulan data format or the LIBSVM dataset format and originate from the publicly available collections of datasets that are provided by the following projects:
- The MEKA project
- The MULAN project
- The LIBSVM project
- The MLDA tool for analyzing multi-label datasets
The datasets are intended for use with the BOOMER machine learning algorithm and have been used for empirical studies that are concerned with this particular multi-label classification method.
In addition, this repositories does also include toy datasets that are well-suited for debugging purposes due to their small size and may be useful to analyze the behavior of learning algorithms.
Moreover, this repository does also contain several synthetic datasets with varying characteristics, as well as a Python script for generating them, originating from this paper. These datasets may be useful to investigate the ability of different classification methods to deal with conditional and marginal label dependence. The motivation to capture hidden dependencies between labels that can be found in most real-world datasets is a driving force of research on multi-label classification.
This repository uses Git Large File Storage (LFS). It is required to have this open source extension to the Git version control system installed on your computer. Once you have installed the software, the repository can be cloned as usual via the following command:
git clone https://github.com/mrapp-ke/Boomer-Datasets.git
In accordance with the Mulan data format, all datasets that are provided by this repository come as an .arff file that specify the feature values and ground truth labels of the examples they entail. In addition, a .xml file that specifies the names of the available labels is provided for each dataset. In most cases, predefined splits of a dataset into training and test data (indicated by including the suffixes _training and _test in the respective file names) are available as well.
In the following, we provide a description of the datasets that are included in this repository, as well as references to the original authors:
K. Lang. 2008. The 20 newsgroup dataset. http://people.csail.mit.edu/jrennie/20Newsgroups/.
A compilation of around 20,000 newsgroup posts on 20 different topics.
Derek Greene and Pádraig Cunningham. A matrix factorization approach for integrating multiple data views. In Proceedings of the European Conference on Machine Learning and Knowledge Discovery in Databases (ECML-PKDD), pp. 423–438, 2009.
These datasets include 948 news articles covering 416 distinct news stories from the period February – April 2009. They have been collected from 3 sources: BBC, Reuters and The Guardian. Of these stories, 169 were reported in all three sources, 194 in two sources, and 53 appeared in a single news source. Each story was manually annotated with one or more of the six topical labels: business, entertainment, health, politics, sport, technology. In this way, three datasets with the news from BBC, Reuters and The Guardian respectively are created. A feature selection method has been performed in order to reduce the feature space and achieve a better performance. Each dataset has been selected 1000 features. Also, a dataset with the intersection (3Sources-Inter) of these three datasets (news which are in all three sources) has been created with the union of the 1000 features of each one of the datasets.
Ioannis Katakis, Grigorios Tsoumakas, and Ioannis Vlahavas. Multilabel Text Classification for Automated Tag Suggestion. In Proceedings of the ECML/PKDD 2008 Discovery Challenge, 2008.
A dataset that is based on the ECML/PKDD 2008 discovery challenge. It contains 7395 Bibtex entries from the BibSonomy social bookmark and publication sharing system, annotated with a subset of the tags that have been assigned by BibSonomy users.
Forrest Briggs et al. The 9th annual MLSP competition: New methods for acoustic classification of multiple simultaneous bird species in a noisy environment. In IEEE International Workshop on Machine Learning for Signal Processing, pp. 1–8, 2013.
A dataset that is aimed at predicting the set of bird species that one can hear in ten-second audio clips.
Lei Tang and Huan Liu. Relational learning via latent social dimensions. In Proc. ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 817–826, 2009.
This datasets is a node classification problem where each node is associated with multiple labels and features are embedding vectors learned from graphs. Embedding vectors are generated by the following representation-learning methods: DeepWalk, LINE and Node2Vec.
Ioannis Katakis, Grigorios Tsoumakas, and Ioannis Vlahavas. Multilabel Text Classification for Automated Tag Suggestion. In Proceedings of the ECML/PKDD 2008 Discovery Challenge, 2008.
A dataset that is based on the ECML/PKDD 2008 discovery challenge. It contains bookmark entries from the BibSonomy social bookmark and publication sharing system.
Douglas Turnbull, Luke Barrington, David Torres and Gert Lanckriet. Semantic Annotation and Retrieval of Music and Sound Effects, IEEE Transactions on Audio, Speech and Language Processing 16(2), pp. 467-476, 2008.
A music dataset that is composed of 502 songs. Each one was manually annotated with a subset of 174 tags that correspond to 6 semantic concepts: instrumentation, vocal characteristics, genres, emotions, acoustic quality and usage terms.
H. Shao, G.Z. Li, G.P. Liu, and Y.Q. Wang. Symptom selection for multi-label data of inquiry diagnosis in traditional chinese medicine. Science China Information Sciences, 56(5):1–13, 2013.
This dataset contains information about coronary heart disease (CHD) in traditional Chinese medicine (TCM). This dataset has been filtered by specialists, keeping only 49 features.
Pinar Duygulu, Kobus Barnard, Nando de Freitas, and David Forsyth, Object recognition as machine translation: Learning a lexicon for a fixed image vocabulary , 7th European Conference on Computer Vision, pp. 97-112, 2002.
An popular benchmark dataset for image classification that is based on 5,000 Corel images.
Kobus Barnard, Pinar Duygulu, Nando de Freitas, David Forsyth, David Blei, and Michael I. Jordan. Matching Words and Pictures. Journal of Machine Learning Research (3), pp. 1107-1135, 2003.
This dataset consists of 10 parts and is derived from the popular benchmark dataset ECCV 2002 by eliminating less frequently appeared labels.
G. Tsoumakas, I. Katakis, and I. Vlahavas. Effective and Efficient Multilabel Classification in Domains with Large Number of Labels. In Proc. ECML/PKDD 2008 Workshop on Mining Multidimensional Data, 2008.
A dataset that contains the textual data of web pages, together with corresponding tags.
G. Tsoumakas, I. Katakis, and I. Vlahavas. Effective and Efficient Multilabel Classification in Domains with Large Number of Labels. In Proc. ECML/PKDD 2008 Workshop on Mining Multidimensional Data, 2008.
A small dataset that is aimed at classifying music into emotions, based on the Tellegen-Watson-Clark model of mood.
Jesse Read, Bernhard Pfahringer, and Geoff Holmes. Multi-label Classification Using Ensembles of Pruned Sets. In Proceedings of the 2008 Eighth IEEE International Conference on Data Mining, pp. 995–1000, 2008.
This datasets includes a subset of the Enron e-mail corpus, annotated with different topical categories.
Jianhua Xu, Jiali Liu, Jing Yin, and Chengyu Sun. A multi-label feature extraction algorithm via maximizing feature variance and feature-label dependence simultaneously. Knowledge-Based Systems (98), pp. 172 — 184, 2016
A dataset from the field of biology that is used to predict the sub-cellular locations of proteins according to 7,766 sequences for the Eukaryote species. Two variants that come with gene onology feature (GO) or include 20 amino acid, 20 pseudo-amino acid and 400 diptide components (Pse-AAC) are available.
Eneldo Loza Mencía and Johannes Fürnkranz. Efficient Pairwise Multilabel Classification for Large-Scale Problems in the Legal Domain. In Proceedings of the European Conference on Machine Learning and Principles and Practice of Knowledge Disocvery in Databases (ECML-PKDD-2008), pp. 50–65. 2008.
A collection of 19,348 documents about European Union law. It contains many different types of documents, such as treaties or case-law and legislative proposals, which are indexed according to several orthogonal categorization schemes. The most important categorization is provided by the EUROVOC descriptors, which form a topical hierarchy with almost 4,000 categories that are concerned with different aspects of European law.
Adriano Rivolli, Larissa C. Parker, and Andre C.P.L.F. de Carvalho. Food Truck Recommendation Using Multi-label Classification. In EPIA 2017: Progress in Artificial Intelligence, pp. 585–596, 2017.
A dataset that was created from a survey among 407 participants who were either approached at fast food festivals and popular events or anonymously received a request to answer a questionnaire, describing their preferences when it comes to their favorite food trucks.
Sotiris Diplaris, Grigorios Tsoumakas, Pericles Mitkas, and Ioannis Vlahavas. Protein Classification with Multiple Algorithms. In Procedings of the Panhellenic Conference on Informatics, pp. 448–456, 2005.
A dataset for protein function classification. Each example represents a protein and each label corresponds to a protein class.
Jianhua Xu, Jiali Liu, Jing Yin, and Chengyu Sun. A multi-label feature extraction algorithm via maximizing feature variance and feature-label dependence simultaneously. Knowledge-Based Systems (98), pp. 172 — 184, 2016.
This dataset is used to predict the sub-cellular locations of proteins according to their sequences. It contains 1392 sequences for Gram negative bacterial (Gnegative) species. Both the GO (Gene ontology) features and PseAAC (including 20 amino acid, 20 pseudo-amino acid and 400 diptide components) are provided. There are 8 subcellular locations (cell inner membrane, cell outer membrane, cytoplasm, extracellular, fimbrium, flagellum, nucleoid and periplasm).
Jianhua Xu, Jiali Liu, Jing Yin, and Chengyu Sun. A multi-label feature extraction algorithm via maximizing feature variance and feature-label dependence simultaneously. Knowledge-Based Systems (98), pp. 172 — 184, 2016
This dataset is used to predict the sub-cellular locations of proteins according to their sequences. It contains 519 sequences for Gram positive species. Both the GO (gene ontology) features and PseAAC (including 20 amino acid, 20 pseudo-amino acid and 400 diptide components) are provided. There are 4 subcellular locations (cell membrane, cell wall, cytoplasm and extracell).
Jianhua Xu, Jiali Liu, Jing Yin, and Chengyu Sun. A multi-label feature extraction algorithm via maximizing feature variance and feature-label dependence simultaneously. Knowledge-Based Systems (98), pp. 172 — 184, 2016.
This dataset is used to predict the sub-cellular locations of proteins according to their sequences. It contains 3106 sequences for Human species. Both the GO (Gene ontology) features and PseAAC (including 20 amino acid, 20 pseudo-amino acid and 400 diptide components) are provided. There are 14 subcellular locations (centriole, cytoplasm, cytoskeleton, endoplasm reticulum, endosome, extracell, golgi apparatus, lysosome, microsome, mitochondrion, nucleus, peroxisome, plasma membrace, and synapse).
Min-Ling Zhang and Zhi-Hua Zhou. ML-kNN: A lazy learning approach to multi-label learning. Pattern Recognition, 40(7): pp. 2038–2048, 2007.
An image classification dataset that is composed of 2,000 images. Each image was first converted to the CIE Luv space. Afterwards, it was divided into 49 blocks using a 7x7 grid, where in each block the first and second moments (mean and variance) of each band are computed, corresponding to a low-resolution image and computationally inexpensive texture features. Finally, each image is transformed into a 294-dimensional feature vector.
Jesse Read. Scalable multi-label classification. PhD Thesis, University of Waikato, 2010.
This dataset contains 120,919 textual summaries of movie plots that have been obtained from the Internet Movie Database (www.imdb.com). Each summary is labeled with one or more movie genres.
Jesse Read. Scalable multi-label classification. PhD Thesis, University of Waikato, 2010.
A dataset that was created from the Language Log Forum, where various topics related to language were discussed.
C.G.M. Snoek, M.Worring, J.C. van Gemert, J.-M. Geusebroek, A.W.M. Smeulders. The Challenge Problem for Automated Detection of 101 Semantic Concepts in Multimedia, In Proceedings of ACM Multimedia, 421-430. 2006.
A multimedia dataset for generic video indexing, which was extracted from the TRECVID 2005/2006 benchmark. This dataset contains 85 hours of international broadcast news data, annotated with 100 labels. Each video is represented as a 120-dimensional feature vector of numerical features.
John P. Pestian, Christopher Brew, Pawel Matykiewicz, D. J. Hovermale, Neil Johnson, K. Bretonnel Cohen, and Wodzislaw Duch. A shared task involving multi-label classification of clinical free text. In Proceedings of the Workshop on BioNLP 2007: Biological, Translational, and Clinical Language Processing (BioNLP ’07), pp. 97–104, 2007.
This dataset is based on the data that was made available during the Computational Medicine Centers 2007 Medical Natural Language Processing Challenge 10. It consists of 978 clinical text reports, labeled with one or more out of 45 disease codes.
Tat-Seng Chua, Jinhui Tang, Richang Hong, Haojie Li, Zhiping Luo, and Yan-Tao Zheng. “NUS-WIDE: A Real-World Web Image Database from National University of Singapore”, ACM International Conference on Image and Video Retrieval, 2009.
An image classification dataset, where images are represented using 128-D cVLAD+ features.
Thorsten Joachims, Text Categorization with Support Vector Machines: Learning with Many Relevant Features. In Proceeding of the European Conference on Machine Learning (ECML), 1998.
This dataset consists of medical abstracts from the MeSH categories of the year 1991 that should be assigned to 23 cardiovascular disease categories.
Jianhua Xu, Jiali Liu, Jing Yin, and Chengyu Sun. A multi-label feature extraction algorithm via maximizing feature variance and feature-label dependence simultaneously. Knowledge-Based Systems (98), pp. 172 — 184, 2016.
This dataset is used to predict the sub-cellular locations of proteins according to their sequences. It contains 978 sequences for Plant species. Both the GO (Gene ontology) features and PseAAC (including 20 amino acid, 20 pseudo-amino acid and 400 diptide components) are provided. There are 12 subcellular locations (cell membrace, cell wall, chloroplast, cytoplasm, endoplasmic reticulum, extracellular, golgi apparatus, mitochondrion, nucleus, peroxisome, plastid, and vacuole).
Grigorios Tsoumakas and Ioannis Vlahavas. Random k-Labelsets: An Ensemble Method for Multilabel Classification. In Proceedings of the European Conference on Machine Learning (ECML), pp. 406–417, 2007.
A well-known benchmark dataset for text classification that was created from the larger Reuters-RCV1 corpus by selecting a subset of 500 features.
David D. Lewis, Yiming Yang, Tony G. Rose, and Fan Li. RCV1: A new benchmark collection for text categorization research. Journal of Machine Learning Research (5), pp. 361-397, 2004.
This dataset is a well-known benchmark for text classification methods. It has 5 subsets, each one with 6.000 articles assigned into one or more of 101 topics.
Matthew R. Boutell, Jiebo Luo, Xipeng Shen, and Christopher M. Brown. Learning multi-label scene classification. Pattern Recognition, 37(9): pp. 1757–1771, 2004.
An image dataset that contains 2,407 images, annotated with up to 6 labels: beach, sunset, fall foliage, field, mountain and urban. Each image is represented by 294 visual features, corresponding to spatial color moments in the LUV space.
Jesse Read. Scalable multi-label classification. PhD Thesis, University of Waikato, 2010.
A dataset that consists of article blurbs with subject categories, mined from http://slashdot.org.
Francisco Charte and David Charte. Working with multilabel datasets in R: The MLDR package. The R Journal, 7(2):149–162, 2015.
A collection of text classification datasets, generated from text that has been obtained from a selection of Stack Exchange forums.
A. Srivastava, B. Zane-Ulman: Discovering recurring anomalies in text reports regarding complex space systems. In: IEEE Aerospace Conference, 2005.
A subset of the Aviation Safety Reporting System dataset. It contains 28,596 aviation safety text reports about events that took place during a flight and that have been submitted by the flight crew afterwards. The goal is to label each document with the types of problems they describe. The dataset includes 49,060 discrete attribute that correspond to the terms that occur in the text reports. The safety reports are provided with 22 labels, each of them representing a problem type that may appear during a flight.
Ilias Chalkidis, Abhik Jana, Dirk Hartung, Michael Bommarito, Ion Androutsopoulos, Daniel Katz, and Nikolaos Aletras. LexGLUE: A benchmark dataset for legal language understanding in English. In Proc. Annual Meeting of the Association for Computational Linguistics, pp. 4310–4330, 2022.
This dataset contains 50 Terms of Service (ToS) from on-line platforms (e.g., YouTube, Ebay, Facebook, etc.). The dataset has been annotated on the sentence-level with 8 types of unfair contractual terms, meaning terms (sentences) that potentially violate user rights according to EU consumer law. TF-IDF features have been calculated from the raw texts.
Jianhua Xu, Jiali Liu, Jing Yin, and Chengyu Sun. A multi-label feature extraction algorithm via maximizing feature variance and feature-label dependence simultaneously. Knowledge-Based Systems (98), pp. 172 — 184, 2016.
This dataset is used to predict the sub-cellular locations of proteins according to their sequences. It contains 207 sequences for virus species. Both the GO (Gene ontology) features and PseAAC (including 20 amino acid, 20 pseudo-amino acid and 400 diptide components) are provided. There are 6 subcellular locations (viral capsid, host cell membrane, host endoplasm reticulum, host cytoplasm, host nucleus and secreted).
H. Blockeel, S. Džeroski, and J. Grbovic. Simultaneous prediction of multiple chemical parameters of river water quality with tilde. Lecture Notes in Computer Science 1704, pp. 32–40, 1999.
This dataset is used to predict the quality of water of Slovenian rivers, given 16 characteristics such as the temperature, ph, hardness, NO2 or C02.
Yahoo-Arts, Yahoo-Business, Yahoo-Computers, Yahoo-Education, Yahoo-Entertainment, Yahoo-Health, Yahoo-Recreation, Yahoo-Reference, Yahoo-Science, Yahoo-Social, Yahoo-Society
N. Ueda, K. Saito: Parametric mixture models for multi-labeled text, In Neural Information Processing Systems (NIPS), pp. 737-744, 2002.
A collection of text classification datasets. The goal is to annotate web pages with top-level and second-level categories.
Andre Elisseeff and Jason Weston. A kernel method for multi-labelled classification. In Advances in Neural Information Processing Systems, 14: pp. 681–687, 2001.
This dataset contains micro-array expressions and phylogenetic profiles for 2,417 yeast genes. Each gene is annoted with a subset of 14 functional categories (e.g., metabolism, energy, etc.) of the top-level of the functional catalogue.
H. Sajnani, V. Saini, K. Kumar, E. Gabrielova, P. Choudary, C. Lopes. Classifying Yelp reviews into relevant categories. Technical Report, 2013.
This dataset has been obtained from more than 10.000 user reviews and ratings about business and services on Yelp. It is concerned with categorizing whether the food, service, ambiance, deals and price of one of these business are good or not.
Lei Tang and Huan Liu. Relational learning via latent social dimensions. In Proc. ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pp. 817–826, 2009.
This is a node classification problem where each node is associated with multiple labels and features are embedding vectors learned from graphs. Embedding vectors are generated by the following representation-learning methods: DeepWalk, LINE and Node2Vec. After embedding vectors are generated, nodes with no labels are discarded.