A Systematic Survey of Business Intelligence Literature Using Machine Learning Techniques

Authors

  • Georgios Fakas Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece Author
  • Elias Houstis Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece Author
  • Manolis Vavalis Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece Author

DOI:

https://doi.org/10.47363/JESMR/2022(3)150

Keywords:

Business Intelligence, Business Intelligence Survey, Research Opportunities, Text Mining, Machine Learning

Abstract

Business intelligence is the field that develops methodologies and tools for analysis of business information to assist the management and decision process of a corporation. The principal aims of this study are a) to complement the existing literature surveys in the BI area by identifying publications for the period 2007 to 2020, b) to classify these publications according to research strategies and various well-defined research topic categories, and c) apply machine learning techniques to assess their ‘Relevance’ with the BI discipline. We have collected 332 papers using ‘Google ‘Scholar’ using a set of related keywords associated with the BI literature. The results show that most papers appeared in 2015 and 2017. The classifications of the literature based on research strategies and topics indicate that most papers address ‘formal theory and/or reviews’ and belong to the ‘benefits’ topic category. For estimating the ‘Relevance’ of the surveyed publications, we extracted information from them using the natural language techniques ‘term-document matrix (TDM)’ and ‘Topics’ and apply machine learning techniques to the generated input feature spaces. The experiments indicate that the overall best individual classifier was the Random Forest with SMOTE sampling on 50% of the original data applied to the ‘Topic’ feature space, achieving 62.12% accuracy. The next best classifier is the Neural Networks with ROSE sampling on 50% of the original data with the ‘TDM input feature space, giving 53% accuracy. The best ensemble type classifier was Neural Networks with ROSE sampling, polynomial SVM without oversampling, and Gradient Boosting without oversampling, which achieved 76.92% accuracy using the ‘Topic’ input feature space

Author Biographies

  • Georgios Fakas, Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece

     Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece

  • Elias Houstis, Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece

     Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece

  • Manolis Vavalis, Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece

     Department of Electrical and Computer Engineering University of Thessaly Sekeri and Heiden Pedion Areos, Zip 383 34, Volos Greece

Downloads

Published

2022-04-02