Classifiers selection based on combination of performance measures for water quality dataset

A classifier is a systematic approach for building a classification model from a dataset.Classification models for certain domain such as medical, toxicology and water quality are difficult to justify its performance because the number of records of the datasets is relatively small. The instances of...

Full description

Bibliographic Details
Main Author: Salisu Yusuf Muhammad (Author)
Corporate Author: Universiti Sultan Zainal Abidin . Faculty of Informatics and Computing
Format: Thesis Book
Language:English
Subjects:
Description
Summary:A classifier is a systematic approach for building a classification model from a dataset.Classification models for certain domain such as medical, toxicology and water quality are difficult to justify its performance because the number of records of the datasets is relatively small. The instances of positive class for the dataset normally are smaller compared to negative class. Thus, by using the accuracy (Ace) alone as a performance measure is not good enough to indicate the performance of the classifiers. Sometimes, there are a number of classifiers from a collection of models having the same Acc, but they may be different. There are other performance measures that can differentiate the models. Thus, there is a need to analyze the details performance of the classifiers and select the most relevant model from the collection. The objective of this thesis is to introduce a new technique that will combine other performance measures (Acc, TPR and TNR) that can provide the detail performance of the classification model in order to overcome the problem of unbalanced datasets. Acc, True Positive Rate (TPR) and True Negative Rate (TNR) can play an important role in providing the detail performance for each model and they are more reliable measurement technique that can be used for comparing the performance of the classification models. This thesis proposes a new technique, Model Selection Technique (MST) for selection and ranking of models from the repository of models by combining three performance measures (Acc, TPR and TNR). The technique provides weightage to each performance measure to find the most suitable model from the repository of models. The knowledge discovery in a database (KDD) has been used in the methodology to generate various classification models for the repository. A number of classification models have been generated to classify water quality using the most significant features and classifiers such as J48, JRip and BayesNet. To validate the technique proposed, the water quality dataset of Kinta River was used in this research. The results demonstrate that the Function classifier is'the optimal model with the most outstanding accuracy of 97.01%, TPR = 0.96 and TNR = 0.98. As a conclusion, the results show that by combining performance measures (Acc, TPR and TNR), as proposed within this thesis, the Acc increased and the distance between TPR and TNR decreased. It gives an important technique for selecting the most robust model in classifying the water quality class of river water.
Physical Description:xiv, 112 leaves : ill. (some col.) ; 30 cm.
Bibliography:Includes bibliographical references (leaves 103-107)