Automatic Hausa text summarization based on feature extraction using naive Bayes model

As a result of advances in information technology, information overload becomes a global problem. .Automatic text summarization, a branch of natural language processing, is one of the techniques that can be used to overcome the challenge. Automatic text summarization is a technique used to summarize...

Full description

Bibliographic Details
Main Author: Bashir, Muazzam (Author)
Corporate Author: Universiti Sultan Zainal Abidin . Faculty of Informatics and Computing
Format: Thesis Book
Subjects:

MARC

LEADER 00000cam a2200000 7i4500
001 0000089934
005 20230112093000.0
008 160811s2016 my eng
040 |a UniSZA   |e rda 
050 0 0 |a P98.5.A87   |b B37 2016 
090 0 0 |a P98.5.A87   |b B37 2016 
100 1 |a Bashir, Muazzam ,   |e author 
245 1 0 |a Automatic Hausa text summarization based on feature extraction using naive Bayes model   |c Muazzam Bashir 
264 0 |c 2016 
300 |a 111leaves ;   |c 30 cm. 
336 |a text  |2 rdacontent 
337 |a unmediated  |2 rdamedia 
338 |a volume  |2 rdacarrier 
502 |a Thesis (Degree of Master of Science in the Faculty of Informatics and Computing) - Universiti Sultan Zainal Abidin, 2016 
504 |a Includes bibliographical references (leaves 90-96) 
505 0 |a 1. Introduction -- 2. Literature review -- 3. Methodology -- 4. Results and discussion -- 5. Conclusion 
520 |a As a result of advances in information technology, information overload becomes a global problem. .Automatic text summarization, a branch of natural language processing, is one of the techniques that can be used to overcome the challenge. Automatic text summarization is a technique used to summarize a text without losing the essential information. Although there are many commercial text summarization tools available online, there is a limited research to summarize text automatically in Rausa. This study was conducted to develop a system to summarize text automatically in Rausa language. Rausa, a Chadic language that is widely spoken in West Africa, is a low resource language. A data set of 10 Rausa documents were extracted from two different newspapers which are 'Aminiya' and 'Leadership Rausa'. Each document was given to three linguistic experts for human made summary. The study adopted five features (keyword, length, title, cue phrases and location of a sentence) in the summarization process. Rausa morphological rules were reviewed, while Porter's algorithm was modified to fit the language. Meanwhile, a stemming algorithm was developed to stem Rausa terms. Term Frequency Inverse Sentence Frequency and Kmixture Probabilistic models were used to weigh each word before and after stemming. A set of words was chosen as keywords based on a threshold value. The keywords were used to produce summaries based on the models. This is to determine the fitness of the models and the impact of stemming on automatic text summarization for the Rausa language. Moreover, Naive Bayes model was employed to weigh each sentence based on its features. The system has produced a set of summary of sentences based on the threshold value. Considering human made summaries are perfect, the researcher has compared the system generated output to human made summaries. The results show that, the Term Frequency Inverse Sentence Frequency model, having an average F-score of 56.0% outclasses K-mixture Probabilistic model with 38.9%. This is based on automatic text summarization with stemming. The Term Frequency Inverse Sentence Frequency model with 34.8% f-score . has also out performed the K-mixture Probabilistic model with 26.3% based on automatic text summarization without stemming. The overall system testing revealed an average Fscore of 78.1 %. The result obtained from the automatic text summarization when tested on the Rausa language, has proven to be better if the text has been stemmed. 
610 2 0 |a Universiti Sultan Zainal Abidin   |x Dissertations 
610 2 0 |a Universiti Sultan Zainal Abidin   |x Faculty of Informatics and Computing   |v Dissertations 
650 0 |a Automatic abstracting 
650 0 |a Hausa Language 
655 0 |a Dissertations, Academic 
710 2 |a Universiti Sultan Zainal Abidin .   |b Faculty of Informatics and Computing 
999 |a 1000166691   |b Thesis   |c Reference   |e Tembila Campus