Design and Implementation of a Simple TextRank Algorithm for Keyword Extraction

Chinedum E. Amaechi1, Nwamaka Cecilia Unegbu2, Nnaemeka Onyemlukwe3, Overcomer Ifeanyi AlexAnusiuba4

1,4 Lecturer, Department of Computer Science, Nnamdi Azikwe University, Awka, Nigeria.

2 Research Assistant, Department of Computer Science, Nnamdi Azikwe University, Awka, Nigeria.

3 Lecturer, Department of Computer Science, University on the Niger, Uumuya, Nigeria.

Abstract

Automatic keyword extraction is a fundamental task in Natural Language Processing (NLP) that facilitates efficient information retrieval and text summarization. While supervised methods exist, they often require large, labelled datasets, which are scarce. This study proposes the design and implementation of a simple, unsupervised keyword extraction system based on the TextRank algorithm. The system constructs a graph from text where words are nodes and edges represent co-occurrence relationships. The importance of each word is then computed using a graph-based ranking algorithm similar to PageRank. Developed using Python and the Object-Oriented Analysis and Design Methodology (OOADM), the system includes modules for text preprocessing, graph construction, and keyword ranking. Evaluation on a sample of research abstracts demonstrated that the proposed TextRank system achieved an average precision of 82% and an F1-Score of 0.37, outperforming a baseline TF-IDF method. The results indicate that the graph-based approach more effectively captures relevant keywords by considering structural relationships over mere frequency. The system provides a lightweight, adaptable, and efficient solution for keyword extraction across various domains without the need for pre-labelled data.   

Keywords: Keyword extraction, TextRank algorithm, Graph-Based model, Unsupervised learning, Natural Language Processing

References

  1. Asrori, R.B., Setyawan, R. and Muljono, M. (2020) ‘Performance analysis graph-based keyphrase extraction in Indonesia scientific paper’, Int. Seminar on Application for Technology of Information and Communication, pp. 185–190.  
  2. Booch, G., Maksimchuk, R.A., Engle, M.W., et al. (2007) Object-oriented analysis and design with applications, 3rd ed. Houstan: Addison-Wesley.   
  3. Devika, R. and Subramaniyaswamy, V. (2021) ‘Semantic graph-based keyword extraction model using ranking method on big social data’, Wirel. Netw., 27, pp. 5447–5459. 
  4. Khan, M.Q., Shahid, A., Uddin, M.I., et al. (2022), ‘Impact analysis of keyword extraction using contextual word embedding’, PeerJ Computer Sci., 8, e967. 
  5. Mihalcea, R. and Tarau, P. ‘TextRank: bringing order into text’, Proc. 2004 Conf. Empirical Methods Natural Lang. Process., pp. 404–411.
  6. Mothe, J., Ramiandrisoa, F. and Rasolomanana, M. (2018) ‘Automatic keyphrase extraction using graph-based methods’, 33th CM Symposium on Applied Computing (SAC 2018), Pau, France, pp. 728-730.  
  7. Sondhi, P. and Jabbar, A. (2021) ‘Survey on keyphrase extraction using machine learning approaches’, Int. J. Trend Sci. Res. Dev., 5(3), pp. 485–489. 
  8. Ushio, A., Liberatore, F. and Camacho-Collados, J. (2021) ‘Back to the basics: a quantitative analysis of statistical and graph-based term weighting schemes for keyword extraction’, Proc. 2021 Conf. Empirical Methods Natural Lang. Process., pp. 8121–8132. 

Rajshahi Medical College and University of Rajshahi, BANGLADESH.



Royal Melbourne Institute of Technology (RMIT), Melbourne, AUSTRALIA.




Agri. Services, Islamabad Model College for Girls, and Riphah International University, PAKISTAN.




Kampala International University, UGANDA; Rivers State University, NIGERIA.


Discover more from International Journal of Technology, Health and Sustainability

Subscribe now to keep reading and get access to the full archive.

Continue reading