-
Unsupervised Broadcast News Summarization; a comparative study on Maximal Marginal Relevance (MMR) and Latent Semantic Analysis (LSA)
Authors:
Majid Ramezani,
Mohammad-Salar Shahryari,
Amir-Reza Feizi-Derakhshi,
Mohammad-Reza Feizi-Derakhshi
Abstract:
The methods of automatic speech summarization are classified into two groups: supervised and unsupervised methods. Supervised methods are based on a set of features, while unsupervised methods perform summarization based on a set of rules. Latent Semantic Analysis (LSA) and Maximal Marginal Relevance (MMR) are considered the most important and well-known unsupervised methods in automatic speech su…
▽ More
The methods of automatic speech summarization are classified into two groups: supervised and unsupervised methods. Supervised methods are based on a set of features, while unsupervised methods perform summarization based on a set of rules. Latent Semantic Analysis (LSA) and Maximal Marginal Relevance (MMR) are considered the most important and well-known unsupervised methods in automatic speech summarization. This study set out to investigate the performance of two aforementioned unsupervised methods in transcriptions of Persian broadcast news summarization. The results show that in generic summarization, LSA outperforms MMR, and in query-based summarization, MMR outperforms LSA in broadcast news summarization.
△ Less
Submitted 5 January, 2023;
originally announced January 2023.
-
Text-based automatic personality prediction: A bibliographic review
Authors:
Ali-Reza Feizi-Derakhshi,
Mohammad-Reza Feizi-Derakhshi,
Majid Ramezani,
Narjes Nikzad-Khasmakhi,
Meysam Asgari-Chenaghlu,
Taymaz Akan,
Mehrdad Ranjbar-Khadivi,
Elnaz Zafarni-Moattar,
Zoleikha Jahanbakhsh-Naghadeh
Abstract:
Personality detection is an old topic in psychology and Automatic Personality Prediction (or Perception) (APP) is the automated (computationally) forecasting of the personality on different types of human generated/exchanged contents (such as text, speech, image, video). The principal objective of this study is to offer a shallow (overall) review of natural language processing approaches on APP si…
▽ More
Personality detection is an old topic in psychology and Automatic Personality Prediction (or Perception) (APP) is the automated (computationally) forecasting of the personality on different types of human generated/exchanged contents (such as text, speech, image, video). The principal objective of this study is to offer a shallow (overall) review of natural language processing approaches on APP since 2010. With the advent of deep learning and following it transfer-learning and pre-trained model in NLP, APP research area has been a hot topic, so in this review, methods are categorized into three; pre-trained independent, pre-trained model based, multimodal approaches. Also, to achieve a comprehensive comparison, reported results are informed by datasets.
△ Less
Submitted 5 September, 2022; v1 submitted 4 October, 2021;
originally announced October 2021.
-
Phraseformer: Multimodal Key-phrase Extraction using Transformer and Graph Embedding
Authors:
Narjes Nikzad-Khasmakhi,
Mohammad-Reza Feizi-Derakhshi,
Meysam Asgari-Chenaghlu,
Mohammad-Ali Balafar,
Ali-Reza Feizi-Derakhshi,
Taymaz Rahkar-Farshi,
Majid Ramezani,
Zoleikha Jahanbakhsh-Nagadeh,
Elnaz Zafarani-Moattar,
Mehrdad Ranjbar-Khadivi
Abstract:
Background: Keyword extraction is a popular research topic in the field of natural language processing. Keywords are terms that describe the most relevant information in a document. The main problem that researchers are facing is how to efficiently and accurately extract the core keywords from a document. However, previous keyword extraction approaches have utilized the text and graph features, th…
▽ More
Background: Keyword extraction is a popular research topic in the field of natural language processing. Keywords are terms that describe the most relevant information in a document. The main problem that researchers are facing is how to efficiently and accurately extract the core keywords from a document. However, previous keyword extraction approaches have utilized the text and graph features, there is the lack of models that can properly learn and combine these features in a best way.
Methods: In this paper, we develop a multimodal Key-phrase extraction approach, namely Phraseformer, using transformer and graph embedding techniques. In Phraseformer, each keyword candidate is presented by a vector which is the concatenation of the text and structure learning representations. Phraseformer takes the advantages of recent researches such as BERT and ExEm to preserve both representations. Also, the Phraseformer treats the key-phrase extraction task as a sequence labeling problem solved using classification task.
Results: We analyze the performance of Phraseformer on three datasets including Inspec, SemEval2010 and SemEval 2017 by F1-score. Also, we investigate the performance of different classifiers on Phraseformer method over Inspec dataset. Experimental results demonstrate the effectiveness of Phraseformer method over the three datasets used. Additionally, the Random Forest classifier gain the highest F1-score among all classifiers.
Conclusions: Due to the fact that the combination of BERT and ExEm is more meaningful and can better represent the semantic of words. Hence, Phraseformer significantly outperforms single-modality methods.
△ Less
Submitted 9 June, 2021;
originally announced June 2021.
-
Multimodal price prediction
Authors:
Aidin Zehtab-Salmasi,
Ali-Reza Feizi-Derakhshi,
Narjes Nikzad-Khasmakhi,
Meysam Asgari-Chenaghlu,
Saeideh Nabipour
Abstract:
Price prediction is one of the examples related to forecasting tasks and is a project based on data science. Price prediction analyzes data and predicts the cost of new products. The goal of this research is to achieve an arrangement to predict the price of a cellphone based on its specifications. So, five deep learning models are proposed to predict the price range of a cellphone, one unimodal an…
▽ More
Price prediction is one of the examples related to forecasting tasks and is a project based on data science. Price prediction analyzes data and predicts the cost of new products. The goal of this research is to achieve an arrangement to predict the price of a cellphone based on its specifications. So, five deep learning models are proposed to predict the price range of a cellphone, one unimodal and four multimodal approaches. The multimodal methods predict the prices based on the graphical and non-graphical features of cellphones that have an important effect on their valorizations. Also, to evaluate the efficiency of the proposed methods, a cellphone dataset has been gathered from GSMArena. The experimental results show 88.3% F1-score, which confirms that multimodal learning leads to more accurate predictions than state-of-the-art techniques.
△ Less
Submitted 2 April, 2021; v1 submitted 9 July, 2020;
originally announced July 2020.
-
Automatic Personality Prediction; an Enhanced Method Using Ensemble Modeling
Authors:
Majid Ramezani,
Mohammad-Reza Feizi-Derakhshi,
Mohammad-Ali Balafar,
Meysam Asgari-Chenaghlu,
Ali-Reza Feizi-Derakhshi,
Narjes Nikzad-Khasmakhi,
Mehrdad Ranjbar-Khadivi,
Zoleikha Jahanbakhsh-Nagadeh,
Elnaz Zafarani-Moattar,
Taymaz Rahkar-Farshi
Abstract:
Human personality is significantly represented by those words which he/she uses in his/her speech or writing. As a consequence of spreading the information infrastructures (specifically the Internet and social media), human communications have reformed notably from face to face communication. Generally, Automatic Personality Prediction (or Perception) (APP) is the automated forecasting of the pers…
▽ More
Human personality is significantly represented by those words which he/she uses in his/her speech or writing. As a consequence of spreading the information infrastructures (specifically the Internet and social media), human communications have reformed notably from face to face communication. Generally, Automatic Personality Prediction (or Perception) (APP) is the automated forecasting of the personality on different types of human generated/exchanged contents (like text, speech, image, video, etc.). The major objective of this study is to enhance the accuracy of APP from the text. To this end, we suggest five new APP methods including term frequency vector-based, ontology-based, enriched ontology-based, latent semantic analysis (LSA)-based, and deep learning-based (BiLSTM) methods. These methods as the base ones, contribute to each other to enhance the APP accuracy through ensemble modeling (stacking) based on a hierarchical attention network (HAN) as the meta-model. The results show that ensemble modeling enhances the accuracy of APP.
△ Less
Submitted 8 June, 2022; v1 submitted 9 July, 2020;
originally announced July 2020.
-
A Model to Measure the Spread Power of Rumors
Authors:
Zoleikha Jahanbakhsh-Nagadeh,
Mohammad-Reza Feizi-Derakhshi,
Majid Ramezani,
Taymaz Akan,
Meysam Asgari-Chenaghlu,
Narjes Nikzad-Khasmakhi,
Ali-Reza Feizi-Derakhshi,
Mehrdad Ranjbar-Khadivi,
Elnaz Zafarani-Moattar,
Mohammad-Ali Balafar
Abstract:
With technologies that have democratized the production and reproduction of information, a significant portion of daily interacted posts in social media has been infected by rumors. Despite the extensive research on rumor detection and verification, so far, the problem of calculating the spread power of rumors has not been considered. To address this research gap, the present study seeks a model t…
▽ More
With technologies that have democratized the production and reproduction of information, a significant portion of daily interacted posts in social media has been infected by rumors. Despite the extensive research on rumor detection and verification, so far, the problem of calculating the spread power of rumors has not been considered. To address this research gap, the present study seeks a model to calculate the Spread Power of Rumor (SPR) as the function of content-based features in two categories: False Rumor (FR) and True Rumor (TR). For this purpose, the theory of Allport and Postman will be adopted, which it claims that importance and ambiguity are the key variables in rumor-mongering and the power of rumor. Totally 42 content features in two categories "importance" (28 features) and "ambiguity" (14 features) are introduced to compute SPR. The proposed model is evaluated on two datasets, Twitter and Telegram. The results showed that (i) the spread power of False Rumor documents is rarely more than True Rumors. (ii) there is a significant difference between the SPR means of two groups False Rumor and True Rumor. (iii) SPR as a criterion can have a positive impact on distinguishing False Rumors and True Rumors.
△ Less
Submitted 17 June, 2022; v1 submitted 18 February, 2020;
originally announced February 2020.