Search | arXiv e-print repository

Input-length-shortening and text generation via attention values

Authors: Neşet Özkan Tan, Alex Yuxuan Peng, Joshua Bensemann, Qiming Bao, Tim Hartill, Mark Gahegan, Michael Witbrock

Abstract: Identifying words that impact a task's performance more than others is a challenge in natural language processing. Transformers models have recently addressed this issue by incorporating an attention mechanism that assigns greater attention (i.e., relevance) scores to some words than others. Because of the attention mechanism's high computational cost, transformer models usually have an input-leng… ▽ More Identifying words that impact a task's performance more than others is a challenge in natural language processing. Transformers models have recently addressed this issue by incorporating an attention mechanism that assigns greater attention (i.e., relevance) scores to some words than others. Because of the attention mechanism's high computational cost, transformer models usually have an input-length limitation caused by hardware constraints. This limitation applies to many transformers, including the well-known bidirectional encoder representations of the transformer (BERT) model. In this paper, we examined BERT's attention assignment mechanism, focusing on two questions: (1) How can attention be employed to reduce input length? (2) How can attention be used as a control mechanism for conditional text generation? We investigated these questions in the context of a text classification task. We discovered that BERT's early layers assign more critical attention scores for text classification tasks compared to later layers. We demonstrated that the first layer's attention sums could be used to filter tokens in a given sequence, considerably decreasing the input length while maintaining good test accuracy. We also applied filtering, which uses a compute-efficient semantic similarities algorithm, and discovered that retaining approximately 6\% of the original sequence is sufficient to obtain 86.5\% accuracy. Finally, we showed that we could generate data in a stable manner and indistinguishable from the original one by only using a small percentage (10\%) of the tokens with high attention scores according to BERT's first layer. △ Less

Submitted 13 March, 2023; originally announced March 2023.

Comments: 7 pages, 4 figures. AAAI23-EMC2

arXiv:1109.4528 [pdf, ps, other]

Δ- convergence on time scale

Authors: M. Seyyit Seyyidoglu, N. Özkan Tan

Abstract: In the present paper we will give some new notions, such as Δ-convergence and Δ-Cauchy, by using the Δ-density and investigate their relations. It is important to say that, the results presented in this work generalize some of the results mentioned in the theory of statistical convergence. In the present paper we will give some new notions, such as Δ-convergence and Δ-Cauchy, by using the Δ-density and investigate their relations. It is important to say that, the results presented in this work generalize some of the results mentioned in the theory of statistical convergence. △ Less

Submitted 21 September, 2011; originally announced September 2011.

Comments: 7 pages

arXiv:1104.5388 [pdf, ps, other]

Integral Transformations Between Some Function Spaces On Time Scales

Authors: Mustafa Seyyit Seyyidoglu, Neset Ozkan Tan

Abstract: In this paper we defined some function spaces on time scale which are Banach spaces respect to supremum norm. We study integral transformations which are carry to some important properties between mentioned above function spaces. In this paper we defined some function spaces on time scale which are Banach spaces respect to supremum norm. We study integral transformations which are carry to some important properties between mentioned above function spaces. △ Less

Submitted 28 April, 2011; originally announced April 2011.

Comments: 9 pages

MSC Class: 39A10

Showing 1–3 of 3 results for author: Tan, N Ö