Skip to main content

Showing 1–23 of 23 results for author: de la Rosa, J

.
  1. arXiv:2412.09460  [pdf, other

    cs.CL

    The Impact of Copyrighted Material on Large Language Models: A Norwegian Perspective

    Authors: Javier de la Rosa, Vladislav Mikhailov, Lemei Zhang, Freddy Wetjen, David Samuel, Peng Liu, Rolv-Arild Braaten, Petter Mæhlum, Magnus Breder Birkenes, Andrey Kutuzov, Tita Enstad, Hans Christian Farsethås, Svein Arne Brygfjeld, Jon Atle Gulla, Stephan Oepen, Erik Velldal, Wilfred Østgulen, Liljia Øvrelid, Aslak Sira Myhre

    Abstract: The use of copyrighted materials in training language models raises critical legal and ethical questions. This paper presents a framework for and the results of empirically assessing the impact of publisher-controlled copyrighted corpora on the performance of generative large language models (LLMs) for Norwegian. When evaluated on a diverse set of tasks, we found that adding both books and newspap… ▽ More

    Submitted 24 January, 2025; v1 submitted 12 December, 2024; originally announced December 2024.

    Comments: 17 pages, 5 figures, 8 tables. Accepted at NoDaLiDa/Baltic-HLT 2025

  2. arXiv:2402.01917  [pdf, ps, other

    cs.CL

    Whispering in Norwegian: Navigating Orthographic and Dialectic Challenges

    Authors: Per E Kummervold, Javier de la Rosa, Freddy Wetjen, Rolv-Arild Braaten, Per Erik Solberg

    Abstract: This article introduces NB-Whisper, an adaptation of OpenAI's Whisper, specifically fine-tuned for Norwegian language Automatic Speech Recognition (ASR). We highlight its key contributions and summarise the results achieved in converting spoken Norwegian into written forms and translating other languages into Norwegian. We show that we are able to improve the Norwegian Bokmål transcription by Open… ▽ More

    Submitted 2 February, 2024; originally announced February 2024.

  3. arXiv:2310.18110  [pdf, ps, other

    eess.SP

    A Control-Bounded Quadrature Leapfrog ADC

    Authors: Hampus Malmberg, Fredrik Feyling, Jose M. de la Rosa

    Abstract: In this paper, the design flexibility of the control-bounded analog-to-digital converter principle is demonstrated. A band-pass analog-to-digital converter is considered as an application and case study. We show how a low-pass control-bounded analog-to-digital converter can be translated into a band-pass version where the guaranteed stability, converter bandwidth, and signal-to-noise ratio are pre… ▽ More

    Submitted 27 October, 2023; originally announced October 2023.

    Comments: 13 pages and 16 figures

  4. arXiv:2307.14784  [pdf

    physics.app-ph

    New approach to designing functional materials for stealth technology: Radar experiment with bilayer absorbers and optimization of the reflection loss

    Authors: Jaume Calvo de la Rosa, Aleix Bou Comas, Joan Manel Hernandez, Pilar Marin, Jose Maria Lopez-Villegas, Javier Tejada, Eugene M. Chudnovsky

    Abstract: Microwave power absorption by a two-layer system deposited on a metallic surface has been studied in the experimental setup emulating the response to a radar signal. Layers containing hexaferrite and iron powder in a dried paint of thickness under 1mm have been used. The data is analyzed within a theoretical model derived for a bilayer system from the transmission line theory. A good agreement bet… ▽ More

    Submitted 1 October, 2024; v1 submitted 27 July, 2023; originally announced July 2023.

  5. arXiv:2307.01672  [pdf, ps, other

    cs.CL

    Boosting Norwegian Automatic Speech Recognition

    Authors: Javier de la Rosa, Rolv-Arild Braaten, Per Egil Kummervold, Freddy Wetjen, Svein Arne Brygfjeld

    Abstract: In this paper, we present several baselines for automatic speech recognition (ASR) models for the two official written languages in Norway: Bokmål and Nynorsk. We compare the performance of models of varying sizes and pre-training approaches on multiple Norwegian speech datasets. Additionally, we measure the performance of these models against previous state-of-the-art ASR models, as well as on ou… ▽ More

    Submitted 4 July, 2023; originally announced July 2023.

    Comments: 10 pages, 10 figures. Published as Proceedings NoDaLiDa 2023, pages 555--564

    Journal ref: 2023. Boosting Norwegian Automatic Speech Recognition. In Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa), pages 555--564, Tórshavn, Faroe Islands. University of Tartu Library

  6. arXiv:2307.01387  [pdf, other

    cs.CL

    ALBERTI, a Multilingual Domain Specific Language Model for Poetry Analysis

    Authors: Javier de la Rosa, Álvaro Pérez Pozo, Salvador Ros, Elena González-Blanco

    Abstract: The computational analysis of poetry is limited by the scarcity of tools to automatically analyze and scan poems. In a multilingual settings, the problem is exacerbated as scansion and rhyme systems only exist for individual languages, making comparative studies very challenging and time consuming. In this work, we present \textsc{Alberti}, the first multilingual pre-trained large language model f… ▽ More

    Submitted 3 July, 2023; originally announced July 2023.

    Comments: Accepted for publication at SEPLN 2023: 39th International Conference of the Spanish Society for Natural Language Processing

  7. arXiv:2303.03915  [pdf, other

    cs.CL cs.AI

    The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

    Authors: Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Šaško, Quentin Lhoest, Angelina McMillan-Major, Gerard Dupont, Stella Biderman, Anna Rogers, Loubna Ben allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa , et al. (29 additional authors not shown)

    Abstract: As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large language models as a values-driven undertaking, putting issues of ethics, harm, and governance in the f… ▽ More

    Submitted 7 March, 2023; originally announced March 2023.

    Comments: NeurIPS 2022, Datasets and Benchmarks Track

    ACM Class: I.2.7

  8. Quadrature Control-Bounded ADCs

    Authors: Hampus Malmberg, Fredrik Feyling, Jose M de la Rosa

    Abstract: In this paper, the design flexibility of the control-bounded analog-to-digital converter principle is demonstrated by considering band-pass analog-to-digital conversion. We show how a low-pass control-bounded analog-to-digital converter can be translated into a band-pass version where the guaranteed stability, converter bandwidth, and signal-to-noise ratio are preserved while the center frequency… ▽ More

    Submitted 19 February, 2024; v1 submitted 12 November, 2022; originally announced November 2022.

    Comments: 5 pages, 6 figures, submitted to ISCAS 2023

    Journal ref: 2023 IEEE 66th International Midwest Symposium on Circuits and Systems (MWSCAS), Tempe, AZ, USA, 2023, pp. 380-384

  9. arXiv:2211.05100  [pdf, other

    cs.CL

    BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

    Authors: BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major , et al. (369 additional authors not shown)

    Abstract: Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed by resource-rich organizations and are frequently kept from the public. As a step towards democratizing this powerful technology, we present BLOOM, a 176B-parameter open-access… ▽ More

    Submitted 27 June, 2023; v1 submitted 9 November, 2022; originally announced November 2022.

  10. arXiv:2207.06814  [pdf, other

    cs.CL cs.AI

    BERTIN: Efficient Pre-Training of a Spanish Language Model using Perplexity Sampling

    Authors: Javier de la Rosa, Eduardo G. Ponferrada, Paulo Villegas, Pablo Gonzalez de Prado Salas, Manu Romero, Marıa Grandury

    Abstract: The pre-training of large language models usually requires massive amounts of resources, both in terms of computation and data. Frequently used web sources such as Common Crawl might contain enough noise to make this pre-training sub-optimal. In this work, we experiment with different sampling methods from the Spanish version of mC4, and present a novel data-centric technique which we name… ▽ More

    Submitted 14 July, 2022; originally announced July 2022.

    Comments: Published at Procesamiento del Lenguaje Natural

    Journal ref: Procesamiento del Lenguaje Natural, 68 (2022): 13-23

  11. arXiv:2204.05211  [pdf, other

    cs.CL

    Entities, Dates, and Languages: Zero-Shot on Historical Texts with T0

    Authors: Francesco De Toni, Christopher Akiki, Javier de la Rosa, Clémentine Fourrier, Enrique Manjavacas, Stefan Schweter, Daniel van Strien

    Abstract: In this work, we explore whether the recently demonstrated zero-shot abilities of the T0 model extend to Named Entity Recognition for out-of-distribution languages and time periods. Using a historical newspaper corpus in 3 languages as test-bed, we use prompts to extract possible named entities. Our results show that a naive approach for prompt-based zero-shot multilingual Named Entity Recognition… ▽ More

    Submitted 11 April, 2022; originally announced April 2022.

  12. arXiv:2109.11588  [pdf, ps, other

    math.GN

    Relationships between some selection principles and star selection principles

    Authors: Javier Casas de la Rosa

    Abstract: Motivated by the definition of classical star selection principles, Cruz-Castillo, Ramírez-Páramo and Tenorio defined some selection principles and posed several questions about relationships between these notions and some of the classical star selection principles. The main goal of this paper is to answer the questions posed by these authors. In addition, we show that some of these defined select… ▽ More

    Submitted 23 September, 2021; originally announced September 2021.

  13. arXiv:2109.08607  [pdf, other

    cs.CL

    The futility of STILTs for the classification of lexical borrowings in Spanish

    Authors: Javier de la Rosa

    Abstract: The first edition of the IberLEF 2021 shared task on automatic detection of borrowings (ADoBo) focused on detecting lexical borrowings that appeared in the Spanish press and that have recently been imported into the Spanish language. In this work, we tested supplementary training on intermediate labeled-data tasks (STILTs) from part of speech (POS), named entity recognition (NER), code-switching,… ▽ More

    Submitted 17 September, 2021; originally announced September 2021.

    Journal ref: ADoBo 2021 Shared Task IberLEFT@SEPLN, CEUR Workshop Proceedings (Vol. 2943, pp. 947-955)

  14. arXiv:2104.09617  [pdf, other

    cs.CL cs.DL

    Operationalizing a National Digital Library: The Case for a Norwegian Transformer Model

    Authors: Per E Kummervold, Javier de la Rosa, Freddy Wetjen, Svein Arne Brygfjeld

    Abstract: In this work, we show the process of building a large-scale training set from digital and digitized collections at a national library. The resulting Bidirectional Encoder Representations from Transformers (BERT)-based language model for Norwegian outperforms multilingual BERT (mBERT) models in several token and sequence classification tasks for both Norwegian Bokmål and Norwegian Nynorsk. Our mode… ▽ More

    Submitted 19 April, 2021; originally announced April 2021.

    Comments: Accepted to NoDaLiDa 2021

  15. Experimental Body-input Three-stage DC offset Calibration Scheme for Memristive Crossbar

    Authors: Charanraj Mohan, L. A. Camuñas-Mesa, Elisa Vianello, Carlo Reita, José M. de la Rosa, Teresa Serrano-Gotarredona, Bernabé Linares-Barranco

    Abstract: Reading several ReRAMs simultaneously in a neuromorphic circuit increases power consumption and limits scalability. Applying small inference read pulses is a vain attempt when offset voltages of the read-out circuit are decisively more. This paper presents an experimental validation of a three-stage calibration scheme to calibrate the DC offset voltage across the rows of the memristive crossbar. T… ▽ More

    Submitted 3 March, 2021; originally announced March 2021.

    Comments: 5 pages, 9 figures, conference paper published in ISCAS20

    ACM Class: B.7

  16. Implementation of binary stochastic STDP learning using chalcogenide-based memristive devices

    Authors: C. Mohan, L. A. Camuñas-Mesa, J. M. de la Rosa, T. Serrano-Gotarredona, B. Linares-Barranco

    Abstract: The emergence of nano-scale memristive devices encouraged many different research areas to exploit their use in multiple applications. One of the proposed applications was to implement synaptic connections in bio-inspired neuromorphic systems. Large-scale neuromorphic hardware platforms are being developed with increasing number of neurons and synapses, having a critical bottleneck in the online l… ▽ More

    Submitted 1 March, 2021; originally announced March 2021.

    Journal ref: 2021 IEEE International Symposium on Circuits and Systems (ISCAS), 2021, pp. 1-5

  17. arXiv:2011.09567  [pdf, ps, other

    cs.CL

    Predicting metrical patterns in Spanish poetry with language models

    Authors: Javier de la Rosa, Salvador Ros, Elena González-Blanco

    Abstract: In this paper, we compare automated metrical pattern identification systems available for Spanish against extensive experiments done by fine-tuning language models trained on the same task. Despite being initially conceived as a model suitable for semantic tasks, our results suggest that BERT-based models retain enough structural information to perform reasonably well for Spanish scansion.

    Submitted 18 November, 2020; originally announced November 2020.

    Comments: LXAI Workshop @ NeurIPS 2020

  18. arXiv:1611.05360  [pdf

    cs.CL

    The Life of Lazarillo de Tormes and of His Machine Learning Adversities

    Authors: Javier de la Rosa, Juan-Luis Suárez

    Abstract: Summit work of the Spanish Golden Age and forefather of the so-called picaresque novel, The Life of Lazarillo de Tormes and of His Fortunes and Adversities still remains an anonymous text. Although distinguished scholars have tried to attribute it to different authors based on a variety of criteria, a consensus has yet to be reached. The list of candidates is long and not all of them enjoy the sam… ▽ More

    Submitted 16 November, 2016; originally announced November 2016.

    Comments: 66 pages, 11 figures

    Journal ref: Lemir: Revista de Literatura Española Medieval y del Renacimiento, 20 (2016)

  19. The Supernovae Analysis Application (SNAP)

    Authors: Amanda J. Bayless, Chris L. Fryer, Brandon Wiggins, Wesley Even, Ryan Wollaeger, Janie de la Rosa, Peter W. A. Roming, Lucy Frey, Patrick A. Young, Rob Thorpe, Luke Powell, Rachel Landers, Heather D. Persson, Rebecca Hay

    Abstract: The SuperNovae Analysis aPplication (SNAP) is a new tool for the analysis of SN observations and validation of SN models. SNAP consists of an open source relational database with (a) observational light curve, (b) theoretical light curve, and (c) correlation table sets, statistical comparison software, and a web interface available to the community. The theoretical models are intended to span a gr… ▽ More

    Submitted 11 November, 2016; originally announced November 2016.

    Comments: Submitted to AAS publishing, 22 pages, 8 figures

  20. arXiv:1606.09025  [pdf, other

    astro-ph.HE astro-ph.SR

    SN 2015bh: NGC 2770's 4th supernova or a luminous blue variable on its way to a Wolf-Rayet star?

    Authors: C. C. Thöne, A. de Ugarte Postigo, G. Leloudas, C. Gall, Z. Cano, K. Maeda, S. Schulze, S. Campana, K. Wiersema, J. Groh, J. de la Rosa, F. E. Bauer, D. Malesani, J. Maund, N. Morrell, Y. Beletsky

    Abstract: Very massive stars in the final phases of their lives often show unpredictable outbursts that can mimic supernovae, so-called, "SN impostors", but the distinction is not always straigthforward. Here we present observations of a luminous blue variable (LBV) in NGC 2770 in outburst over more than 20 years that experienced a possible terminal explosion as type IIn SN in 2015, named SN 2015bh. This po… ▽ More

    Submitted 29 June, 2016; originally announced June 2016.

    Comments: 29 pages, 20 figures

    Journal ref: A&A 599, A129 (2017)

  21. FRIDA: diffraction-limited imaging and integral-field spectroscopy for the GTC

    Authors: Alan M. Watson, José A. Acosta-Pulido, Luis C. Álvarez-Núñez, Vicente Bringas-Rico, Nicolás Cardiel, Salvador Cuevas, Oscar Chapa, José Javier Díaz García, Stephen S. Eikenberry, Carlos Espejo, Rubén A. Flores-Meza, Jorge Fuentes-Fernández, Jesús Gallego, José Leonardo Garcés Medina, Francisco Garzón López, Peter Hammersley, Carolina Keiman, Gerardo Lara, José Alberto López, Pablo L. López, Diana Lucero, Heidy Moreno Arce, Sergio Pascual Ramirez, Jesús Patrón Recio, Almudena Prieto , et al. (5 additional authors not shown)

    Abstract: FRIDA is a diffraction-limited imager and integral-field spectrometer that is being built for the adaptive-optics focus of the Gran Telescopio Canarias. In imaging mode FRIDA will provide scales of 0.010, 0.020 and 0.040 arcsec/pixel and in IFS mode spectral resolutions of 1500, 4000 and 30,000. FRIDA is starting systems integration and is scheduled to complete fully integrated system tests at the… ▽ More

    Submitted 31 May, 2016; originally announced May 2016.

    Comments: To appear in the proceedings of SPIE conference 9908 "Ground-based and Airborne Instrumentation for Astronomy VI". 8 pages

  22. arXiv:1002.3582  [pdf

    astro-ph.IM astro-ph.SR

    T35: a small automatic telescope for long-term observing campaigns

    Authors: Susana Martin-Ruiz, Francisco J. Aceituno, Miguel Abril, Luis P. Costillo, Antonio Garcia, Jose Luis de la Rosa, Isabel Bustamante, Juan Gutierrez-Soto, Hector Magan, Jose Luis Ramos, Marcos Ubierna

    Abstract: The T35 is a small telescope (14") equipped with a large format CCD camera installed in the Sierra Nevada Observatory (SNO) in Southern Spain. This telescope will be a useful tool for the detecting and studying pulsating stars, particularly, in open clusters. In this paper, we describe the automation process of the T35 and show also some images taken with the new instrumentation.

    Submitted 18 February, 2010; originally announced February 2010.

    Comments: 13 pages, 9 figures. Accepted for publication in the special issue "Robotic Astronomy" of Advances of Astronomy

  23. arXiv:0807.4035  [pdf

    physics.ins-det

    UDP: an integral management system of embedded scripts implemented into the IMaX instrument of the Sunrise mission

    Authors: R. Morales Munoz, P. Mellado, J. Marco de la Rosa, IMaX Team

    Abstract: The UDP (User Defined Program) system is a scripting framework for controlling and extending instrumentation software. It has been specially designed for air- and space-borne instruments with flexibility, error control, reuse, automation, traceability and ease of development as its main objectives. All the system applications are connected through a database containing the valid script commands… ▽ More

    Submitted 25 July, 2008; originally announced July 2008.

    Comments: This paper has been presented in the SPIE 2008, Marselle, France

    Journal ref: Proc.SPIEInt.Soc.Opt.Eng.7019:701916,2008