Skip to main content

Showing 1–9 of 9 results for author: Jiménez, Á B

.
  1. arXiv:2503.18594  [pdf, other

    cs.CL cs.AI

    ClinText-SP and RigoBERTa Clinical: a new set of open resources for Spanish Clinical NLP

    Authors: Guillem García Subies, Álvaro Barbero Jiménez, Paloma Martínez Fernández

    Abstract: We present a novel contribution to Spanish clinical natural language processing by introducing the largest publicly available clinical corpus, ClinText-SP, along with a state-of-the-art clinical encoder language model, RigoBERTa Clinical. Our corpus was meticulously curated from diverse open sources, including clinical cases from medical journals and annotated corpora from shared tasks, providing… ▽ More

    Submitted 24 March, 2025; originally announced March 2025.

  2. arXiv:2503.08188  [pdf, other

    cs.CL cs.AI

    RigoChat 2: an adapted language model to Spanish using a bounded dataset and reduced hardware

    Authors: Gonzalo Santamaría Gómez, Guillem García Subies, Pablo Gutiérrez Ruiz, Mario González Valero, Natàlia Fuertes, Helena Montoro Zamorano, Carmen Muñoz Sanz, Leire Rosado Plaza, Nuria Aldama García, David Betancur Sánchez, Kateryna Sushkova, Marta Guerrero Nieto, Álvaro Barbero Jiménez

    Abstract: Large Language Models (LLMs) have become a key element of modern artificial intelligence, demonstrating the ability to address a wide range of language processing tasks at unprecedented levels of accuracy without the need of collecting problem-specific data. However, these versatile models face a significant challenge: both their training and inference processes require substantial computational r… ▽ More

    Submitted 11 March, 2025; originally announced March 2025.

  3. arXiv:2501.16011  [pdf, other

    cs.CL

    MEL: Legal Spanish Language Model

    Authors: David Betancur Sánchez, Nuria Aldama García, Álvaro Barbero Jiménez, Marta Guerrero Nieto, Patricia Marsà Morales, Nicolás Serrano Salas, Carlos García Hernán, Pablo Haya Coll, Elena Montiel Ponsoda, Pablo Calleja Ibáñez

    Abstract: Legal texts, characterized by complex and specialized terminology, present a significant challenge for Language Models. Adding an underrepresented language, such as Spanish, to the mix makes it even more challenging. While pre-trained models like XLM-RoBERTa have shown capabilities in handling multilingual corpora, their performance on domain specific documents remains underexplored. This paper pr… ▽ More

    Submitted 27 January, 2025; originally announced January 2025.

    Comments: 8 pages, 6 figures, 3 tables

  4. arXiv:2501.15990  [pdf, other

    cs.CL

    3CEL: A corpus of legal Spanish contract clauses

    Authors: Nuria Aldama García, Patricia Marsà Morales, David Betancur Sánchez, Álvaro Barbero Jiménez, Marta Guerrero Nieto, Pablo Haya Coll, Patricia Martín Chozas, Elena Montiel Ponsoda

    Abstract: Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded project lead by the Universidad Politécnica de Madrid (UPM) and developed by Instituto de Ingeniería del Conocimiento (IIC) to create a series of state-of-the-art… ▽ More

    Submitted 27 January, 2025; originally announced January 2025.

    Comments: 12 pages, 13 figures, 6 tables

  5. arXiv:2410.16292  [pdf, other

    cs.SE cs.AI cs.CL

    An evaluation of LLM code generation capabilities through graded exercises

    Authors: Álvaro Barbero Jiménez

    Abstract: Large Language Models have shown prominent capabilities in generating functional code from natural language descriptions. However, a standardized way to evaluate these capabilities in an objective and unbiased manner is still to be found. In this paper we review the current evaluation methods available to this end, and run a new evaluation of the performance of one state-of-the-art model (GPT4-o-m… ▽ More

    Submitted 6 October, 2024; originally announced October 2024.

  6. arXiv:2308.02199  [pdf, other

    cs.CL cs.AI cs.LG

    A Survey of Spanish Clinical Language Models

    Authors: Guillem García Subies, Álvaro Barbero Jiménez, Paloma Martínez Fernández

    Abstract: This survey focuses in encoder Language Models for solving tasks in the clinical domain in the Spanish language. We review the contributions of 17 corpora focused mainly in clinical tasks, then list the most relevant Spanish Language Models and Spanish Clinical Language models. We perform a thorough comparison of these models by benchmarking them over a curated subset of the available corpora, in… ▽ More

    Submitted 4 August, 2023; originally announced August 2023.

  7. arXiv:2305.14115  [pdf, other

    cs.LG

    RLBoost: Boosting Supervised Models using Deep Reinforcement Learning

    Authors: Eloy Anguiano Batanero, Ángela Fernández Pascual, Álvaro Barbero Jiménez

    Abstract: Data quality or data evaluation is sometimes a task as important as collecting a large volume of data when it comes to generating accurate artificial intelligence models. In fact, being able to evaluate the data can lead to a larger database that is better suited to a particular problem because we have the ability to filter out data obtained automatically of dubious quality. In this paper we prese… ▽ More

    Submitted 23 May, 2023; originally announced May 2023.

    Comments: 25 pages, 14 figures

  8. arXiv:2302.02412  [pdf, other

    cs.CV cs.AI cs.LG

    Mixture of Diffusers for scene composition and high resolution image generation

    Authors: Álvaro Barbero Jiménez

    Abstract: Diffusion methods have been proven to be very effective to generate images while conditioning on a text prompt. However, and although the quality of the generated images is unprecedented, these methods seem to struggle when trying to generate specific image compositions. In this paper we present Mixture of Diffusers, an algorithm that builds over existing diffusion models to provide a more detaile… ▽ More

    Submitted 5 February, 2023; originally announced February 2023.

    ACM Class: I.2.6

  9. arXiv:2205.10233  [pdf, other

    cs.CL

    RigoBERTa: A State-of-the-Art Language Model For Spanish

    Authors: Alejandro Vaca Serrano, Guillem Garcia Subies, Helena Montoro Zamorano, Nuria Aldama Garcia, Doaa Samy, David Betancur Sanchez, Antonio Moreno Sandoval, Marta Guerrero Nieto, Alvaro Barbero Jimenez

    Abstract: This paper presents RigoBERTa, a State-of-the-Art Language Model for Spanish. RigoBERTa is trained over a well-curated corpus formed up from different subcorpora with key features. It follows the DeBERTa architecture, which has several advantages over other architectures of similar size as BERT or RoBERTa. RigoBERTa performance is assessed over 13 NLU tasks in comparison with other available Spani… ▽ More

    Submitted 3 June, 2022; v1 submitted 27 April, 2022; originally announced May 2022.