Skip to main content

Showing 1–4 of 4 results for author: Martinson, S

.
  1. arXiv:2503.22719  [pdf, other

    cs.AI

    LLM-based Agent Simulation for Maternal Health Interventions: Uncertainty Estimation and Decision-focused Evaluation

    Authors: Sarah Martinson, Lingkai Kong, Cheol Woo Kim, Aparna Taneja, Milind Tambe

    Abstract: Agent-based simulation is crucial for modeling complex human behavior, yet traditional approaches require extensive domain knowledge and large datasets. In data-scarce healthcare settings where historic and counterfactual data are limited, large language models (LLMs) offer a promising alternative by leveraging broad world knowledge. This study examines an LLM-driven simulation of a maternal mobil… ▽ More

    Submitted 25 March, 2025; originally announced March 2025.

  2. arXiv:2501.14249  [pdf, other

    cs.LG cs.AI cs.CL

    Humanity's Last Exam

    Authors: Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Dmitry Dodonov, Tung Nguyen, Jaeho Lee, Daron Anderson, Mikhail Doroshenko, Alun Cennyth Stokes , et al. (1084 additional authors not shown)

    Abstract: Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of… ▽ More

    Submitted 19 April, 2025; v1 submitted 24 January, 2025; originally announced January 2025.

    Comments: 29 pages, 6 figures

  3. arXiv:2410.09988  [pdf, other

    cs.LG cs.AI

    HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics

    Authors: Jingxuan Fan, Sarah Martinson, Erik Y. Wang, Kaylie Hausknecht, Jonah Brenner, Danxian Liu, Nianli Peng, Corey Wang, Michael P. Brenner

    Abstract: Advanced applied mathematics problems are underrepresented in existing Large Language Model (LLM) benchmark datasets. To address this, we introduce HARDMath, a dataset inspired by a graduate course on asymptotic methods, featuring challenging applied mathematics problems that require analytical approximation techniques. These problems demand a combination of mathematical reasoning, computational t… ▽ More

    Submitted 13 December, 2024; v1 submitted 13 October, 2024; originally announced October 2024.

    Comments: Code and the HARDMath dataset is available at https://github.com/sarahmart/HARDMath

  4. arXiv:2406.05200  [pdf

    physics.ins-det nucl-ex

    Decay Energy Spectrometry for Improved Nuclear Material Analysis at the IAEA NML

    Authors: G. B. Kim, A. R. L. Kavner, T. Parsons-Davis, S. Friedrich, O. B. Drury, D. Lee, X. Zhang, N. Hines, S. T. P. Boyd, S. Weidenbenner, K. Schreiber, S. Martinson, C. Smith, D. McNeel, S. Salazar, K. Koehler, M. Carpenter, M. Croce, D. Schmidt, J. Ullom

    Abstract: Decay energy spectrometry (DES) is a novel radiometric technique for high-precision analysis of nuclear materials. DES employs the unique thermal detection physics of cryogenic microcalorimeters with ultra-high energy resolution and 100$\%$ detection efficiency to accomplish high precision decay energy measurements. Low-activity nuclear samples of 1 Bq or less, and without chemical separation, are… ▽ More

    Submitted 11 July, 2024; v1 submitted 7 June, 2024; originally announced June 2024.

    Comments: This was submitted to 2022 IAEA symposium on nuclear safeguards (https://www.iaea.org/events/sg-2022), and posted at https://media.superevent.com/documents/20221027/668fdac0ee8d895ec6bcf293b1c42e6a/id-145.pdf