Discrete Latent Structure in Neural Networks

Niculae, Vlad; Corro, Caio F.; Nangia, Nikita; Mihaylova, Tsvetomila; Martins, André F. T.

Computer Science > Machine Learning

arXiv:2301.07473 (cs)

[Submitted on 18 Jan 2023 (v1), last revised 22 Nov 2024 (this version, v2)]

Title:Discrete Latent Structure in Neural Networks

Authors:Vlad Niculae, Caio F. Corro, Nikita Nangia, Tsvetomila Mihaylova, André F. T. Martins

View PDF

Abstract:Many types of data from fields including natural language processing, computer vision, and bioinformatics, are well represented by discrete, compositional structures such as trees, sequences, or matchings. Latent structure models are a powerful tool for learning to extract such representations, offering a way to incorporate structural bias, discover insight about the data, and interpret decisions. However, effective training is challenging, as neural networks are typically designed for continuous computation.
This text explores three broad strategies for learning with discrete latent structure: continuous relaxation, surrogate gradients, and probabilistic estimation. Our presentation relies on consistent notations for a wide range of models. As such, we reveal many new connections between latent structure learning strategies, showing how most consist of the same small set of fundamental building blocks, but use them differently, leading to substantially different applicability and properties.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
ACM classes:	I.2.6
Cite as:	arXiv:2301.07473 [cs.LG]
	(or arXiv:2301.07473v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2301.07473

Submission history

From: Vlad Niculae [view email]
[v1] Wed, 18 Jan 2023 12:30:44 UTC (10,146 KB)
[v2] Fri, 22 Nov 2024 12:22:39 UTC (10,241 KB)

Computer Science > Machine Learning

Title:Discrete Latent Structure in Neural Networks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Discrete Latent Structure in Neural Networks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators