A Comparison of Two Smoothing Methods for Word Bigram Models

Peto, Linda Bauman

Computation and Language

arXiv:cmp-lg/9410034 (cmp-lg)

[Submitted on 31 Oct 1994]

Title:A Comparison of Two Smoothing Methods for Word Bigram Models

Authors:Linda Bauman Peto (University of Toronto)

View PDF

Abstract: A COMPARISON OF TWO SMOOTHING METHODS FOR WORD BIGRAM MODELS
Linda Bauman Peto
Department of Computer Science
University of Toronto Abstract Word bigram models estimated from text corpora require smoothing methods to estimate the probabilities of unseen bigrams. The deleted estimation method uses the formula:
Pr(i|j) = lambda f_i + (1-lambda)f_i|j, where f_i and f_i|j are the relative frequency of i and the conditional relative frequency of i given j, respectively, and lambda is an optimized parameter. MacKay (1994) proposes a Bayesian approach using Dirichlet priors, which yields a different formula:
Pr(i|j) = (alpha/F_j + alpha) m_i + (1 - alpha/F_j + alpha) f_i|j where F_j is the count of j and alpha and m_i are optimized parameters. This thesis describes an experiment in which the two methods were trained on a two-million-word corpus taken from the Canadian _Hansard_ and compared on the basis of the experimental perplexity that they assigned to a shared test corpus. The methods proved to be about equally accurate, with MacKay's method using fewer resources.

Comments:	this http URL. thesis, 57 pages, compressed and uucencoded postscript file
Subjects:	Computation and Language (cs.CL)
Report number:	CSRI-304
Cite as:	arXiv:cmp-lg/9410034
	(or arXiv:cmp-lg/9410034v1 for this version)
	https://doi.org/10.48550/arXiv.cmp-lg/9410034

Submission history

From: Linda Bauman Peto [view email]
[v1] Mon, 31 Oct 1994 19:55:57 UTC (100 KB)

Computation and Language

Title:A Comparison of Two Smoothing Methods for Word Bigram Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computation and Language

Title:A Comparison of Two Smoothing Methods for Word Bigram Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators