On the class of coding optimality of human languages and the origins of Zipf's law

Ferrer-i-Cancho, Ramon

Computer Science > Computation and Language

arXiv:2505.20015 (cs)

[Submitted on 26 May 2025 (v1), last revised 2 Sep 2025 (this version, v5)]

Title:On the class of coding optimality of human languages and the origins of Zipf's law

Authors:Ramon Ferrer-i-Cancho

View PDF HTML (experimental)

Abstract:Here we present a new class of optimality for coding systems. Members of that class are displaced linearly from optimal coding and thus exhibit Zipf's law, namely a power-law distribution of frequency ranks. Within that class, Zipf's law, the size-rank law and the size-probability law form a group-like structure. We identify human languages that are members of the class. All languages showing sufficient agreement with Zipf's law are potential members of the class. In contrast, there are communication systems in other species that cannot be members of that class for exhibiting an exponential distribution instead but dolphins and humpback whales might. We provide a new insight into plots of frequency versus rank in double logarithmic scale. For any system, a straight line in that scale indicates that the lengths of optimal codes under non-singular coding and under uniquely decodable encoding are displaced by a linear function whose slope is the exponent of Zipf's law. For systems under compression and constrained to be uniquely decodable, such a straight line may indicate that the system is coding close to optimality. We provide support for the hypothesis that Zipf's law originates from compression and define testable conditions for the emergence of Zipf's law in compressing systems.

Comments:	a few typos corrected, in press in Europhysics Letters
Subjects:	Computation and Language (cs.CL); Physics and Society (physics.soc-ph)
Cite as:	arXiv:2505.20015 [cs.CL]
	(or arXiv:2505.20015v5 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2505.20015

Submission history

From: Ramon Ferrer-i-Cancho [view email]
[v1] Mon, 26 May 2025 14:05:45 UTC (22 KB)
[v2] Tue, 3 Jun 2025 17:00:20 UTC (23 KB)
[v3] Wed, 4 Jun 2025 11:35:43 UTC (23 KB)
[v4] Fri, 18 Jul 2025 14:57:19 UTC (24 KB)
[v5] Tue, 2 Sep 2025 18:22:53 UTC (24 KB)

Computer Science > Computation and Language

Title:On the class of coding optimality of human languages and the origins of Zipf's law

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:On the class of coding optimality of human languages and the origins of Zipf's law

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators