Universal Graph Compression: Stochastic Block Models

Bhatt, Alankrita; Wang, Ziao; Wang, Chi; Wang, Lele

Computer Science > Information Theory

arXiv:2006.02643 (cs)

[Submitted on 4 Jun 2020 (v1), last revised 21 May 2024 (this version, v3)]

Title:Universal Graph Compression: Stochastic Block Models

Authors:Alankrita Bhatt, Ziao Wang, Chi Wang, Lele Wang

View PDF HTML (experimental)

Abstract:Motivated by the prevalent data science applications of processing large-scale graph data such as social networks and biological networks, this paper investigates lossless compression of data in the form of a labeled graph. Particularly, we consider a widely used random graph model, stochastic block model (SBM), which captures the clustering effects in social networks. An information-theoretic universal compression framework is applied, in which one aims to design a single compressor that achieves the asymptotically optimal compression rate, for every SBM distribution, without knowing the parameters of the SBM. Such a graph compressor is proposed in this paper, which universally achieves the optimal compression rate with polynomial time complexity for a wide class of SBMs. Existing universal compression techniques are developed mostly for stationary ergodic one-dimensional sequences. However, the adjacency matrix of SBM has complex two-dimensional correlations. The challenge is alleviated through a carefully designed transform that converts two-dimensional correlated data into almost i.i.d. submatrices. The sequence of submatrices is then compressed by a Krichevsky--Trofimov compressor, whose length analysis is generalized to identically distributed but arbitrarily correlated sequences. In four benchmark graph datasets, the compressed files from competing algorithms take 2.4 to 27 times the space needed by the proposed scheme.

Subjects:	Information Theory (cs.IT); Databases (cs.DB); Statistics Theory (math.ST)
Cite as:	arXiv:2006.02643 [cs.IT]
	(or arXiv:2006.02643v3 [cs.IT] for this version)
	https://doi.org/10.48550/arXiv.2006.02643

Submission history

From: Ziao Wang [view email]
[v1] Thu, 4 Jun 2020 04:51:26 UTC (26 KB)
[v2] Sat, 6 Feb 2021 02:12:38 UTC (98 KB)
[v3] Tue, 21 May 2024 21:51:08 UTC (111 KB)

Computer Science > Information Theory

Title:Universal Graph Compression: Stochastic Block Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Information Theory

Title:Universal Graph Compression: Stochastic Block Models

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators