Skip to main content

Showing 1–1 of 1 results for author: Sapio, A

Searching in archive stat. Search in all archives.
.
  1. arXiv:1903.06701  [pdf, other

    cs.DC cs.LG cs.NI stat.ML

    Scaling Distributed Machine Learning with In-Network Aggregation

    Authors: Amedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson, Panos Kalnis, Changhoon Kim, Arvind Krishnamurthy, Masoud Moshref, Dan R. K. Ports, Peter Richtárik

    Abstract: Training machine learning models in parallel is an increasingly important workload. We accelerate distributed parallel training by designing a communication primitive that uses a programmable switch dataplane to execute a key step of the training process. Our approach, SwitchML, reduces the volume of exchanged data by aggregating the model updates from multiple workers in the network. We co-design… ▽ More

    Submitted 30 September, 2020; v1 submitted 22 February, 2019; originally announced March 2019.