Fast and scalable all-optical network architecture for distributed deep learning-Reference-Cited by-同舟云学术

Fast and scalable all-optical network architecture for distributed deep learning

Published:2024-02-22 Issue:3 Volume:16 Page:342
ISSN:1943-0620
Container-title:Journal of Optical Communications and Networking
language:en
Short-container-title:J. Opt. Commun. Netw.

Author:

Li Wenzhe¹,Yuan Guojun¹,Wang Zhan¹,Tan Guangming¹,Zhang Peiheng¹,Rouskas George N.²^ORCID

Affiliation:

1. CAS

2. North Carolina State University

Abstract

With the ever-increasing size of training models and datasets, network communication has emerged as a major bottleneck in distributed deep learning training. To address this challenge, we propose an optical distributed deep learning (ODDL) architecture. ODDL utilizes a fast yet scalable all-optical network architecture to accelerate distributed training. One of the key features of the architecture is its flow-based transmit scheduling with fast reconfiguration. This allows ODDL to allocate dedicated optical paths for each traffic stream dynamically, resulting in low network latency and high network utilization. Additionally, ODDL provides physically isolated and tailored network resources for training tasks by reconfiguring the optical switch using LCoS-WSS technology. The ODDL topology also uses tunable transceivers to adapt to time-varying traffic patterns. To achieve accurate and fine-grained scheduling of optical circuits, we propose an efficient distributed control scheme that incurs minimal delay overhead. Our evaluation on real-world traces showcases ODDL’s remarkable performance. When implemented with 1024 nodes and 100 Gbps bandwidth, ODDL accelerates VGG19 training by 1.6× and 1.7× compared to conventional fat-tree electrical networks and photonic SiP-Ring architectures, respectively. We further build a four-node testbed, and our experiments show that ODDL can achieve comparable training time compared to that of an ideal electrical switching network.

Funder

National Key Research and Development Program of China

National Natural Science Foundation of China

Jiangsu Science and Technology Project

National Science Foundation

Publisher

Optica Publishing Group

Reference48 articles.

1. BlueConnect: Decomposing all-reduce for deep learning on heterogeneous network hierarchy

2. Scalable Deep Learning on Distributed Infrastructures

3. PipeDream: generalized pipeline parallelism for DNN training;Narayanan,2019

4. Aluminum: an asynchronous, GPU-aware communication library optimized for large-scale training of deep neural networks on HPC systems;Dryden,2018