Federated Learning enables collaborative machine learning in distributed data environments without the consolidation of raw data. To avoid the communication bottleneck and the single point of failure found in Centralized Federated Learning, Decentralized Federated Learning (DFL) eliminates the coordinating central server and distributes its responsibilities among the participants. However, the challenge of communication efficiency persists in DFL, resulting in a trade-off between communication overhead and convergence rate. Despite efforts already invested in improving communication efficiency, the existing literature lacks a comprehensive overview of communication aspects in DFL. To fill this gap, this work systemizes recent DFL publications with a particular focus on communication cost of core constituents in horizontal DFL. We establish a taxonomy of essential DFL components that influence communication by identifying classes of strategies and linking them to implementations in the current literature. Next, we examine popular techniques to optimize the communication cost in DFL and categorize these methods into sub-classes. To demonstrate the applicability of our taxonomy, we then compare the methods of recent DFL publications by taking these components and optimization techniques into account. This paper provides an examination and categorization of key communication aspects in DFL, valuable for both researchers and practitioners engaged in resource-constrained distributed optimization.
@article{WEISSTEINER2026108666,title={Communication efficiency in Decentralized Federated Learning: A taxonomy of core components and optimization techniques},journal={Computer Communications},volume={259},pages={108666},year={2026},issn={0140-3664},doi={10.1016/j.comcom.2026.108666},url={https://doi.org/10.1016/j.comcom.2026.108666},author={Weissteiner, David and Demelius, Lea and Trügler, Andreas},keywords={Decentralized federated learning (DFL), Communication efficiency, Resource-constrained distributed optimization, Communication optimization},}
2025
Decentralized Federated Learning: Framework Design, Communication Efficiency, and Dynamic Synchronization
Federated Learning enables collaborative machine learning in distributed data environments without consolidation of raw data. In Centralized Federated Learning, devices share updates to the model parameters with a central server that aggregates them and redistributes the aggregated state. However, the central server poses a single point of failure and a major communication bottleneck. Decentralized Federated Learning (DFL) overcomes these problems by eliminating the central role and distributing the responsibilities to individual devices (actors). In this decentralized setting, communication is no longer centered on one point, allowing actors to communicate arbitrarily with their peers. The amplified possibility of network links and potential device limitations, however, raises the challenge of communication efficiency. In particular in resource-constrained environments, heavy communication workloads slow down the entire system, sometimes even rendering the application infeasible due to bandwidth limitations. In this thesis, we approach the challenge of communication efficiency in DFL. We start by outlining the fundamentals of DFL to clarify the general paradigm and identify the key challenges. Finding that a DFL system consists of numerous components, we recognize the difficulty of comparing different methods of the individual components with each other. Therefore, we implemented MoDeFL, a Modular Decentralized Federated Learning framework to facilitate the evaluation of DFL methods and to foster their comparability. MoDeFL aims to enable complete configuration and independent interchangeability of individual component methods. The implementation of a multitude of communication optimization methods and the support of communication-relevant evaluation metrics in MoDeFL has sparked our interest in methods for communication efficiency. In response, we survey the DFL literature with a particular focus on communication efficiency. Beyond establishing a taxonomy for the various communication-efficient techniques in DFL, this survey reveals the topic of dynamic synchronization among the underexplored research directions in DFL. In light of this research gap, we design Gradient Thresholding, a novel synchronization algorithm to reduce communication cost while preventing actors from diverging. This algorithm is able to detect divergent actors without the need for additional communication and triggers the synchronization process accordingly. Thus, Gradient Thresholding reduces communication cost by omitting exchanges of model updates between actors. Our experiments demonstrate the superiority of Gradient Thresholding over baselines and further illustrate its concept.
@mastersthesis{Weissteiner2025DecenFL,author={Weissteiner, David},title={Decentralized Federated Learning: Framework Design, Communication Efficiency, and Dynamic Synchronization},school={Graz University of Technology},year={2025},type={Master's Thesis},address={Rechbauerstraße 12, 8010 Graz, AUSTRIA},month=dec,doi={10.3217/xzyye-svh31},url={https://doi.org/10.3217/xzyye-svh31},note={Supervised by Andreas Tr{\"u}gler},}
2022
Federated Data Preparation, Learning, and Debugging in Apache SystemDS
Sebastian
Baunsgaard
, Matthias
Boehm
, Kevin
Innerebner
, and
7 more authors
In Proceedings of the 31st ACM International Conference on Information & Knowledge Management , Dec 2022
Federated learning allows training machine learning (ML) models without central consolidation of the raw data. Variants of such federated learning systems enable privacy-preserving ML, and address data ownership and/or sharing constraints. However, existing work mostly adopt data-parallel parameter-server architectures for mini-batch training, require manual construction of federated runtime plans, and largely ignore the broad variety of data preparation, ML algorithms, and model debugging. Over the last years, we extended Apache SystemDS by an additional federated runtime backend for federated linear-algebra programs, federated parameter servers, and federated data preparation. In this paper, we share the system-level compiler and runtime integration, new features such as multi-tenant federated learning, selected federated primitives, multi-key homomorphic encryption, and our monitoring infrastructure. Our demonstrator showcases how composite ML pipelines can be compiled into federated runtime plans with low overhead.
@inproceedings{10.1145/3511808.3557162,author={Baunsgaard, Sebastian and Boehm, Matthias and Innerebner, Kevin and Kehayov, Mito and Lackner, Florian and Ovcharenko, Olga and Phani, Arnab and Rieger, Tobias and Weissteiner, David and Wrede, Sebastian Benjamin},title={Federated Data Preparation, Learning, and Debugging in Apache SystemDS},year={2022},isbn={9781450392365},publisher={Association for Computing Machinery},address={New York, NY, USA},url={https://doi.org/10.1145/3511808.3557162},doi={10.1145/3511808.3557162},booktitle={Proceedings of the 31st ACM International Conference on Information \& Knowledge Management},pages={4813–4817},numpages={5},keywords={monitoring, federated raw data, federated learning},location={Atlanta, GA, USA},series={CIKM '22},}
Federated Learning (FL) is being increasingly studied and became a common application in many domains. Existing work about FL mostly focuses on a system architecture with a single centralized actor (coordinator) deploying continuous computation requests over multiple decentralized federated workers. However, due to the limitation of a single coordinator within the federated infrastructure, only one actor can address the data of the federated worker at the same time.
The emerging field of FL created many application scenarios allowing to learn centralized models from private data, and some systems even making it feasible to perform certain ML pipelines. However, these applications miss the benefits of multiple coordinators within the same federated infrastructure.
In this work, we introduce Multi-tenant Federated Learning (MuTeFL), which provides the possibility to share data and computing resources of multiple federated workers among several distinct coordinators. To preserve the independence of the coordinators, we create an autonomous, server-like federated worker that can serve multiple coordinators simultaneously. Moreover, we isolate the coordinator-instances at the federated worker to avoid revealing data to other coordinators, we introduce parallelization strategies to utilize the full potential of the worker, and we eliminate redundancies at the worker by using reuse techniques.
We thereby create the possibility to deploy the federated worker as a long-running server process that can be addressed by different parties at any time. This enables, for example, simultaneous model training by different data scientists on the same federated workers, as well as providing learning capabilities on sensitive data to the public domain.
@thesis{Weissteiner2022MuTeFL,author={Weissteiner, David},title={Multi-tenant Federated Learning},school={Graz University of Technology},year={2022},type={Bachelor's Thesis},address={Rechbauerstraße 12, 8010 Graz, AUSTRIA},month=aug,url={https://ywcb00.ywcb.org/assets/pdf/MuTeFL.pdf},note={Supervised by Matthias Boehm and Arnab Phani},}