A short counterexample property for safety and liveness verification of fault-tolerant distributed algorithms

DOI10.1145/3009837.3009860MaRDI QIDQ5370906zbMATHOpenAlexFDO

Authors

Publication date 20 October 2017

Published in Proceedings of the 44th ACM SIGPLAN Symposium on Principles of Programming Languages (Search for Journal in Brave)

Full work available at URL https://arxiv.org/abs/1608.05327

zbMATH Keywords

reliable broadcast Byzantine faults fault-tolerant distributed algorithms parameterized model checking

Mathematics Subject Classification ID

Specification and verification (program logics, model checking, etc.) (68Q60) Distributed algorithms (68W15)

Abstract: Distributed algorithms have many mission-critical applications ranging from embedded systems and replicated databases to cloud computing. Due to asynchronous communication, process faults, or network failures, these algorithms are difficult to design and verify. Many algorithms achieve fault tolerance by using threshold guards that, for instance, ensure that a process waits until it has received an acknowledgment from a majority of its peers. Consequently, domain-specific languages for fault-tolerant distributed systems offer language support for threshold guards. We introduce an automated method for model checking of safety and liveness of threshold-guarded distributed algorithms in systems where the number of processes and the fraction of faulty processes are parameters. Our method is based on a short counterexample property: if a distributed algorithm violates a temporal specification (in a fragment of LTL), then there is a counterexample whose length is bounded and independent of the parameters. We prove this property by (i) characterizing executions depending on the structure of the temporal formula, and (ii) using commutativity of transitions to accelerate and shorten executions. We extended the ByMC toolset (Byzantine Model Checker) with our technique, and verified liveness and safety of 10 prominent fault-tolerant distributed algorithms, most of which were out of reach for existing techniques.

Recommendations

Cited in

(21)

Describes a project that uses

Uses Software

This page was built for publication: A short counterexample property for safety and liveness verification of fault-tolerant distributed algorithms

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q5370906)