Dependable Systems

Reliability from mixed-criticality embedded systems to GPU supercomputers — the thread behind everything else

Goal: computing systems that keep their promises under constraints — the research line where I started, and the lens my edge-AI work inherits.

Supercomputer failure analysis

Multi-year failure and repair logs from the Tsubame supercomputers, characterizing how GPU, node, and interconnect failures evolve across generations of multi-GPU compute nodes and how large systems actually recover. Published at DSN 2021 (Taherin et al., 2021).

Reliability-aware energy management

Energy management for mixed-criticality embedded systems where saving power must never compromise a safety guarantee: exploiting slack and controlled service-level degradation of low-criticality tasks while preserving reliability for critical ones (Taherin et al., 2018) (Taherin et al., 2015).

References

2021

  1. DSN
    Amir Taherin, Tirthak Patel, Giorgis Georgakoudis, and 2 more authors
    In 51st Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN), 2021

2018

  1. TSUSC
    Amir Taherin, Mohammad Salehi, and Alireza Ejlali
    IEEE Transactions on Sustainable Computing, 2018

2015

  1. RTEST
    Amir Taherin, Mohammad Salehi, and Alireza Ejlali
    In CSI Symposium on Real-Time and Embedded Systems and Technologies (RTEST), 2015