Reading Coded Deep Learning: Framework and Algorithm.
I suspect that the official implementation differs slightly from the paper. Setting the initial alpha of the activation quantizer to 500.0 leads to gradient explosion at the very beginning. The entropy term is not normalized and dominates the training. I cannot reproduce their results.
Distributed under the terms of the LICENSE.
© All rights reserved by FelysNeko