Implement the reduction at the heart of distributed data-parallel training: average one gradient vector from every worker.
Required APIdef average_gradients(gradients):
return averaged_vectorBehaviorThink through the mechanism first if you want the extra reasoning step. It never blocks the editor.
average_gradients([[1.0, 3.0], [3.0, 5.0]])synchronized workers apply the same average update
average_gradients([[2.0, -1.0]])one worker is the identity reduction
2 hidden edge tests run after the visible contract passes.
Run Tests to see the contract verdicts here.