espnet2.speechlm.model.speechlm.lm.loss.memory_efficient_load_balancing_loss
Less than 1 minute
espnet2.speechlm.model.speechlm.lm.loss.memory_efficient_load_balancing_loss
espnet2.speechlm.model.speechlm.lm.loss.memory_efficient_load_balancing_loss(gate_logits, num_experts=None, top_k=2)
Compute the router loss without allocating token-by-expert one-hot masks.
