ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
arXiv:2602.21140v1 Announce Type: new Abstract: As LLM deployments scale over more hardware, the probability of a single failure in a system increases significantly, and cloud...