News
Xiaomi releases MiMo-V2.5-DFlash weights on Hugging Face
Xiaomi uploaded MiMo-V2.5-DFlash, a 300B-parameter model featuring DFlash optimization for faster inference. The release includes dedicated DFlash weights and a separate MTP model. Early reports suggest DFlash could potentially double inference speed on consumer hardware with vram offloading.
Xiaomi has released MiMo-V2.5-DFlash on Hugging Face, a 300-billion-parameter language model featuring DFlash optimization for enhanced inference speed. The release includes dedicated DFlash weights stored in a separate directory alongside the standard model files, allowing developers to test the optimized variant.
Based on initial user reports, the base MiMo-V2.5 model achieves 8–10 tokens per second on dual 24GB graphics cards with 96–128GB DDR5 RAM offloading. The DFlash optimization is expected to roughly double this throughput, making the model substantially more practical for consumer-grade hardware deployments.
The release also includes a separate MTP (Multi-Token Prediction) model, addressing prior implementation challenges. While previous shared MTP heads did not function properly in llama.cpp due to difficulties identifying MTP layers, the standalone MTP model may offer better compatibility with existing inference frameworks. DFlash itself is expected to work within current tools, providing an immediate performance boost for users seeking faster inference on local systems.
Readers’ Forum
No contributions yet — open the debate.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Waiting for your click …
·