I've got a working recipe to run this model on Dual DGX Spark: <a href="https://github.com/volfco/spark-vllm-docker/blob/main/recipes/mimo-v2.6-flash.yaml" rel="nofollow">https://github.com/volfco/spark-vllm-docker/blob/main/recipe...
Averages ~25-35tok/s which isn't bad for a first attempt.
volf_ · · focus · HN ↗
Averages ~25-35tok/s which isn't bad for a first attempt.