Its trained at native MXFP4 with MXFP8 activations layer, so you need around 1.5TB of VRAM to fully offload it without taking into account the context cache.
It might be doable to do some smart expert offloading and swapping, but expect minimum 500GB of VRAM and 1TB of system RAM minimum and the t/s would be reduced.
Its only realistically runnable on datacenter grade gpu at decent speed for now (and judging from the price of ram for the next few years)
Its trained at native MXFP4 with MXFP8 activations layer, so you need around 1.5TB of VRAM to fully offload it without taking into account the context cache. It might be doable to do some smart expert offloading and swapping, but expect minimum 500GB of VRAM and 1TB of system RAM minimum and the t/s would be reduced.
Its only realistically runnable on datacenter grade gpu at decent speed for now (and judging from the price of ram for the next few years)