跳到正文
Comfy· Bryan Wade·· 15 小时前精选

如何在单张 GPU 上用 MiniMax H3 生成实时视频

How I Generated Live Video with MiniMax H3 on a Single GPU

AI 导读

作者用 FastVideo 的 FastH3 V2 checkpoint 配合四项采样、VSA 稀疏注意力、ClipProj 小文本编码器、INT8 剪枝 checkpoint 与 FP4 融合 MLP 等优化,在单张 RTX 5090 上把 15 秒 448×256 视频的生成时间压到 15 秒以内,显存需求从 80GB 降到 30GB 以下。

推荐理由

作者公开了在单张 RTX 5090 上实时生成视频的完整优化组合,可迁移到其他视频模型的部署场景。

正文 · 原文

ComfyStreamerH3

Background

Live video has been the talk of the town on AI twitter these days. From Fal’s post trained H3 Max model that competed on quality and topped the charts last month, to the FastH3 model that optimized on cost. Video models have come a long way since people joked about AI models being unable to animate Will Smith Eating Spaghetti!

As impressive as these recent model are, there exists a problem. And that problem is cost. Running these models continuously can cost an upwards of 6k per day! This makes these models completely unusable for the average consumer.

Thanks for reading ComfyUI Newsletter! Subscribe for free to receive new posts and support my work.

The goal

So the question I asked is how far can these models be pushed? What if you had an agent search the internet for every optimization, every speed up, every cost savings, and definitely every dirty hack that sacrificed quality for speed. And then what if you put it all together into a single custom comfy node that runs on a single consumer GPU? My requirements were the following:

  • Hardware: one RTX 5090 with 32 GB of VRAM

  • Speed: generate 15 seconds of video in 15 seconds or less

  • Resolution: 448×256 resolution

  • Quality: keep it watchable, even if the clips have rough edges

The RTX 5090 was a useful target: some high-end gaming PCs have one, and rentals can be found for around 70¢ an hour at the cheapest end of the market.

Benchmarks

Examples: Same spaghetti. Three styles.

Anime

Stylized 3D

Watercolor

The clips clearly have some rough edges and artifacts. But its not a completely horrible result for a single GPU you can rent for under 70 cents an hour!

But enough with the talk, lets get into the nitty gritty details. Casual viewers be warned, I don’t even understand some of these, as they were simply what my agent was able to pull from other projects. Astra Codex be praised!

Implementation Details

I used FastVideo’s FastH3 V2 checkpoint in a four-step setup. For comparison, FastVideo’s Preview V1 generated a 15-second, 1344×768 clip in 47.2 seconds on one B200 (hint: much more expensive GPU).

Here’s what we optimized to run the workflow on one RTX 5090:

  • Four sampling steps. Each step is another round of model work to refine the video. Using four means less work for each clip.

  • Sparse attention. Attention lets parts of the video share information as it is generated. VSA focuses full attention on the selected 20% of video blocks—skipping about 80%—while still including the prompt.

  • A smaller text encoder. The encoder turns your prompt into information H3 can use. NicoLab28’s ClipProj adapts Qwen3-VL-4B to stand in for H3’s 32B encoder: it uses about 4.5 GB instead of 15.7 GB.

  • A smaller checkpoint. FastVideo’s pruned INT8 checkpoint combines pruning (removing selected weights) with INT8 storage (8-bit numbers), reducing the GPU memory the model needs.

  • A fused FP4 MLP. The MLP handles a lot of the model’s repeated calculations. Using four-bit math for this part and fusing its operations means less memory for the weights and fewer temporary results to move around; this builds on ByronLeeeee’s H3 optimizations.

  • Decode and encode together. Kijai’s INT8 H3 video decoder rebuilds the video in tiles. NVENC starts writing each finished tile while the decoder prepares the next, so the steps overlap. The workflow keeps only the last frame for the next job instead of caching every decoded frame.

These optimizations both improved speed and memory uses, bringing VRAM requirements down from 80GB to < 30GB, safely fitting within a RTX 5090’s 32 GB of GPU memory.

Demo

If you want to test out the results for yourself the custom comfy node is released open source here: https://github.com/Comfy-Org/comfystreamerh3

Additionally you can rent GPUs from the Comfy Developer Platform.

Comfy API handles GPU hosting, models, and dependencies, so your app can request video without needing a GPU of your own. And you can follow the instructions to sign up and deploy this demo to Comfy API in the readme here.

Final Thoughts

There is clearly significant room to improve image detail, consistency, and resolution of these videos. But getting H3 to generate even partially legible video in real time on one RTX 5090 shows how far these models can be pushed when it comes to the frontier of single-GPU live video. One next step would be to pass each clip through a cost optimized video upscaling model, such as the recently release SolRefinery : IE generate quick initial videos at 448 × 256 and then cheaply upscale to a higher resolution. My preliminary testing with this model shows a 7× pixel density increase using only 3X the GPUs. So given the possibility of future optimizations, we might have cost efficient, semi-usable live video sooner than you’d think!

Upscaled Results

Thanks for reading ComfyUI Newsletter! Subscribe for free to receive new posts and support my work.

来源:Comfy · blog.comfy.org