|
|
马上注册,结交更多好友,享用更多功能,让你轻松玩转社区。
您需要 登录 才可以下载或查看,没有帐号?免费注册
x
本帖最后由 明心亘青 于 2026-8-26 18:32 编辑
最近正在试着入坑无畏契约,看完教程觉得有收获而且还不嫌弃带唐比的可以考虑带带我(?)
推荐与站内其他用户的教程共同使用,你只看我的也可以,本文也不是必须看的文章。
注意:
1、文章会与站内其他教程在部分观点以及做法上存在一定程度的偏差,但没有绝对错误,你认为谁是对的取决于你在自己的实践中积累的判断。
2、本文会夹杂部分甚至大量常见英文名词,这是笔者习惯所致。
3、本文完全手写,不掺杂几乎任何 AI 生成的文字,不可避免地会出现错字、语句不通顺等问题。同时掺杂了一点 Markdown 语法,看看论坛能不能渲染,不能的话人脑渲染吧(
4、本文在快要写完的时候因为手误点到了别的地方导致页面被刷新了一次,导致几乎 70% 内容都要重写一遍,有点红,所以一些地方会写得较为粗略,甚至不会校对名词的大小写问题。在 Word 里写一遍再粘贴过来是对的。什么时候给编辑帖子页面加个切换网点或者刷新时提示有未保存的内容提醒……
Minimax H3 官方开源的版本便没有审核一说,但相关知识库方面有所欠缺。正好就在近日,Wan 2.2 Remix 的作者与 Wan 2.2 Dasiwa 的作者分别发布了 Minimax H3 Remix(FL2VA)和 Minimax H3 Dasiwa(Ref2VA),感兴趣的读者可自行下载测试。笔者只测试了 Remix 的 FL2VA,NSFW 方面相比有所加强。出于教程考量,文章在讲述 FL2VA 时仍然会使用官方模型。
(你问为什么只在讲 FL2VA 时用?因为我硬盘不够了哈哈)
就在笔者截图时,笔者忽然发现 Dasiwa 的作者又将其模型删除了,原因未知。亟待其后续更新吧。
另外 Redcraft 即红潮也有其微调,但笔者觉得很难用。长篇大论一大堆,写了特大量技术名词,似乎是想彰显自身技术水平。但实际使用效果却极差。
本文更加侧重于「使用方法」而不是「使用示例」。当然,这不意味着本文就不会有任何实操范例,不仅有,而且还有二次元图像实操,这也是本人认为 H3 相较于以往最大的优势的体现区域,知识库的庞大甚至使得 H3 能够理解「Vtuber」这种概念,其运用的 Live2D 骨骼动画更是不在话下,音频也比 MM Audio 生成的效果好一大截。不夸张地说,H3 几乎是过往所有开源视频生成模型(LTX2.5 在 H3 之后,听别人说相比 H3 除了速度快没有优势,Bernini 是视频编辑模型,且听说是云平台都能爆显存的抽象玩意,这两个模型我没有测试,所以此处用于是「几乎」)的平替甚至上位替代。当然,存储空间也是。在实操前,请确保你模型文件存储的路径至少有 55G 的富余空间可供使用。
(二次元 AI 动图终于摆脱伪人感了有没有懂的)
得益于 MoE——MoE 意为 Mixture of Expert,即混合专家模型,在调度资源时不会完整地激活模型所有参数,而是小部分地激活,这将使得配置不够高的设备也能运行参数量较大的模型——的资源调度架构,本人配置为 5070Ti Laptop 12G+32G DDR5+50G 虚拟内存,使用的模型为 int8 的量化版本。低于这个参数的应该也可以跑,高于的一定可以跑,实测中 FL2VA 1MP 5s 可以同时进行 2 倍超分+RIFE 补帧,若时间延长到 10 秒则只支持到 0.6MP。
文前解释一些后文可能会用到的术语,后文不会再解释。基本上是自己的理解,如有纰漏欢迎指正:
FL2VA:First last Frame to video and audio,意为「首尾帧生视频音频」。
注:Minimax H3 相较于 Wan 2.2,有一个特点便是生成视频时也同时支持音频的生成,也因此,在官方文档中,相比过去常说的 I2V,其缩写会额外增加一个 A(Audio)。
Ref2VA:Reference to video and audio,意为「参考生视频」。
注:此处不说「参考图」是因为 H3 的性能强悍,除了支持最大 9 张图像的参考以外还支持音频参考与视频参考,通过 REF2VA 的参考视频功能,你可以实现类似动作迁移等功能。但视频参考极其吃配置(笔者的电脑只够生成),且根据笔者的道听途说来看 H3 一段视频只能参考最多 81 帧,但具体是通过抽帧还是截断实现这 81 帧的参考就不得而知了。
MP:Million Pixel,百万像素。MP 具体代指的分辨率在官方工作流中有,此处暂且按下不表。
注:REF2VA 画质很差。根据笔者的道听途说,官方承认 REF2VA 的 1MP≈FL2VA 的 0.4MP。不过笔者暂时没有找到官方承认该说法的出处,但 REF2VA 画质不如 FL2VA 在笔者的体感中是存在的。
LLM:Large Language Model,大语言模型。常说的 DeepSeek,Gemini、ChatGPT 以及 Claude 都属于这一类。
这里补充一个常见的谬误,所谓的多模态功能并不包含生图,而是支持多种类型的文件输入,输出依旧是文字。不管是 Google 的 Nano Banana 还是 ChatGPT 的 GPT-Image 2 亦或是 Grok 的 Grok Imagine 2 都是独立的生图模型,只是其语义理解是这些 LLM 而已。说这些生图模型是 LLM 的一部分,就像说指着 Seedance 2 这是豆包一样。
Agent:常译智能体,笔者更喜欢使用 Agent 这一原文称呼。作用是调用 LLM,支持 MCP、Skills 等功能。后文中并不会涉及 MCP。
Skills:即技能。顾名思义,能让 Agent 中的 LLM 针对某项能力进行学习。
说到这儿,想必已经有人猜出正文会干什么了。我们的确会使用 Agent 调用 LLM 书写 Prompt。碍于配置问题,笔者无法使用 Qwen 3.8 27B 哪怕是 iQ3 的量化,为此只能给在线 API 爆金币了。至于使用什么 Agent 笔者不会做任何限定,你可以使用自己觉得顺手的。在文中,笔者将会使用 ZCode。至于 LLM 也不会做限制,只要你会破甲即可。文中将会使用 DeepSeek。
---
打开 Modelscope, 搜索 Minimax H3。不要点左边的 Minimax 发布的仓库,那是开源的模型权重而非我们需要的可在 ComfyUI 中使用的模型文件。
点进模型仓库后,选择我们需要的模型。
前面已经讲过 fl2va、ref2va 的区别了,这里就说说这些bf16什么的是什么意思。
bf16 通常被认为是一个模型比较完整的状态,在这种情况下的模型可被视为完全体类似的也有 fp16 和 fp32。而 int8 和 fp8 都是通过量化算法对其体积大小进行压缩。其中 int8 基于整数,fp8基于浮点运算。通常来说,在运行环境为 CU130 及以上的设备上,int8 的速度表现要优于 fp8。bf16 由于过于完整,存储占用来到了恐怖的 66.28G,一般人的电脑吃不消,所以我们下载 int8 即可。如果你的运行环境不够先进,那么 fp8 也是可以的。
至于是否要选择 Pruned,我的建议是要,这也是官方的选择,也就是 int8 pruned 版本。Minimax H3 的参数量是 33.1B,在 Comfy-Org 的 Pruned——剪枝,即「剪」去模型中无用的参数,优化体积占用——下为 20.1B。Pruned 模型减去的是 Time Steps 的冗余计算方式,从原理上来讲并不会对画面质量产生肉眼可见的影响。在 BF16 的完整参数下已有视频对 Pruned 与原生模型进行生成对比,B 站 BV 号为 BV17a8z6SE3H,感兴趣的读者可自行观看。
来到 loras 文件夹,这是加速用的 Turbo lora。第二、三个 Lora 来自同一团队,与第一个 Lora 的团队不同。Fl2va 从一二中任选一个即可。官方与笔者均采用第二个 Lora。
注:Turbo lora 刚问世时笔者便进行过测试,得出的结论是可以用,但是 4-Steps 太少。即便加载的是 4-Steps lora,也仍然需要将步数调整至 6~8 步才可使用。在实际使用中,若出现噪点爆音等问题,将步数从 4 步提升至 6~8 步说不定可以解决。
来到 text encoders 文件夹,官方选择的是 nvfp4 的版本,笔者则选择了 int8.
Nvfp4 是在原本 fp8 基础上进一步量化压缩的版本,主要在 Blackwell 架构(50 系显卡)上进行了优化。非 50 系显卡也可以使用。
另外不要使用 Huggingface 上所谓的 uncensored 的版本。这三个 te 本质上都是无甲的,uncensored 不仅多此一举,其作者本人在 Reddit 上更是提到该版本会产生性能下降的问题。
VAE 没什么好讲的,两个都下了就行。
模型文件存放目录不做赘述,参考该图像即可。
另外如果你进错仓库到了 Minimax 官方的版本,倒也不用急,可以顺便看看这两个文件,它们会教你如何书写 Prompt。
来到 ComfyUI。首先请确保你的 ComfyUI 版本在 0.30.0 及以上。可以说这个版本就是为了 Minimax H3 进行更新的。而 Turbo Lora 作者在 Huggingface 的 Modelcard 页面提到建议将版本更新至 0.31.0 及以上。拿不准主意的话直接更新到最新版即可。
进入 ComfyUI,点击左上角,进入浏览模板页面。
这里就是官方工作流。我们先使用 I2VA。
不出意外的话,你见到的页面应该是这样的。将模型替换为你放入模型目录的文件即可。
如图所示。此处我们选择开启 Turbo 模式。不开启 Turbo 模式的话,耗时最起码是 Turbo 模式的一倍以上。
Duration 滑块就是视频秒数。Minimax H3 固定 24 帧,不可更改。
官方工作流中有 MP 与实际分辨率的换算表格,参考即可。
另外强烈建议按照这个进行修改连接方式。
报错的话就改成这样。
这样一来有利于统一视频与图像的分辨率比例。对于 832*1216 这样的常用分辨率,4:3 16:9 这种常用的比例是无法描述的。而强制使用这些分辨率去生成视频的话会使得成品变得十分怪异,模型会强制裁剪、拉伸原图以匹配视频分辨率。
随便塞张图进去:
- Prompt:采用中焦镜头从正侧面平视略低角度拍摄,镜头横在她与灶台之间,她搭在灶台上的腿、前俯的腰背与身后男人的半身轮廓在同一平面内层层叠进,浅景深把厨房背景柔化,画面充满家常烟火气里偷欢的紧张张力。厨房暖白操作灯从头顶洒下,侧光沿她裸露的背脊、翘起的臀线与他扶在她腰侧的前臂勾勒出清晰轮廓,明暗对比分明,整体暖橙低饱和,氛围普通而暧昧。甜美可爱的20岁年轻成年女性,东亚美少女,白皙肌肤,娇小苗条,第一印象是乖巧里藏着放逸。黑色双马尾垂在肩侧,随她前后晃动的节奏轻轻摆荡,几缕碎发贴在泛红的耳边。妆容为清透淡妆,粉唇微微泛亮,眼下薄薄腮红。她低头凝视着锅里的食物,眼神专注却掩不住失神,嘴唇轻抿压下声音,脸颊绯红,耳根烧得通红,只有睫毛微微颤动出卖了她。She wears nothing but a white bib apron tied around her neck and waist, her back and buttocks completely bare behind the tied strings, leaving her fully exposed from behind, a pair of thin black ankle socks on her feet. She stands at the kitchen counter with one bare foot propped up on the countertop, leaning her upper body forward over the stove as she stirs a pan with one hand, while behind her a man's hands grip her hips and thrust into her from behind with slow deep strokes, his chest and abdomen pressing against her arched back but his face staying out of frame, her hips rocking against her stirring hand. 现代化开放式厨房,灶台上一口锅冒着蒸汽,旁边菜板与调料瓶整齐,头顶射灯暖光,橱柜与瓷砖在焦外虚化,窗外夜色沉静。构图采用对角线,她搭腿、前俯的腰线与身后男人的躯干共同构成一条斜向主线,顶灯暖光沿灶台把视线引向她翘起的臀线与交合处,镜头始终落在焦点处的她身上。使用模型:Moody Krea 2 Mix V7 NVFP4
复制代码
- <div>Prompt:</div><div>For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.</div><div>
- </div><div>integrated_multimodal_description: [Shot 1] Live-action, photoreal, cinematic lighting, a fixed medium-wide shot beginning exactly on the first frame of <Picture 1>: a young East Asian woman with long dark-brown hair and bangs sits on the kitchen counter with her legs spread and one bare foot resting on the lower cabinet, wearing a white apron tied loosely around her waist over nothing, her hands holding a spatula over a pan on the stove. Behind her, a tall East Asian man with short dark hair and visible upper body stands close, gripping her hips and buttocks with both hands while a thick penis is already deeply buried inside her vagina, her labia stretched tightly around the shaft, her body glistening with oil and sweat. The camera remains completely static. From that locked starting position the man begins thrusting with large amplitude: he pulls back until only the glans remain inside her, then drives forward with full force, burying the entire penis to the hilt again and again, his hips slapping heavily against her ass on every deep insertion, her body rocking forward and back on the counter, her breasts bouncing, her long hair swaying, and her face showing natural reactions—eyes half-closed in pleasure, mouth slightly open, lips parted with soft moans escaping as her thighs tremble and her hands grip the spatula tighter.</div><div>
- </div><div>overall_soundscape: Continuous sizzling from the frying pan, steam hissing, and the rhythmic wet slapping of skin on skin from the deep thrusts, with the woman's soft involuntary moans and gasps growing louder and more frequent on each deep insertion.</div><div>
- </div><div>non_diegetic_music: N/A</div>
复制代码
可以看到效果非常好。但真正让读者觉得核弹爆炸,体现其知识库带来的力大砖飞效果的,笔者认为还是二次元图像。
这就是 H3 庞大的知识库带来的好处,尽管作为 Live2D 仍然不是很完美,但一致性、动态效果方面已经远超过去的模型了。
它甚至在这里做了碰撞的物理效果。
其实官方工作流没什么好讲的,因为官方工作流基本上不会使用第三方节点,导致你并不需要配置什么其他的东西就能直接填提示词塞图片以后点击运行一键生成。可能要做的也就这样多加几个节点,以便超分和补帧。后面其他的工作流也会这样,不再赘述。
放大使用 SeedVR2 也是可以的,电脑吃得消就行。
接下来看看 Ref2VA。
提示词输入在左下角。
我们可以试着随便塞两张三视图进去。
效果:
- <div>Prompt:</div><div>subject_definitions: <Subject 1> is the girl shown in <Picture 1>, with long silver-white hair, red eyes, two gold star hairpins on the left side, and a white sailor-style school uniform with a red collar and red neck tie, a red pleated skirt, a black belt, black opaque tights, and black loafers. <Subject 2> is the girl shown in <Picture 2>, with long turquoise twin-tails, blue eyes, a small navy beret with a red bow, a white short-sleeve sailor uniform with a red collar scarf, a navy pleated skirt, dark-navy over-knee socks, and black loafers.</div><div>
- </div><div>summary: [reference generation] The target video shows <Subject 1>, the silver-haired girl from <Picture 1>, pushing <Subject 2>, the turquoise twin-tail girl from <Picture 2>, down onto a bed and holding her there for the full five-second duration.</div><div>
- </div><div>retention_analysis: <Subject 1> (appears in [Shot 1]): fully_preserved - the long silver-white hair, red eyes, gold star hairpins, white sailor uniform with red collar and tie, red pleated skirt, black belt, black tights, and black loafers are retained. <Subject 2> (appears in [Shot 1]): fully_preserved - the long turquoise twin-tails, blue eyes, navy beret with red bow, white short-sleeve sailor uniform with red collar scarf, navy pleated skirt, over-knee socks, and black loafers are retained.</div><div>
- </div><div>detailed_description: The target video is in a clean 2D-anime style with soft indoor lighting and a warm, slightly saturated color palette, matching the two reference sheets. [Shot 1] A medium shot opens in a bright bedroom, a bed with light-colored sheets in the background, as <Subject 1>, the silver-haired girl with gold star hairpins in her white sailor uniform with a red collar and red pleated skirt, steps forward and grips <Subject 2>'s shoulders with both hands. <Subject 2>, the turquoise twin-tail girl in the white short-sleeve sailor uniform and navy skirt, stumbles backward with a start, her eyes widening and her mouth falling open. The camera pushes in with small amplitude at slow speed, following <Subject 1> as she thrusts <Subject 2> down onto the bed; <Subject 2>'s twin-tails fan out across the mattress and her skirt flares as she lands on her back. <Subject 1> climbs over and leans down close, pinning both of <Subject 2>'s wrists against the sheets beside her head, her silver hair falling forward as she holds <Subject 2> down while both girls breathe quickly.</div><div>
- </div><div>overall_soundscape: Quick footsteps cross the floor, the soft rustle of fabric, and a muffled thump as a body lands on the mattress, followed by the fast, shallow breathing of both girls and a short surprised gasp from <Subject 2>.</div><div>
- </div><div>non_diegetic_music: N/A</div>
复制代码
不难发现,H3 的 Ref2VA 其实就是 Seedance 那样的「图片 1……图片 2……」在参考图像合理以及提示词的作用下,理论上可以用于 AI 短剧的制作。
碍于配置原因,笔者的电脑只够支持笔者跑 0.2MP 的视频参考,甚至才跑了 5 秒,再往上就会爆显存。笔者在此不会演示视频替换的效果。
另外 H3 在面对高动态场景时会略显乏力,加上 Ref2VA 的画质较差,所以会有些画面糊得不像样子。
~~这他妈是啥~~
进入 T2VA 的工作流,我们会发现这相当眼熟:
是的,这只是删掉了图像输入的 I2VA 工作流。
- <div>Prompt:</div><div>integrated_multimodal_description: [Shot 1] 2D-animated Live2D style, a front medium shot frames a cute VTuber girl centered on screen in a cozy pastel streaming room decorated with soft neon lights and plush toys; she has long petal-pink hair tied into twin-tails with tiny star hairpins, large sparkling violet eyes, and a white-and-pink idol-style dress with frills. The camera holds a static shot as she leans forward slightly, blinking and smiling with an eager expression, her twin-tails and dress swaying gently with her subtle idle motion. She brings one hand up to her chest and says in a bright, clear, energetic young female voice (S1): `<d>[Chinese] 大家好!我是新人Vtuber小柚子,请多多关照!</d>` Her lips move in sync with the words, a soft blush spreads across her cheeks, and she finishes with a cheerful little bow and a wave, her rig animating naturally throughout.</div><div>
- </div><div>overall_soundscape: Soft room ambience and gentle keyboard clicks continue underneath, joined by the light rustle of fabric as she bows and the faint tap of her hand against her chest.</div><div>
- </div><div>non_diegetic_music: N/A</div>
复制代码
T2VA 是最能体现模型知识库的地方,可以看到它甚至认识 Vtuber 这个概念。所以我不建议使用 H3 进行 NSFW 的 T2VA 创作。哪怕是微调。
- <div>Prompt:</div><div>integrated_multimodal_description: [Shot 1] Live-action, realistic, cinematic, a medium shot frames a young East Asian woman in her early twenties with fair, smooth skin, a slim figure, long silver-lavender straight hair tied with a red ribbon, bright blue eyes, and soft makeup with slightly parted lips. She stands in a crowded convention hall with posters, panel walls, and blurred attendees filling the background. She is bent forward with her hands braced on a display table, her white sailor-school uniform top pulled open to expose her bare back and the side of her breast, while her short pleated skirt is hiked up over her raised hips to reveal black thigh-high stockings. A man behind her, mostly out of frame, grips her hips and thrusts into her slowly and rhythmically; the camera pushes in with small amplitude at slow speed as her body rocks forward with each motion, her hair and skirt swaying, her breasts moving with realistic weight, and her lips parting as she releases soft, rising moans with a deepening flush across her face, while the out-of-focus crowd continues around her through the end of the shot.</div><div>
- </div><div>overall_soundscape: The young woman's soft, throaty moans and quickened breathing grow louder and more rhythmic, layered over the low murmur of the convention crowd, distant camera shutters, and the faint rustle of fabric.</div><div>
- </div><div>non_diegetic_music: N/A</div><div>
- </div><div>模型:Minimax H3 Remix</div>
复制代码
NSFW 还是老老实实 I2VA 吧。当然也因此,建议图像中出现生殖器等器官。另外官方模型在一些姿势上的效果不好,例如传教士。但是笔者在测试时使用的是手写提示词,不确定使用 LLM 进行更详细的描述能否得到改善。
工作流其实需要讲的并不多。除了需要加入超分补帧外,基本都是写了提示词就能点运行一键生成的。在这方面浪费笔墨并不是很划算的事。所以在此不再进行赘述了。
---
现在让我们看看大概是最多人关心,也是笔者认为最难的部分:提示词撰写。如果你有观察笔者前面给出的提示词案例,你会发现笔者几乎在所有提示词中加入了overall_soundscape non_diegetic_music integrated_multimodal_description 这些东西。这些东西是怎么来的呢?答案就在官方的 Prompt Guide 里。
在这里你可以阅读官方对于提示词书写的规范要求。当然,全是英文看着有点累,你可以在附件中下载我使用 DeepSeek V4 Flash 翻译的版本。
当然,我知道一定有读者不愿意当能工智人。可能是英语不好,可能是无从落笔,更多的是懒。笔者也很懒,所以再介绍第二种提示词书写方式,也就是使用 LLM。
打开 Github,搜索 Minimax H3.(此处不提供 Github 链接最主要的原因是论坛会认为 Github 属于不明来路的链接,所以请读者自行搜索。)
点进 Skills 文件夹。
点进去后,将顶上的链接复制发送给 LLM 即可。
之后你就可以直接描述要求,让 LLM 一键完成提示词书写的任务了。
当然,偶尔 LLM 还是会犯蠢,需要肘击:
这里再说一次,Skills 的加载目前是个 Agent 就能完成,即便你不使用笔者同款的 ZCode 也是可以完成的。例如这是 DeepSeek Harness 的效果:
使用 Skills 只需要读者拥有合格的表达能力即可,相比手写确实省了不少事儿。
到这里,笔者就没有什么还能再写的东西了。类似 Anima、Krea 2 都能使用这个方法进行提示词的书写,读者可自行探索。
本文完。
|
|