ComfyUI Workflows
Browse real AI workflows from Flux, SDXL, Pony, ControlNet, video generation and commercial AI creators.
AI Audio Maker Workflow
Have you eever wanted to add ai generated audio to your wan2.2 generations make them as immersive as LTX2.3 well here is the workflow.Download:https://github.com/kijai/ComfyUI-MMAudiodownload the models put them in the folder:ComfyUI/models/mmaudioYou also need Kijay nodes
Wan 2.2 Fast Img2Video + Auto-Captioning (Low VRAM Friendly)
### 🚀 Wan 2.2 Fast Image-to-Video WorkflowThis workflow is designed for speed and ease of use. It generates high-quality 2.5s videos in under 90 seconds (depending on GPU) using the Wan 2.2 5B model.✨ Key Features:* Auto-Captioning: Uses WD14 Tagger to automatically analyze your input image and generate tags.* Smart Prompting: Automatically concatenates your manual prompt with the image tags to ensure the video stays true to the source image.* Fast Generation: Optimized for speed (10 steps) without sacrificing too much quality.* Dual Output: Saves in both WebP and WebM formats.🛠️ How to use:1. Load Image: Upload your starting image in the "Load Image" node.2. Add Motion Prompt: In the "ttN text" node (green box), describe the movement you want (e.g., "smiling, wind blowing, blinking").3. Run: The workflow will combine your text with the visual description of the image and generate the video.📦 Requirements (ComfyUI): Please use *ComfyUI Manager** to "Install Missing Custom Nodes".* Required Models: * UNET: wan2.2_ti2v_5B_fp16.safetensors * VAE: wan2.2_vae.safetensors * CLIP: umt5_xxl_fp16.safetensors💡 Performance:Tested on mid-range GPUs. Generates a 2.5-second clip in approx 1:30 minutes.---Created to help you animate static images quickly! Enjoy.
Wan 2.2 5B Image-to-Video Simple ComfyUI Workflow
How to use workflow refer to the video
EASY Text-to-Video with Wan 2.2 5B in ComfyUI
How to use the workflow please refer to the video
Project t2i2v
A comfyUI workflows start from generating an image and then use that image to generate a video.all work done on my 8gb vram laptop
Ovi v1.1 New 10s Audio + Video Model That Says Anything You Type
all the links are in the description of the video
Ovi (only 10 Steps with turbo lora) Open-source AI video with audio! Veo 3.1 In ComfyUI Locally
Ovi Open-source AI video with audio! Create Talking Character Video Veo 3.1 In ComfyUI Locally00:00 Demo 00:44 Image 2 Video 07:44 Text 2 VideoDownload All required Models here
HuMo Audio-Driven Text-to-Speech Digital Human Stable Revision V3
You can click on the link below and try it out directly. If the effect is good, you can deploy it locallyhttps://www.runninghub.cn/post/1975450428768956418?inviteCode=rh-v1121Fan benefits, register to get 1000 points, daily login 100 points, play 4090! Experience the super power of 48G!B-site video: https://www.bilibili.com/video/BV1vMxozeEY1/Knowledge Planet: https://t.zsxq.com/7F90AGraphic tutorial: https://t.zsxq.com/lNK8vThe workflow is very simple to use, and I have uploaded a tutorial on how to use it on YouTube. You can try it out and follow the tutorial and Civitai's workflow to easily reproduce the results. Thank you. If you have any other questions, please leave a comment in the comment sectionThis tutorial is very important, not watching it will surely lead to a lot of problems! Watch the video tutorial before trying !!!! Very important! Especially a key point at the end!
WAN 2.2 5b i2v Workflow
WAN 2.2 5B WorkflowA workflow for WAN 2.2 5B with multiple improvements for speed, flexibility, and usability:Auto resolution calculator – automatically adjusts resolution based on a maximum dimension.Video length in seconds – define duration in seconds instead of frames.Fast interrupt – integrates BlehModelPatchFastTerminate to stop WAN generations more quickly.FastWan LoRA – accelerates sampling. Recommended settings: steps = 4–6, CFG = 1.0.Taew2.2 Preview – extremely fast (~1s) preview decode to check quality before running the slower wan2.2_vae.2-Step Process – first run generates a preview with taew2_2 , second run applies the full higher quality VAE decode with upscale and interpolation.2-Step ProcessThe wan2.2_vae is slooooooooow...Sometimes the VAE decode process takes longer than the generation itself.This workflow allows partially decoding the latent video using the taew2_2 model, which is extremely fast (usually ~1 second, depending on video length and resolution), but slightly reduces video quality.By using taew2_2 to create a preview, you can quickly decide if the video is worth keeping or if it needs prompt/setting changes before running the slow wan2.2_vae. This saves a lot of time, especially on lower-end hardware.How to UseStep 1: On the first run, the workflow generates the video and decodes it with the taew2_2 model (fast preview).Step 2: If nothing is changed, the next run will decode the same latent video using the full wan2.2_vae model. After that, further processing (upscale/interpolation) will be applied, if enabled.ImportantThe seed must be fixed. If it changes, a new video will be generated, and the workflow won’t decode the same latent with the full wan2.2_vae.DownloadGet the taew2_2 model here:https://github.com/madebyollin/taehv/blob/main/taew2_2.pthPlace the file in:ComfyUI\models\vae_approx\FastWan LoraIn order to use lower step count and speedup the generation, download the FastWan lora and select it on the workflow.https://huggingface.co/Kijai/WanVideo_comfy/blob/main/FastWan/Wan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensorsAfter enabling it, set:Steps: 4–6CFG: 1.0Source: KijaiWorkflow OptionsThe workflow includes several features you can enable or disable to customize the output:Enable Fast taew2_2 Preview – Uses taew2_2 for fast preview decoding.Enable Sage Attention – Enables the use of Sage Attention during generation.Enable Two Steps Process – Runs first with taew2_2 preview, then with full wan2.2_vae decoding.Enable Use Tiled VAE – Uses tiled VAE decoding for handling larger resolutions.Enable Color Correction – Calibrates video colors using the first frame as reference.Enable Interpolate – Interpolates the video to 50 FPS for smoother motion.Enable Upscale – Upscales the video using the model selected in the “upscale” group.Enable Upscaled Video Show/Save – Displays and saves the upscaled video.
Lucy+Edit Video Editor (Global Alteration - Probability Card Draw)
You can click on the link below and try it out directly. If the effect is good, you can deploy it locallyhttps://www.runninghub.ai/post/1952381519627259905/?/?inviteCode=rh-v1121Fan benefits, register to get 1000 points, daily login 100 points, play 4090! Experience the super power of 48G!B-site video: https://www.bilibili.com/video/BV1EknczKEWF/Knowledge Planet: https://t.zsxq.com/7F90AGraphic tutorial: https://t.zsxq.com/lNK8vThe workflow is very simple to use, and I have uploaded a tutorial on how to use it on YouTube. You can try it out and follow the tutorial and Civitai's workflow to easily reproduce the results. Thank you. If you have any other questions, please leave a comment in the comment sectionThis tutorial is very important, not watching it will surely lead to a lot of problems! Watch the video tutorial before trying !!!! Very important! Especially a key point at the end!
Lucy+Edit Video Editing (Object Addition)
You can click on the link below and try it out directly. If the effect is good, you can deploy it locallyhttps://www.runninghub.ai/post/1952381519627259905/?/?inviteCode=rh-v1121Fan benefits, register to get 1000 points, daily login 100 points, play 4090! Experience the super power of 48G!B-site video: https://www.bilibili.com/video/BV1EknczKEWF/Knowledge Planet: https://t.zsxq.com/7F90AGraphic tutorial: https://t.zsxq.com/lNK8vThe workflow is very simple to use, and I have uploaded a tutorial on how to use it on YouTube. You can try it out and follow the tutorial and Civitai's workflow to easily reproduce the results. Thank you. If you have any other questions, please leave a comment in the comment sectionThis tutorial is very important, not watching it will surely lead to a lot of problems! Watch the video tutorial before trying !!!! Very important! Especially a key point at the end!
Lucy+Edit Video Editing (Subject Replacement)
You can click on the link below and try it out directly. If the effect is good, you can deploy it locallyhttps://www.runninghub.ai/post/1952381519627259905/?/?inviteCode=rh-v1121Fan benefits, register to get 1000 points, daily login 100 points, play 4090! Experience the super power of 48G!B-site video: https://www.bilibili.com/video/BV1EknczKEWF/Knowledge Planet: https://t.zsxq.com/7F90AGraphic tutorial: https://t.zsxq.com/lNK8vThe workflow is very simple to use, and I have uploaded a tutorial on how to use it on YouTube. You can try it out and follow the tutorial and Civitai's workflow to easily reproduce the results. Thank you. If you have any other questions, please leave a comment in the comment sectionThis tutorial is very important, not watching it will surely lead to a lot of problems! Watch the video tutorial before trying !!!! Very important! Especially a key point at the end!
WAN 2.2 5b WhiteRabbit InterpLoop
喜欢中文的你看这边:英文看完后就是中文版WAN 2.2 5b WhiteRabbit Interp-LoopThis ready-to-run ComfyUI workflow turns one image into a short looping video with WAN 2.2 5b. Then, it cleans the loop seam so it feels natural. Optionally, you can also boost the frame rate and upscale with ESRGAN.In other words, this is an Image to Video workflow that creates loops with WAN 2.2 5b!Why is this so complicated?!WAN 2.2 5b does not fully support injecting frames after the first. If you try to inject a last frame, it will create a looping animation but the last 4 frames will be "dirty" with a strange "flash" at the end of the loop.This workflow leverages custom nodes I designed to overcome this limitation. We trim out the dirty frames and then interpolate over the seam.Model Setup (WAN 2.2 5b)Install these into the usual ComfyUI folders. FP16 = best quality. FP8 = faster and lighter, with some trade-offs.Diffusion model → models/diffusion_models/- FP16: wan2.2_ti2v_5B_fp16.safetensors- FP8: Wan2_2-TI2V-5B_fp8_e5m2_scaled_KJ.safetensorsText encoder → models/text_encoders/- FP16: umt5_xxl_fp16.safetensors- FP8: umt5_xxl_fp8_e4m3fn_scaled.safetensorsVAE → models/vae/- wan2.2_vae.safetensorsOptional LoRA → models/lora/- Recommended: Live Wallpaper StyleTip: keep subfolders like models/vae/wan2.2/ so your growing collection stays tidy.How It Works- Seam prep: we take the very last and first frames and generate new in-betweens that bridge them smoothly. Only those new frames get appended — no duplicate of frame 1.- Full-clip interpolation (optional): multiply in-betweens across the whole video, then resample to any FPS you want.- Upscale (optional): add an upscaler pass before full-clip interpolation using an ESRGAN model of your choice.- Output: saved to your ComfyUI/output/ folder, filename prefix LoopVid.Controls You’ll Care AboutDefaults are set for “safe on most GPUs.” Tweak if you have more VRAM.Full-Clip Interpolation- Roll & Multiply: add more in-betweens everywhere (e.g., ×3).- Reample Framerate: convert to an exact FPS (e.g., 60). Great after Multiply, but you could use it on its own.Other handy knobs- Duration: WAN cost climbs past ~3s (2.2 is tuned up to ~5s).- Working Size: long edge in pixels (shape comes from your input image).- Steps: ~30 is WAN 2.2’s sweet spot.- CFG: WAN default is 5, I have it bumped a little higher. Higher = more “prompt strength,” sometimes more motion.- Schedule Shift: motion vs. stability. Higher = more motion.- Upscale: choose model/target size; reduce tile/batch if you hit OOM.You can find more detailed information on all these settings in the workflow itself.Using Vision Models for Prompts (optional but handy)If writing movement prompts feels daunting, you can use a vision model to get a great starting point. You have a few options.Free Cloud OptionsGoogle's Gemini or OpenAI's ChatGPT are free and will get the job done for most people.- Upload your image and paste the prompt below.- Copy the model’s description and paste it into this workflow’s Prompt field....however, these services are not exactly private and might censor lewd/NSFW requests. That's why you might prefer to explore the other two options.Paid Cloud OptionsThere are many services that offer cloud model access which is a more reliable way to get uncensored access to models.You could pay for credits on OpenRouter for example. Personally, I prefer Featherless because they charge a flat monthly fee which keeps my costs predictable, and they have a strict no log policy. If you decide you want to give them a try, you could always use my referral link which helps me out!If you decide to go the API/Paid Cloud route, you might find my app, CloudInterrogator, useful. It's designed to make prompting cloud vision models as easy as possible and it's fully free and open source!Local Inference OptionI know a lot of people on CivitAI are local-or-nothing types. For you, there is Ollama.Here's the best guide I could find on setting it up. You will want to look at Google's Gemma-3 family of models and look at which size is appropriate for your card.If you use Ollama, there's nothing stopping you from using CloudInterrogator as your access point since Ollama creates OpenAI compatible endpoints, or you could customize this workflow with Ollama nodes for ComfyUI. I don't recommend doing the latter unless you can set it to lock the prompt.Many workflows for WAN build Gemma3/Ollama nodes into the workflow. I decided not to do that, because I think 99% of people are going to be well serviced by Gemini or ChatGPT.Suggested prompt:Analyze the content of this video frame and write a concise, single-paragraph description of your predictions around what movement takes place throughout the video sequence that follows. Your description should include the details of the character and scene as a whole but only as they related to the movement that occurs in the scene. In addition, make note of the movements of the particles, blinking of the eyes if any, movement of the hair... this is a moment captured in time, and you are describing these few seconds encompassed by the image. Everything that can move, does move - even minute details of the scene. Do not describe ‘pauses’. Don't minimize the motion with words like ‘slight’ or ‘subtle’. Do not use metaphorical language. Your description must be direct and decisive. Use simple, common language. Be specific, and describe how each detail in the scene moves, but do not be verbose; each word in your description must have purpose. Use the present tense, as if your predictions are coming true as you type them. You will deliver one paragraph without any additional information and without any special characters that format this response, avoid ‘The image sequence depicts the character’ and describe what happens, without saying ‘the video ....’"You might also have good luck with the suggested prompt from AmazingSeek's workflow depending on the model you use or what you're looking for!Tips & TroubleshootingWAN Framerate: WAN 2.2 is 24 fps. On WAN 2.1, if you decide to try it, set fps to 12 instead. There is a slider for this near the model loader node. The workflow auto-calculates what to do with your framerate (for multiplication and resampling) based on this number.Seam looks off? Try switching between Simple/Fancy seam interpolation; increase the auto-crop search in Fancy; or re-render with a slightly different prompt/CFG.Out-of-memory (OOM)?- Lower tile size (x and y) in the WanVideo Decode node.- Lower Upscale tile size and/or batch size.- Reduce Working Size or Duration.- Enable “Use Tiled Encoder”.AttributeError: type object 'CompiledKernel' has no attribute 'launch_enter_hook'I'm not sure what causes this, though my assumption is it has something to do with the WAN Video Nodes. This should fix it for you:1. Open "🧩 Manager"2. Click "Install PIP Packages"3. Install these two, leave the quotes out: "SageAttention", "Triton-Windows".3.1 Obviously Triton-Windows is only for Windows users. If you get this error on Linux, I would guess the package for Triton is just "Triton".If this doesn't fix it for you, it may be that your ComfyUI Python environment is messed up for some reason or the version of Comfy you're using doesn't work with the Manager "Install PIP Packages" module. In that case, you might find this advice from the comments section helpful:From alex223:"i spent almost a day, but made it work. this thing helped, but also, for some reason my embedded python missed include and libs folder, I copied them from standalone version - that was essential for triton to work. Maybe my comment will help someone."If you're still having problems, you can leave a comment. I don't mind trying to help people troubleshoot but I don't think the issue is with my workflow or with WhiteRabbit (my custom nodes).Acknowledgements- It occurred to me that interpolating over a loop seam might be a good solution to the "dirty frames" problem when I was first experimenting, but it was this workflow by AmazingSeek that really made me decide to go for it.- It appears that Ekafalain should get some credit here, too, for their seamless loop workflow on which AmazingSeek's is based.- While I didn't end up using any of their ideas directly, I want to shout out Caravel for their excellent, multi-step process you can have a look at over here that seems to primarily target WAN 2.2 14b. The level of documentation in this workflow alone is laudable.- My recommended vision prompt is built off of NRDX'. You can find the original workflow it's from over on his patreon. This is the guy who is training LiveWallpaper LoRA for various WAN models, too!P.S. 💖If this workflow helps you, I’d love to see what you make! I put a lot of hard work into making it, including designing custom nodes to bring it all together and trying to document as much as possible so it is maximally useful to you.Links- Have a look at the WhiteRabbit repository for node documentation and atomic workflows if you want a better idea of how to build with the custom nodes here or tweak this workflow.- My Website & Socials: See my art, poetry, and other dev updates at artificialsweetener.ai- Buy Me a Coffee: You can help fuel more projects like this at my Ko-fi pageThis workflow is dedicated to my beloved Cubby 🥰- Find her artwork all over the internet- She has many excellent LoRA on CivitAI for you to explore :3WAN 2.2 5b WhiteRabbit 插值循环这个开箱即用的 ComfyUI 工作流可将一张图片转换为使用 WAN 2.2 5b 生成的短循环视频。随后,它会清理循环衔接处的“接缝”,让过渡更自然。可选地,你还可以提升帧率并用 ESRGAN 进行放大。换句话说,这是一个利用 WAN 2.2 5b 生成循环效果的“图像转视频”工作流!为什么会这么复杂?!WAN 2.2 5b 并不完全支持在首帧之后继续注入帧。如果你尝试注入最后一帧,它虽会生成循环动画,但最后 4 帧会出现“脏帧”,在循环结束处出现奇怪的“闪烁”。此工作流通过我设计的自定义节点来规避这一限制。我们先裁掉脏帧,然后对接缝进行插帧插值。工作流内同时提供了“简单版”和“进阶版”的裁剪/插值流程,并配有切换开关,便于你分别试用。模型设置(WAN 2.2 5b)按常规 ComfyUI 目录安装这些文件。FP16 = 质量最佳;FP8 = 更快更省显存,但有一定取舍。扩散模型 → models/diffusion_models/FP16:wan2.2_ti2v_5B_fp16.safetensorsFP8:Wan2_2-TI2V-5B_fp8_e5m2_scaled_KJ.safetensors文本编码器 → models/text_encoders/FP16:umt5_xxl_fp16.safetensorsFP8:umt5_xxl_fp8_e4m3fn_scaled.safetensorsVAE → models/vae/wan2.2_vae.safetensors可选 LoRA → models/lora/推荐:Live Wallpaper Style提示:使用诸如 models/vae/wan2.2/ 这类子文件夹,便于管理不断增长的模型集合。工作原理接缝准备:取最后一帧与第一帧,生成新的过渡中间帧以实现平滑衔接。只会追加这些新帧——不会重复追加第 1 帧。全片插值(可选):在整段视频中增加倍数级的中间帧,然后重采样到任意 FPS。放大(可选):在全片插值之前加入一次放大流程,使用你选择的 ESRGAN 模型。输出:保存到你的 ComfyUI/output/ 文件夹,文件名前缀为 LoopVid。你会关心的控制项默认设置为“对多数 GPU 安全”。如果你显存更充裕,可以适当调高。全片插值滚动倍增 ("Roll & Multiply"):在全片范围增加更多中间帧(例如 ×3)。重采样帧率 ("Resample Framerate"):转换到精确的 FPS(例如 60)。在倍增后使用效果更佳,但也可单独使用。其他实用旋钮时长 ("Duration"):超过 ~3 秒成本上涨(2.2 调校到 ~5 秒)。工作尺寸 ("Working Size"):以长边像素为准(纵横比来自输入图)。步数 ("Steps"):~30 是 WAN 2.2 的甜点区。CFG:WAN 默认 5,这里略微上调。数值越高=“提示强度”更高,有时也会带来更多运动。日程偏移(Schedule Shift):运动 vs 稳定。数值越高=运动更强。放大 ("Upscale"):选择模型/目标尺寸;如遇 OOM,降低 tile/batch。关于这些设置的更多细节,可在工作流中直接查看。使用视觉模型来生成提示(可选但好用)如果编写“运动提示”让你犯难,可以借助视觉模型获得一个很好的起点。你有多种选择。免费云端方案Google 的 Gemini 或 OpenAI 的 ChatGPT 是免费的,对多数人来说足够用了。上传你的图片并粘贴下方提示词。复制模型给出的描述,将其粘贴到本工作流的 Prompt 字段。……不过,这些服务的私密性并不理想,并且可能会审查低俗/NSFW 类请求。这也是你或许想尝试其他两种方案的原因。付费云端方案有很多服务提供云端模型访问,这是获取未审查模型的更可靠方式。例如,你可以在 OpenRouter 购买点数。就我个人而言更偏好 Featherless,因为它按月固定收费、成本可预期,而且有严格的“无日志”政策。如果你想试试,也可以使用我的推荐链接来支持我!如果你选择 API/付费云路线,我的应用 CloudInterrogator 可能会对你有用。它旨在尽可能简化云端视觉模型的提示流程,而且完全免费开源!本地推理方案我知道 CivitAI 上有不少“只用本地”的用户。你可以选择 Ollama。这里有我能找到的最佳安装指南。你可以关注 Google 的 Gemma-3 模型家族,并选择与你显卡匹配的规模。如果使用 Ollama,你完全可以把 CloudInterrogator 当作访问入口,因为 Ollama 提供 OpenAI 兼容的端点;或者你也可以为 ComfyUI 加上 Ollama 节点来定制本工作流。除非你能把提示锁定,否则我并不推荐后者。许多 WAN 工作流会把 Gemma3/Ollama 节点直接内置进去。我选择不这样做,因为我认为 99% 的人用 Gemini 或 ChatGPT 就已经足够。建议的提示词:分析该视频帧的内容,用一个简洁的单段落描述你对随后的整段视频序列中将发生哪些运动的预测。 你的描述应覆盖角色与场景的整体细节,但只限于与场景中“运动”相关的部分。另外,请记录粒子的运动、如果有的话眼睛的眨动、头发的摆动……这是一个被时间定格的瞬间,你要描述的是这张图像所涵盖的这几秒内发生的事。凡是可能运动的,都在运动——包括场景中微小的细节。 不要描述“停顿”。不要用“轻微”“细微”这类词来弱化运动。不要使用隐喻性语言。你的描述必须直接而明确。使用简单、常用的语言。要具体,说明场景中每个细节是如何运动的,但不要冗长;你写下的每个词都要有用处。使用现在时,好像你的预测在你输入时正在成真。 你将输出一个段落,不包含任何额外信息,也不要使用会改变格式的特殊字符;避免用“图像序列描绘了角色……”之类的说法,直接描述发生了什么,不要说“视频……”。根据你所用的模型或目标,你也许会发现 AmazingSeek 工作流提供的提示词同样好用!技巧与故障排查WAN 帧率:WAN 2.2 为 24 fps。若尝试 WAN 2.1,请将 fps 设为 12。模型加载节点附近有对应滑块。工作流会基于该数值自动计算帧率相关流程(倍增与重采样)。接缝看起来不对?试试在“简单/进阶”接缝插值之间切换;在进阶模式中增加自动裁剪搜索范围;或用略微不同的提示/CFG 重新渲染。显存不足?在 WanVideo Decode 节点降低 tile 尺寸(x 和 y)。降低放大(Upscale)的 tile 尺寸和/或批大小。减小工作尺寸或时长。启用“Use Tiled Encoder”。致谢最初试验时,我想到在循环接缝处做插值可能解决“脏帧”问题,但真正让我决定上手的是 AmazingSeek 的这个工作流。看起来 Ekafalain 也应在此获得一些认可,AmazingSeek 的无缝循环工作流是基于其成果之上的。虽然我最终没有直接采用他们的想法,但仍想致敬 Caravel——他们面向 WAN 2.2 14b 的多步流程非常出色,你可以在这里查看,文档水准就值得称赞。我推荐的视觉提示是基于 NRDX 的版本改写而来。你可以在他 Patreon 上找到原始工作流。他也是为多种 WAN 模型训练 LiveWallpaper LoRA 的那位!附言 💖如果这个工作流对你有帮助,我很想看看你的作品!我为此投入了大量精力,包括设计自定义节点把一切串起来,并尽量详细地撰写文档,以便它对你尽可能有用。链接若想更好地了解如何用这些自定义节点搭建,或如何微调本工作流,请查看 WhiteRabbit 仓库中的节点文档与原子工作流。个人网站与社交:在 artificialsweetener.ai 查看我的艺术、诗歌及开发动态请我喝咖啡:在我的 Ko-fi 页面支持更多类似项目本工作流献给我挚爱的 Cubby 🥰你可以在全网各处看到她的作品她在 CivitAI 上也有许多优秀的 LoRA 供你探索 :3
Wan 2.2 5B - Latent Video Upscaler & Enhancer / Transform Low-Res Videos into HD Masterpieces — The Intelligent Way
Transform Low-Res Videos into HD Masterpieces — The Intelligent WayIntroduction: Beyond Traditional UpscalingTraditional AI upscalers like RealESRGAN are great for images, but they often struggle with videos. They can introduce artifacts, fail to add meaningful detail, and leave footage looking blurry and unconvincing.This workflow, "Wan 2.2 5B - Latent Video Upscaler," offers a paradigm shift. Instead of just guessing pixels, it uses the immense power of the Wan 2.2 5B Text-to-Video model to intelligently reinterpret and reconstruct your video in high definition. It doesn't just scale up; it dreams up the missing details, resulting in a cleaner, more detailed, and more coherent HD video than any conventional upscaler can achieve.TL;DR: Stop using image upscalers on video. Use a diffusion model to truly enhance and upscale your footage with intelligent detail.Key Features & Highlights🤖 Intelligent Enhancement: Leverages the Wan 2.2 5B model to add semantically correct details, textures, and coherence, far surpassing the capabilities of traditional upscalers.⚡ Fast & Efficient: Built on the lightweight 5B parameter model, this workflow performs latent upscaling and denoising significantly faster than generating from scratch.🎨 Quality Preservation: Applies a light touch (denoise=0.2) to enhance and upscale without altering the original motion or content of the video drastically.📈 2x Resolution Boost: Doubles the resolution of your input video directly in the latent space before decoding.🎬 Smooth Final Output: Includes an optional RIFE frame interpolation pass to double the frame rate (from 16fps to 32fps) for buttery-smooth motion in the final render.🔊 Audio Passthrough: Automatically carries over the original audio track from your source video to the final enhanced output.Workflow Overview & StrategyThis workflow is a sophisticated video processing chain:Input: Load your low-resolution source video using VHS_LoadVideo.Initial Upscale: The video is immediately 2x upscaled using a Lanczos filter to get to the target size. This provides a better starting point for the model.Latent Processing: The upscaled frames are encoded into the latent space.Intelligent Enhancement: The core of the workflow. The Wan 2.2 5B model, guided by quality-positive and detail-negative prompts, gently denoises (denoise=0.2) the latents over just 8 steps with UniPC. This step is where the "magic" happens—the model fills in plausible, high-quality details.Decoding: The enhanced latents are decoded back into a high-resolution image sequence.Final Output:Option A: Save the immediately upscaled video at 16fps.Option B (Recommended): Pass the sequence through RIFE VFI to interpolate frames to 32fps, creating a final video that is both high-resolution and super smooth.Technical Details & Requirements🧰 Models Required:Base Model: (GGUF Format)Wan2.2-TI2V-5B-Q8_0.ggufSource: Likely from HuggingFace or other model repositories.LoRA:Wan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensors (Applied at strength 0.5)VAE:Wan2.2_VAE.safetensorsCLIP Vision: (For GGUF Loader)umt5-xxl-encoder-q4_k_m.ggufInterpolation Model:rife47.pth (For RIFE VFI node)⚙️ Recommended Hardware:A GPU with a good amount of VRAM (e.g., 12GB+) is recommended for comfortable operation, especially when processing longer videos.🔌 Custom Nodes:This workflow uses:comfyui-videohelpersuite (VHS) - For video loading/combiningcomfyui-frame-interpolation - For RIFE VFIcomfyui-gguf / gguf - For model loadingcomfyui-easy-use - For memory managementcomfyui-kjnodes - For performance patches (Sage Attention)Usage InstructionsLoad the JSON: Import the provided .json file into your ComfyUI.Load the Models: Ensure all required models are in their correct folders. Check the paths in the LoaderGGUF, VAELoader, and LoraLoaderModelOnly nodes.Select Your Video: In the VHS_LoadVideo node, click the video icon to select your low-resolution input video.Queue Prompt: Run the workflow!Retrieve Output: Find your two enhanced videos in the output directory:.../Wan 2.2 5B Upscales/Denoise 0.2_xxxxx.mp4 (16fps).../Wan 2.2 5B Upscales/Denoise 0.2_32fps_xxxxx.mp4 (32fps - Smoother)Tips & TricksDenoise Strength: The denoise parameter in the KSampler (default 0.2) is key.~0.1-0.3: Best for upscaling/enhancement. Preserves the original content while improving quality.>0.5: Will start to significantly alter the content and style, moving towards a new generation based on your video.Source Quality: This workflow excels at breathing new life into low-quality, pixelated, or noisy source videos from older generators.Prompt Engineering: The positive prompt (high detail, high quality...) is generic to encourage enhancement. For stylistic changes, you can modify this prompt (e.g., "cinematic, film grain, photorealism").Conclusion: The Future of Video UpscalingThis workflow demonstrates a powerful new application for diffusion models: not just as generators, but as intelligent enhancement tools. By leveraging the knowledge within the Wan 2.2 model, we can upscale videos with a level of coherence and detail that traditional methods simply cannot match. It’s faster than full generation and smarter than simple scaling.Upload your low-res clips and witness the intelligent upscaling revolution.Credit: Crafted by the ComfyUI community. Special thanks to the creators of the Wan 2.2 models and the FastWanFullAttn LoRA.
Wan2.2 5B Fun Control - Fast Video ControlNet
Workflow OverviewThis is a sophisticated ComfyUI workflow designed for high-quality, controllable video generation using the powerful Wan2.2 5B Fun model. It leverages ControlNet (via Canny edge detection) to transform a driving motion video and a starting reference image into a stunning, coherent animated sequence. Perfect for creating dynamic character animations with consistent style and precise motion transfer.Core Concept: Use a "control video" (e.g., a person dancing) to guide the motion, and a "reference image" (e.g., a character design) to define the style and subject. The workflow intelligently merges them into a new, AI-generated video.Key Features & Highlights🚀 State-of-the-Art Model: Utilizes the Wan2.2-Fun-5B-Control-Q8_0.gguf quantized model for a balance of incredible quality and manageable hardware requirements.🎨 Precision Control: Implements a Canny Edge ControlNet. The workflow extracts edges from your input video, ensuring the generated animation perfectly follows the original motion.⚡ Optimized for Speed: Integrates a custom LoRA (Wan2_2_5B_FastWanFullAttn), allowing for high-quality results in just 8 sampling steps without significant quality loss.🧠 Efficient LLM Inference: Uses a separate, quantized umt5-xxl-encoder CLIP model for text encoding, reducing VRAM load on your GPU.🔧 Complete Pipeline: Everything from model loading, video preprocessing, conditioning, sampling, to final video encoding is included in one seamless, organized graph.📁 Ready-to-Use: Pre-configured with optimal settings, including a detailed positive/negative prompt. Just load your own image and video to start creating.Workflow StructureThe workflow is neatly grouped into logical sections for easy understanding and customization:Step1 - Load models: Loads the main Wan2.2 5B model, its VAE, the CLIP text encoder, and the FastWan LoRA.Step 2 - Start_image: Loads your initial reference image. This defines the character and style for the first frame.Step 3 - Control video and video preprocessing: Loads your motion video and processes it through the Canny node to extract edge maps.Step 4 - Prompt: Where you input your positive and negative prompts to guide the generation.Step 5 - Video size & length: The Wan22FunControlToVideo node packages everything, setting the output video dimensions and length based on the control video.Sampling & Decoding: The KSampler runs for 8 steps with UniPC, and the VAE decodes the latents into final images.Video Output: The VHS_VideoCombine node encodes the image sequence into an MP4 video file.How to Use This WorkflowDownload & Install:Ensure you have ComfyUI Manager to easily install missing custom nodes.Required Custom Nodes: ComfyUI-VideoHelperSuite, ComfyUI-GGUF (for loading the .gguf models).Download the .json file from this post.Load the Models:Main Model: Place Wan2.2-Fun-5B-Control-Q8_0.gguf in your ComfyUI/models/gguf/ folder.CLIP Model: Place umt5-xxl-encoder-q4_k_m.gguf in the same gguf/ folder.VAE: The workflow points to Wan2.2_VAE.safetensors. Ensure it's in your models/vae/ folder.LoRA: Place Wan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensors in your models/loras/ folder. Adjust the path in the LoraLoader node if yours is in a subfolder (e.g., wan_loras/).Load Your Assets:Reference Image: In the LoadImage node, change the image name to your own file (e.g., my_character.png).Control Video: In the LoadVideo node, change the video name to your own motion clip (e.g., my_dance_video.mp4).Customize Your Prompt:Edit the text in the Positive Prompt node to describe your desired character and scene.The provided negative prompt is already comprehensive, but you can modify it as needed.Run the Workflow:Queue the prompt in ComfyUI. The final video will be saved to your ComfyUI/output/video/ folder.Tips for Best ResultsControl Video: Use a video with clear, strong motion and good contrast for the Canny detector to work best. Silhouettes or videos with a plain background work excellently.Reference Image: The first frame of your output will closely match this image. Use a high-quality image of your character in a pose similar to the first frame of your control video.Length: The length in Wan22FunControlToVideo is set to 121 based on the original video. If your video is a different length, you must update this value to match the number of frames.Experiment: Try adjusting the LoRA strength (e.g., between 0.4 - 0.7) or the Canny thresholds to fine-tune the balance between motion fidelity and creative freedom.Required Models (Download Links)Wan2.2-Fun-5B-Control-Q8_0.gguf: https://huggingface.co/QuantStack/Wan2.2-Fun-5B-Control-GGUFumt5-xxl-encoder-q4_k_m.gguf: https://huggingface.co/city96/umt5-xxl-encoder-gguf/tree/mainWan2.2_VAE.safetensors: https://huggingface.co/QuantStack/Wan2.2-Fun-5B-InP-GGUF/tree/main/vaeWan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensors: https://huggingface.co/Kijai/WanVideo_comfy/blob/main/FastWan/Wan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensorsConclusionThis workflow demonstrates the powerful synergy between the Wan2.2 model, ControlNet, and efficient LoRAs. It abstracts away the complexity, providing you with a robust, one-click solution for creating amazing AI-powered animations. Enjoy creating!If you use this workflow, please share your results! I'd love to see what you create.
Wan 2.2 Fun 5B Inpainting - Seamless Image Morphing
🎬 IntroductionUnleash a new form of AI creativity! This workflow harnesses the specialized power of the Wan 2.2 Fun 5B Inpainting model to generate stunning, seamless video animations that morph one image into another.Forget text prompts driving the motion. Here, you provide a starting image and an ending image. The AI intelligently analyzes both and generates a fluid video transition between them, interpreting the content and creating a natural, often dreamlike, transformation. Perfect for creating mesmerizing loops, concept art evolution, or simply bringing two ideas together in a video.This workflow is optimized for accessibility, utilizing GGUF quantization to run this powerful model efficiently on consumer hardware.✨ Key Features & HighlightsImage-to-Image Morphing: The core feature. Input a start and end image, and let the AI generate the transitional video.GGUF Quantization Support: Powered by the LoaderGGUF and ClipLoaderGGUF nodes, making the 5B parameter model runnable without a top-tier GPU.Lightning LoRA for Speed: Integrates a 4-step LoRA, significantly speeding up the generation process compared to standard sampling.Simple & Intuitive Setup: The workflow is cleanly grouped into logical steps: Load Models, Upload Images, Set Prompt, and Generate.High-Quality Output: Configured for a high resolution (944x944) and smooth frame rate (24 FPS), packaged into an MP4 file by the robust VHS_VideoCombine node.🧩 How It Works (The Magic Behind the Scenes)This workflow is elegant in its execution:Load Models (GGUF): The LoaderGGUF node loads the quantized Wan2.2-Fun-5B-InP model, and the ClipLoaderGGUF node loads the UMT5 text encoder. A standard VAELoader node loads the Wan VAE for decoding.Upload Your Image Pair: This is the crucial step. You provide two inputs:start_image: The initial state of your animation.end_image: The final state you want to transition to.Define the Vibe (Prompt): While the images drive the motion, the text prompt helps define the style and quality of the entire generated sequence. The included positive prompt creates a "dreamy, Q-style" look, while the negative prompt filters out common artifacts.Inpainting Magic (WanFunInpaintToVideo): This specialized node is the engine. It takes your two images, encodes them along with your prompts, and prepares a latent video representation that transitions from the start to the end image.Fast Sampling: The prepared latent is passed to the KSampler, which uses the 4-step Lightning LoRA to rapidly denoise the sequence, creating the final frames of the animation.Decode & Export: The VAE decodes the latent frames into images, and the VHS_VideoCombine node seamlessly compiles them into a final, high-quality MP4 video file.⚙️ Instructions & UsagePrerequisite: Download ModelsYou must download the following model files and place them in your ComfyUI models directory.Essential Models:Wan2.2-Fun-5B-InP-Q8_0.gguf → Place in /models/unet/ (or /models/diffusion/)umt5-xxl-encoder-q4_k_m.gguf → Place in /models/clip/wan_2.2_vae.safetensors → Place in /models/vae/For the 4-Step Lightning Pipeline:Wan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensors → Place in /models/loras/ (Note: The workflow currently points to this, but the note mentions a different official LoRA. You may want to clarify which one to use in your upload).Loading the WorkflowDownload the provided video_wan2_2_5B_fun_inpaint.json file.In ComfyUI, drag and drop the JSON file into the window or use the Load button.Running the WorkflowUpload Your Image Pair:In the "LoadImage" node on the left, upload your start_image.png.In the "LoadImage" node on the right, upload your end_image.png.Tip: For best results, use images that are similar in composition or theme for a more coherent morph.Set Your Prompt (Optional but Recommended):Modify the text in the "CLIP Text Encode (Positive Prompt)" node to influence the style of the entire generated video (e.g., "watercolor style", "cyberpunk", "realistic").The negative prompt is pre-filled and is great for general use.Queue Prompt! Watch as the AI dreams up a unique animation that bridges your two images.⚠️ Important Notes & TipsImage Guidance: The strength of this model is in the images. The text prompt plays a secondary role in styling. The core narrative of the video is the transition between your two uploaded images.Length Setting: The WanFunInpaintToVideo node has a length parameter set to 121 frames. At 24 FPS, this will yield roughly a 3.3-second video. You can adjust this value for shorter or longer animations, but be mindful of VRAM constraints.Resolution: The workflow is set to 944x944. You can adjust the width and height in the WanFunInpaintToVideo node, but this will also impact VRAM usage and performance.Lightning LoRA: The 4-step LoRA is used for speed. If you encounter quality issues or want to try a different style, you can adjust the strength in the LoraLoaderModelOnly node or try the official wan2.2_i2v_lightx2v_4steps_lora mentioned in the node's info.🎭 Example ResultsStart Image: A closed flower bud.End Image: The same flower in full bloom.Prompt: "A beautiful time-lapse of a flower blooming, macro photography, sharp focus, cinematic lighting."(You would embed a short video example generated by this workflow here)Another idea: Start with a sketch, end with the fully rendered artwork.📁 Download & LinksWan 2.2 Fun 5B Inpaint GGUF Model: HuggingFace - QuantStack/Wan2.2-Fun-5B-InP-GGUFumt5-xxl-encoder-q4_k_m.gguf: https://huggingface.co/city96/umt5-xxl-encoder-gguf/tree/mainWan2.2_VAE.safetensors: https://huggingface.co/QuantStack/Wan2.2-Fun-5B-InP-GGUF/tree/main/vaeWan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensors: https://huggingface.co/Kijai/WanVideo_comfy/blob/main/FastWan/Wan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensorsConclusion💎 ConclusionThis workflow opens up a fascinating and less-explored avenue of AI video generation. By moving beyond text-to-video to image-guided-video, it allows for precise and creative control over the starting and ending points of an animation. The use of GGUF makes this creative power accessible to a wide audience.It's perfect for artists, designers, and anyone looking to create unique, seamless transitions and visual stories. Experiment with different image pairs and prompts to discover the full potential of this "Fun" model.We can't wait to see what you morph! Share your creations in the comments below.
WAN2.2 5B - First-Last-Frame to Video (FLF2V)
Unlock the power of 14B-style animation control on the 5B model! This workflow uses a custom node to perform First-Last-Frame conditioning, letting you guide the video's start and end point with precision.Breakthrough Accessibility: This workflow brings a flagship feature from the demanding WAN2.2 14B model to the efficient and popular 5B variant. By leveraging the powerful Wan22FirstLastFrameToVideoLatent custom node, you can now define both the starting frame AND the ending frame of your generated video, giving you unprecedented control over the animation's narrative arc.🎯 Precision-Guided Video Generation:First-Last-Frame (FLF) Conditioning: The core innovation. You provide two images:Start Image: The first frame of your video.End Image: The desired final frame.The model then generates a coherent video that smoothly transitions between these two defined points.Unmatched Creative Control: This is far more powerful than standard image-to-video. Want a flower to bloom? A character to turn and smile? An object to transform? Define the start and end states, and let the AI handle the complex in-between motion.Perfect for the 5B Model: Enjoy this advanced capability without the massive VRAM requirements of the 14B model. This makes high-control animation accessible to a much wider audience.Simple & Efficient: The workflow is clean, focused, and designed for simplicity. Load your two images, set your prompt, and generate. It's perfect for both experimentation and production.⚙️ How It Works:Input: You load two key images: the initial state and the target state.Encoding: The custom node Wan22FirstLastFrameToVideoLatent encodes both images into the model's latent space alongside your text prompt.Generation: The KSampler uses this combined information to generate a video latent that respects both the starting condition and the desired outcome.Output: Decode the latent into a video file that seamlessly animates from your first frame to your last.✨ Key Features:Custom Node Power: Utilizes stduhpf/ComfyUI--Wan22FirstLastFrameToVideoLatent to enable this advanced feature.Full 5B Compatibility: Works with wan2.2_ti2v_5B_fp16.safetensors and standard WAN components (umt5 CLIP, wan2.2_vae).Complete Pipeline: Includes everything from model loading to final MP4 video output with the Video Helper Suite node.Flexible: The note node reminds you that you can use full GPU models if you have the VRAM, making it adaptable to different hardware setups.🎨 Ideal For:Storyboarding: Plan and visualize scenes with exact beginning and ending frames.Controlled Transformations: Create precise morphing, growth, rotation, or state-change animations.Artistic Exploration: Experiment with having the AI solve the "how" of getting from point A to point B.Users who want the creative control of the 14B model but need the practicality of the 5B model.⚠️ Requirements:Custom Node: You must install the Wan22FirstLastFrameToVideoLatent node from its repository (e.g., stduhpf/ComfyUI--Wan22FirstLastFrameToVideoLatent).Standard WAN Dependencies: The usual suspects: ComfyUI-VideoHelperSuite, and the required WAN2.2 5B model files.This workflow is a game-changer for WAN2.2 5B users. It dramatically expands the creative possibilities of the model, moving it from a simple text-to-video tool to a powerful keyframe-based animation system.Stop just generating videos. Start directing them. Download the workflow and define your story's beginning and end.
WAN2.2 5B Ultimate Suite - T2V, I2V & T2I2V Pro
The most advanced and comprehensive WAN2.2 5B workflow on CivitAI. This all-in-one suite masterfully combines Text-to-Video, Image-to-Video, and Text-to-Image-to-Video generation, powered by local LLMs (Ollama) for intelligent, dynamic prompt enhancement. Stop using basic prompts; generate cinematic, fluid animations with intelligent motion design.Workflow DescriptionUnleash the full potential of the WAN2.2 5B model with this meticulously designed, feature-packed ComfyUI workflow. This isn't just a simple pipeline; it's a professional content creation suite that intelligently bridges the gap between your ideas and stunning AI-generated video.Why This Workflow Stands Out:* 🤖 AI-Powered Intelligence: Integrated Ollama LLMs analyze your text or images to generate richly detailed, dynamic prompts specifically engineered for WAN2.2's video capabilities. It translates static concepts into descriptions full of motion, lighting, and cinematic life.* 🎬 Multi-Modal Mastery: Seamlessly switch between three powerful generation modes without changing workflows.* ⚙️ Optimized & Robust: Built with stability and efficiency in mind. Includes automatic GPU memory management, frame interpolation, and a professional video output system.* 🔄 All-in-One Pipeline: From a simple idea or image to a final, smooth video file, everything is connected and automated.Features & Technical Details🧩 Core Components:* Model: wan2.2_ti2v_5B_fp16.safetensors* VAE: wan2.2_vae.safetensors* Key Loras: Wan2.2_5B_FastWanFullAttn (style LoRA)* Upscaler: Integrated for pre-processing input images.* Frame Interpolation: RIFE VFI for buttery-smooth 2x frame generation (outputs 24fps and 48fps videos).🔧 Integrated AI Engines (Ollama):* For Text (T2V): huihui_ai/gemma3-abliterated:12b-q8_0 - Analyzes your simple text and generates detailed video prompts with motion, camera work, and atmosphere.* For Vision (I2V): qwen2.5vl:7b-q8_0 - Analyzes any image you provide and writes the perfect animation prompt based on its content.* For T2I (Flux Group): gemma3:latest - Enhances simple text descriptions for high-quality image generation which can then be animated.📊 Output:* Resolution: Adapts to your input image size or defined latent size.* Frames: Configurable length (default: 121 frames).* Format: MP4 (H.264) with proper metadata.* Dual Output: Standard 24fps and interpolated 48fps videos are saved automatically.How to Use / Steps to RunPrerequisites:1. ComfyUI Manager: Essential for installing missing custom nodes.2. Ollama: Installed and running on your system. You must pull the required LLM models gemma3, qwen2.5vl).3. All Models/LoRAs: Ensure all paths in the workflow point to files you actually have. The most common error is a missing model!4. Custom Nodes: The workflow will prompt you to install any missing nodes via ComfyUI Manager. Key node suites include: * comfyui-ollama * comfyui-videohelpersuite * comfyui-frame-interpolation * comfyui-easy-use * gguf (for Flux loading)Usage Instructions:1. TEXT-to-VIDEO (T2V)1. Locate the green "Enter simple prompt here" node.2. Replace the text with your simple idea (e.g., "a knight drawing his sword in a rainy forest").3. Ensure the OllamaConnectivityV2 node points to your Ollama server (default: http://192.168.0.210:11434).4. Queue Prompt. Watch the Ollama node generate a detailed cinematic prompt, which is then used to create the video.2. IMAGE-to-VIDEO (I2V)1. In the "Load Image" node, upload your starting image.2. The image will be automatically analyzed by the Qwen vision model.3. The Ollama node will generate a motion prompt tailored to the image's content.4. Queue Prompt. The workflow will animate your image based on the AI-generated description.3. TEXT-to-IMAGE-to-VIDEO (T2I2V)1. Use the Flux/Krea group (on the left side of the workflow).2. In the PrimitiveStringMultiline node, enter a description for the image you want to generate (e.g., "a gorilla in the jungle eating a banana").3. Run the prompt. This group will generate a high-quality image.4. Once the image is generated, you can manually connect it to the main I2V pipeline or use the provided "Auto-last frame extract" group to automatically find the latest generated image and animate it.⏯️ Output: Your finished videos will be saved to your ComfyUI output/video/ folder. The workflow also saves a preview of the first frame.Tips & Tricks* Ollama Server: The workflow is pre-configured for IP 192.168.0.210. You MUST change this in all three OllamaConnectivityV2 nodes to http://localhost:11434 or your server's IP.* Speed vs. Quality: Adjust the steps in the KSampler (default 8). Lower is faster, higher may yield better quality.* Control: You can bypass the Ollama nodes entirely. Just plug your own expertly crafted positive prompt directly into the "CLIP Text Encode (Positive Prompt)" node.* Troubleshooting: If you get errors, check the ComfyUI console. Most issues are due to incorrect Ollama server addresses or missing model files.This workflow represents the cutting edge of accessible AI video generation. It demonstrates how leveraging multiple AI systems together (diffusion models + LLMs) creates results far beyond what any single model can achieve alone.Enjoy creating, and please share your amazing results!
WAN2.2 5B - Unlimited Long Video Generation Loop
Unlock the potential for infinite video narratives with this powerful ComfyUI workflow. Designed specifically for the WAN2.2 5B text-to-video model, this setup automates the creation of long, coherent video sequences by implementing an intelligent feedback loop. It doesn't just string clips together; it creates a visually consistent and dynamically evolving story.✨ Key Features & Highlights:AI-Powered Prompt Chaining: The core of this workflow. An Ollama multi-modal LLM (like Qwen2.5-VL) analyzes the last frame of each generated video clip and automatically creates a new, detailed prompt for the next segment. This ensures each new clip logically continues from the previous one.Perfect for Long-Form Content: Generate multi-part scenes, evolving transformations, or endless walking cycles without manual intervention. The loop is configurable to run for any number of iterations.Superior Visual Consistency: Incorporates a color matching node (easy imageColorMatch) to harmonize the colors and tones between segments, preventing jarring visual jumps and creating a seamless flow.Built-In Quality Enhancement: Includes a RIFE VFI frame interpolation node that doubles the frame rate of the final assembled video, resulting in buttery-smooth motion.Fully Automated Pipeline: From loading the initial image to rendering the final high-quality video, the process is hands-free after the initial setup.🛠️ How It Works:Preparation: The workflow starts with your initial image, which is scaled and analyzed.Ollama Vision Analysis: The LLM examines the image and generates a dynamic, movement-focused prompt tailored for the WAN2.2 model.Video Generation: The WAN2.2 5B model generates a short video clip (~5 seconds) based on this AI-crafted prompt.Loop & Refine: The last frame is extracted, color-corrected, and fed back to Ollama to generate the next prompt. This loop repeats for your set number of iterations.Final Assembly: All individual clips are combined into a single, smooth, long-form video file.📦 What's Included:.json Workflow file for ComfyUI.A detailed breakdown of the node groups and their functions.Recommended settings for optimal results.⚙️ Recommended Models:Text-Image-to-Video: wan2.2_ti2v_5B_fp16.safetensorsLoRA: Wan2_2_5B_FastWanFullAttn_lora_rank_128_bf16.safetensors (for faster generation)VAE: wan2.2_vae.safetensorsLLM (for Ollama): A vision-capable model like qwen2.5-vl:7b or llava-1.6🎯 Ideal For:Creating music videos with evolving visuals.Generating long animations and story sequences.Producing dynamic social media content loops.Experimenting with AI-driven storytelling and scene progression.Disclaimer: This workflow requires a properly configured ComfyUI environment with the necessary custom nodes (ComfyUI-Easy-Use, Video-Helper-Suite, ComfyUI-Ollama, ComfyUI-Frame-Interpolation) and an Ollama server running with a vision model.
Low vram wan 2.2 ti2v 5b gguf workflow
Low vram gguf workflow
Auto last frame extractor (copypaste wf)
This workflow is made for who want make long video with low spec PC.It's a simple way to extract last frame from generated video, then drag it into your I2V workflow to keep generation consistent.This workflow is made as copypaste , for don't destroy all your workflow.How to use.Select all node in group (CTRL+drag area with mouse)Copy (CTRL+C)Open your video workflowClick on an empty areaPaste (CTRL+V)Connect your VideoCombine-Filenames to "filenames select" node.All done. Now when you run generation, you get automatic last frame from video.Tested on wan2.2 and ltx0.9.8