SeedVR2 VAE TensorRT Engines for NVIDIA Blackwell
High-performance, pre-compiled NVIDIA TensorRT engine plans (.rtxplan) built for the official SeedVR2 VAE (seedvr2_ema_vae_fp16), specifically compiled and optimized for the NVIDIA Blackwell GPU architecture (RTX 5090, RTX 5080, RTX 5070, RTX 5060, RTX 5050, and Blackwell workstation / data center GPUs, Compute Capability sm_100 / sm_120).
π Overview
SeedVR2 is a state-of-the-art diffusion-based video restoration and super-resolution model developed by ByteDance. In video upscaling workflows, spatial-temporal VAE encoding (RGB frames to latents) and decoding (latents back to RGB frames) represent primary computational and memory bottlenecks.
This repository provides 61 pre-compiled TensorRT execution engine plans (.rtxplan) covering both VAE Decoding and VAE Encoding, delivering full end-to-end acceleration on NVIDIA Blackwell hardware:
- Target Architecture: NVIDIA Blackwell (
sm_100/sm_120, RTX 50-Series: RTX 5090, RTX 5080, RTX 5070, RTX 5060, RTX 5050). - Dual-Phase VAE Coverage: Dedicated engines for VAE Encoding (Phase 1) and VAE Decoding (Phase 3).
- Dual Spatial Tile Tiers:
- Tile 256 (256Γ256): Low-VRAM deterministic execution; mandatory for standard decoders and extended long-sequence encoders (105fβ185f) to prevent OOM on 4K/8K.
- Tile 512 (512Γ512): High spatial fidelity; standard for encoders (1fβ97f) and high-quality decoders (1fβ73f) for seamless patch boundary blending.
- Temporal Range: Comprehensive $4n+1$ sequences from single-image (
1f) up to 185 frames (185f). - Base Checkpoint: Comfy-Org/SeedVR2 (seedvr2_ema_vae_fp16) / ByteDance SeedVR2.
- ComfyUI Custom Node: ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT-Decoder
π¦ Available TensorRT Engine Plans (61 Engines)
π Engine Matrix Overview
| Module | Spatial Tile | Temporal Frames ($4n+1$) | Count | Engine Size | Naming Format |
|---|---|---|---|---|---|
| VAE Decoder | 256Γ256 | 5f, 21f, 25f, 29f, 33f, 37f, 41f, 45f, 49f, 53f, 57f, 61f, 65f, 69f, 73f, 77f, 81f, 85f, 89f, 93f, 97f |
21 | ~306β308 MB | vae_decoder_tile_256_<N>f.rtxplan |
| VAE Decoder | 512Γ512 | 1f, 5f, 21f, 25f, 29f, 33f, 37f, 41f, 45f, 49f, 53f, 65f, 69f, 73f |
14 | ~310β320 MB | vae_decoder_tile_512_<N>f.rtxplan |
| VAE Encoder | 512Γ512 | 1f, 5f, 17f, 21f, 25f, 29f, 33f, 37f, 41f, 45f, 49f, 53f, 57f, 61f, 65f, 69f, 73f, 77f, 81f, 85f, 89f, 93f, 97f |
23 | ~220β257 MB | vae_encoder_<N>f_tile512.rtxplan |
| VAE Encoder | 256Γ256 | 105f, 165f, 185f |
3 | ~220β230 MB | vae_encoder_<N>f_tile256.rtxplan |
π Click to expand complete list of all 61 individual engine files
1. VAE Decoder β Tile 256 (21 engines)
vae_decoder_tile_256_5f.rtxplan vae_decoder_tile_256_41f.rtxplan vae_decoder_tile_256_73f.rtxplan
vae_decoder_tile_256_21f.rtxplan vae_decoder_tile_256_45f.rtxplan vae_decoder_tile_256_77f.rtxplan
vae_decoder_tile_256_25f.rtxplan vae_decoder_tile_256_49f.rtxplan vae_decoder_tile_256_81f.rtxplan
vae_decoder_tile_256_29f.rtxplan vae_decoder_tile_256_53f.rtxplan vae_decoder_tile_256_85f.rtxplan
vae_decoder_tile_256_33f.rtxplan vae_decoder_tile_256_57f.rtxplan vae_decoder_tile_256_89f.rtxplan
vae_decoder_tile_256_37f.rtxplan vae_decoder_tile_256_61f.rtxplan vae_decoder_tile_256_93f.rtxplan
vae_decoder_tile_256_65f.rtxplan vae_decoder_tile_256_97f.rtxplan
vae_decoder_tile_256_69f.rtxplan
2. VAE Decoder β Tile 512 (14 engines)
vae_decoder_tile_512_1f.rtxplan vae_decoder_tile_512_29f.rtxplan vae_decoder_tile_512_49f.rtxplan
vae_decoder_tile_512_5f.rtxplan vae_decoder_tile_512_33f.rtxplan vae_decoder_tile_512_53f.rtxplan
vae_decoder_tile_512_21f.rtxplan vae_decoder_tile_512_37f.rtxplan vae_decoder_tile_512_65f.rtxplan
vae_decoder_tile_512_25f.rtxplan vae_decoder_tile_512_41f.rtxplan vae_decoder_tile_512_69f.rtxplan
vae_decoder_tile_512_45f.rtxplan vae_decoder_tile_512_73f.rtxplan
3. VAE Encoder β Tile 512 (23 engines)
vae_encoder_1f_tile512.rtxplan vae_encoder_37f_tile512.rtxplan vae_encoder_69f_tile512.rtxplan
vae_encoder_5f_tile512.rtxplan vae_encoder_41f_tile512.rtxplan vae_encoder_73f_tile512.rtxplan
vae_encoder_17f_tile512.rtxplan vae_encoder_45f_tile512.rtxplan vae_encoder_77f_tile512.rtxplan
vae_encoder_21f_tile512.rtxplan vae_encoder_49f_tile512.rtxplan vae_encoder_81f_tile512.rtxplan
vae_encoder_25f_tile512.rtxplan vae_encoder_53f_tile512.rtxplan vae_encoder_85f_tile512.rtxplan
vae_encoder_29f_tile512.rtxplan vae_encoder_57f_tile512.rtxplan vae_encoder_89f_tile512.rtxplan
vae_encoder_33f_tile512.rtxplan vae_encoder_61f_tile512.rtxplan vae_encoder_93f_tile512.rtxplan
vae_encoder_65f_tile512.rtxplan vae_encoder_97f_tile512.rtxplan
4. VAE Encoder β Tile 256 (3 engines)
vae_encoder_105f_tile256.rtxplan vae_encoder_165f_tile256.rtxplan vae_encoder_185f_tile256.rtxplan
Note on 4n+1 Temporal Structure: SeedVR2 utilizes a 3D causal temporal downsampling VAE architecture. The temporal latent dimension corresponds to $(T - 1) / 4 + 1$. Consequently, video frame sequences are compiled and processed in $4n + 1$ frame blocks. Single image inference uses the dedicated
1fengine.
π οΈ Architectural & Performance Advantages
Native Blackwell Kernel Tuning: Built and tuned with the TensorRT Blackwell compilation pipeline, fully utilizing Blackwell 5th-Generation Tensor Cores, enhanced memory bandwidth, and fused convolution-activation operators across the entire RTX 50-Series lineup (RTX 5090, RTX 5080, RTX 5070, RTX 5060, RTX 5050).
Decoupled Dual-Phase VAE Configuration: Enables independent selection of engine frame sizes for encoding and decoding (e.g. encode with 81f tile512 for maximum patch fidelity, and decode with 65f tile256 for optimal VRAM headroom).
Dynamic 1-Shot Pad & Crop Execution: Short or non-matching batch lengths are dynamically padded to the nearest valid $4n+1$ sequence, executed in a single high-throughput TensorRT pass, and cropped back with pixel precision without falling back to slow FP16 CPU/PyTorch loops.
Studio-Grade Memory & Execution Safety:
- ExecutionContext Lock & Stream Synchronization: Dedicated mutex locks (
_ENCODE_LOCK,_DECODE_LOCK) and explicit stream synchronization eliminate race conditions and memory corruption (black tiles) during concurrent or chained executions. - Deterministic Tri-Partite Memory Cleanup: Immediate release of intermediate tensors (
del,gc.collect(),torch.cuda.empty_cache()) prevents progressive VRAM fragmentation across multi-batch sequences.
- ExecutionContext Lock & Stream Synchronization: Dedicated mutex locks (
π Usage in ComfyUI
These TensorRT engine plans (.rtxplan) are directly loaded and executed via the dedicated custom node:
- Loader & Node Repository: ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT-Decoder
1. Installation
Clone the custom node repository into your ComfyUI custom_nodes/ directory:
cd ComfyUI/custom_nodes
git clone https://github.com/ussoewwin/ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT-Decoder.git
2. Engine Placement
Place the downloaded .rtxplan engine files into either of the following supported directories:
ComfyUI/custom_nodes/ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT-Decoder/tensorrt_backend/artifacts/
or
ComfyUI/models/tensorrt/seedvr2/
Upon launching ComfyUI, all detected .rtxplan files will automatically populate the engine_frames dropdown list of the respective loader nodes:
vae_decoder_tile_*_*f.rtxplanpopulatesSeedVR2 Load TensorRT VAE Decodervae_encoder_*f_tile*.rtxplanpopulatesSeedVR2 Load TensorRT VAE Encoder
3. Node Integration
Connect both loader nodes directly to the SeedVR2 Video Upscaler node:
- Connect the
SEEDVR2_VAEoutput ofSeedVR2 Load TensorRT VAE Encoderto thevae_encodeinput. - Connect the
SEEDVR2_VAEoutput ofSeedVR2 Load TensorRT VAE Decoderto thevae_decodeinput.
+------------------------------------+
| SeedVR2 Load TensorRT VAE Encoder |
| [engine_frames: auto / 81f / ...] |--(SEEDVR2_VAE)---+
+------------------------------------+ |
v [vae_encode]
+----------------------------+
| SeedVR2 Video Upscaler |
+----------------------------+
^ [vae_decode]
+------------------------------------+ |
| SeedVR2 Load TensorRT VAE Decoder | |
| [engine_frames: auto / 65f / ...] |--(SEEDVR2_VAE)---+
+------------------------------------+
4. Workflow Examples
Complete Video Upscaler Workflow (TensorRT VAE & Quantized Models)
Workflow JSON: example_workflows/SeedVR2_tensorrt_decode.json
TensorRT VAE Decoder Node Integration
Connect the TRT_VAE output from SeedVR2 Load TensorRT VAE Decoder directly into the SeedVR2 Video Upscaler node:
π§ Building Custom Frame Engines on Demand
The custom node repository also includes a dedicated SeedVR2 Build TensorRT VAE Engines node to compile custom frame sizes directly inside ComfyUI:
- Supported Frames: Any custom frame count following the $4n+1$ format (e.g. 101f, 185f, 205f).
- Supported Tile Sizes:
256(mandatory for Decoder; optimal VRAM for long Encoder sequences) and512(recommended for standard Encoder sequences). - Supported Targets:
decoder,encoder, orboth.
Built engines are saved directly into tensorrt_backend/artifacts/ and become immediately selectable upon restarting ComfyUI.
π Credits & References
- ComfyUI TensorRT Loader: ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT-Decoder
- SeedVR / SeedVR2 Foundation: ByteDance Seed Team (ByteDance-Seed/SeedVR)
- Official VAE Checkpoint: Comfy-Org/SeedVR2 (
seedvr2_ema_vae_fp16.safetensors) - ComfyUI Implementation: NumZ & AInVFX (ComfyUI-SeedVR2_VideoUpscaler)
- TensorRT VAE Architecture Inspiration: VRGDG-SeedVR2-TensorRT-Studio
- Acceleration Framework: NVIDIA TensorRT
- License: Apache-2.0 License
- Downloads last month
- -

