SeedVR2 VAE TensorRT Engines for NVIDIA Blackwell

High-performance, pre-compiled NVIDIA TensorRT engine plans (.rtxplan) built for the official SeedVR2 VAE (seedvr2_ema_vae_fp16), specifically compiled and optimized for the NVIDIA Blackwell GPU architecture (RTX 5090, RTX 5080, RTX 5070, RTX 5060, RTX 5050, and Blackwell workstation / data center GPUs, Compute Capability sm_100 / sm_120).


🌟 Overview

SeedVR2 is a state-of-the-art diffusion-based video restoration and super-resolution model developed by ByteDance. In video upscaling workflows, spatial-temporal VAE encoding (RGB frames to latents) and decoding (latents back to RGB frames) represent primary computational and memory bottlenecks.

This repository provides 61 pre-compiled TensorRT execution engine plans (.rtxplan) covering both VAE Decoding and VAE Encoding, delivering full end-to-end acceleration on NVIDIA Blackwell hardware:

  • Target Architecture: NVIDIA Blackwell (sm_100 / sm_120, RTX 50-Series: RTX 5090, RTX 5080, RTX 5070, RTX 5060, RTX 5050).
  • Dual-Phase VAE Coverage: Dedicated engines for VAE Encoding (Phase 1) and VAE Decoding (Phase 3).
  • Dual Spatial Tile Tiers:
    • Tile 256 (256Γ—256): Low-VRAM deterministic execution; mandatory for standard decoders and extended long-sequence encoders (105f–185f) to prevent OOM on 4K/8K.
    • Tile 512 (512Γ—512): High spatial fidelity; standard for encoders (1f–97f) and high-quality decoders (1f–73f) for seamless patch boundary blending.
  • Temporal Range: Comprehensive $4n+1$ sequences from single-image (1f) up to 185 frames (185f).
  • Base Checkpoint: Comfy-Org/SeedVR2 (seedvr2_ema_vae_fp16) / ByteDance SeedVR2.
  • ComfyUI Custom Node: ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT-Decoder

πŸ“¦ Available TensorRT Engine Plans (61 Engines)

πŸ“Š Engine Matrix Overview

Module Spatial Tile Temporal Frames ($4n+1$) Count Engine Size Naming Format
VAE Decoder 256Γ—256 5f, 21f, 25f, 29f, 33f, 37f, 41f, 45f, 49f, 53f, 57f, 61f, 65f, 69f, 73f, 77f, 81f, 85f, 89f, 93f, 97f 21 ~306–308 MB vae_decoder_tile_256_<N>f.rtxplan
VAE Decoder 512Γ—512 1f, 5f, 21f, 25f, 29f, 33f, 37f, 41f, 45f, 49f, 53f, 65f, 69f, 73f 14 ~310–320 MB vae_decoder_tile_512_<N>f.rtxplan
VAE Encoder 512Γ—512 1f, 5f, 17f, 21f, 25f, 29f, 33f, 37f, 41f, 45f, 49f, 53f, 57f, 61f, 65f, 69f, 73f, 77f, 81f, 85f, 89f, 93f, 97f 23 ~220–257 MB vae_encoder_<N>f_tile512.rtxplan
VAE Encoder 256Γ—256 105f, 165f, 185f 3 ~220–230 MB vae_encoder_<N>f_tile256.rtxplan
πŸ” Click to expand complete list of all 61 individual engine files

1. VAE Decoder β€” Tile 256 (21 engines)

vae_decoder_tile_256_5f.rtxplan     vae_decoder_tile_256_41f.rtxplan    vae_decoder_tile_256_73f.rtxplan
vae_decoder_tile_256_21f.rtxplan    vae_decoder_tile_256_45f.rtxplan    vae_decoder_tile_256_77f.rtxplan
vae_decoder_tile_256_25f.rtxplan    vae_decoder_tile_256_49f.rtxplan    vae_decoder_tile_256_81f.rtxplan
vae_decoder_tile_256_29f.rtxplan    vae_decoder_tile_256_53f.rtxplan    vae_decoder_tile_256_85f.rtxplan
vae_decoder_tile_256_33f.rtxplan    vae_decoder_tile_256_57f.rtxplan    vae_decoder_tile_256_89f.rtxplan
vae_decoder_tile_256_37f.rtxplan    vae_decoder_tile_256_61f.rtxplan    vae_decoder_tile_256_93f.rtxplan
                                    vae_decoder_tile_256_65f.rtxplan    vae_decoder_tile_256_97f.rtxplan
                                    vae_decoder_tile_256_69f.rtxplan

2. VAE Decoder β€” Tile 512 (14 engines)

vae_decoder_tile_512_1f.rtxplan     vae_decoder_tile_512_29f.rtxplan    vae_decoder_tile_512_49f.rtxplan
vae_decoder_tile_512_5f.rtxplan     vae_decoder_tile_512_33f.rtxplan    vae_decoder_tile_512_53f.rtxplan
vae_decoder_tile_512_21f.rtxplan    vae_decoder_tile_512_37f.rtxplan    vae_decoder_tile_512_65f.rtxplan
vae_decoder_tile_512_25f.rtxplan    vae_decoder_tile_512_41f.rtxplan    vae_decoder_tile_512_69f.rtxplan
                                    vae_decoder_tile_512_45f.rtxplan    vae_decoder_tile_512_73f.rtxplan

3. VAE Encoder β€” Tile 512 (23 engines)

vae_encoder_1f_tile512.rtxplan      vae_encoder_37f_tile512.rtxplan     vae_encoder_69f_tile512.rtxplan
vae_encoder_5f_tile512.rtxplan      vae_encoder_41f_tile512.rtxplan     vae_encoder_73f_tile512.rtxplan
vae_encoder_17f_tile512.rtxplan     vae_encoder_45f_tile512.rtxplan     vae_encoder_77f_tile512.rtxplan
vae_encoder_21f_tile512.rtxplan     vae_encoder_49f_tile512.rtxplan     vae_encoder_81f_tile512.rtxplan
vae_encoder_25f_tile512.rtxplan     vae_encoder_53f_tile512.rtxplan     vae_encoder_85f_tile512.rtxplan
vae_encoder_29f_tile512.rtxplan     vae_encoder_57f_tile512.rtxplan     vae_encoder_89f_tile512.rtxplan
vae_encoder_33f_tile512.rtxplan     vae_encoder_61f_tile512.rtxplan     vae_encoder_93f_tile512.rtxplan
                                    vae_encoder_65f_tile512.rtxplan     vae_encoder_97f_tile512.rtxplan

4. VAE Encoder β€” Tile 256 (3 engines)

vae_encoder_105f_tile256.rtxplan    vae_encoder_165f_tile256.rtxplan    vae_encoder_185f_tile256.rtxplan

Note on 4n+1 Temporal Structure: SeedVR2 utilizes a 3D causal temporal downsampling VAE architecture. The temporal latent dimension corresponds to $(T - 1) / 4 + 1$. Consequently, video frame sequences are compiled and processed in $4n + 1$ frame blocks. Single image inference uses the dedicated 1f engine.


πŸ› οΈ Architectural & Performance Advantages

  1. Native Blackwell Kernel Tuning: Built and tuned with the TensorRT Blackwell compilation pipeline, fully utilizing Blackwell 5th-Generation Tensor Cores, enhanced memory bandwidth, and fused convolution-activation operators across the entire RTX 50-Series lineup (RTX 5090, RTX 5080, RTX 5070, RTX 5060, RTX 5050).

  2. Decoupled Dual-Phase VAE Configuration: Enables independent selection of engine frame sizes for encoding and decoding (e.g. encode with 81f tile512 for maximum patch fidelity, and decode with 65f tile256 for optimal VRAM headroom).

  3. Dynamic 1-Shot Pad & Crop Execution: Short or non-matching batch lengths are dynamically padded to the nearest valid $4n+1$ sequence, executed in a single high-throughput TensorRT pass, and cropped back with pixel precision without falling back to slow FP16 CPU/PyTorch loops.

  4. Studio-Grade Memory & Execution Safety:

    • ExecutionContext Lock & Stream Synchronization: Dedicated mutex locks (_ENCODE_LOCK, _DECODE_LOCK) and explicit stream synchronization eliminate race conditions and memory corruption (black tiles) during concurrent or chained executions.
    • Deterministic Tri-Partite Memory Cleanup: Immediate release of intermediate tensors (del, gc.collect(), torch.cuda.empty_cache()) prevents progressive VRAM fragmentation across multi-batch sequences.

πŸš€ Usage in ComfyUI

These TensorRT engine plans (.rtxplan) are directly loaded and executed via the dedicated custom node:

1. Installation

Clone the custom node repository into your ComfyUI custom_nodes/ directory:

cd ComfyUI/custom_nodes
git clone https://github.com/ussoewwin/ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT-Decoder.git

2. Engine Placement

Place the downloaded .rtxplan engine files into either of the following supported directories:

ComfyUI/custom_nodes/ComfyUI-SeedVR2-VideoUpscaler-with-TensorRT-Decoder/tensorrt_backend/artifacts/

or

ComfyUI/models/tensorrt/seedvr2/

Upon launching ComfyUI, all detected .rtxplan files will automatically populate the engine_frames dropdown list of the respective loader nodes:

  • vae_decoder_tile_*_*f.rtxplan populates SeedVR2 Load TensorRT VAE Decoder
  • vae_encoder_*f_tile*.rtxplan populates SeedVR2 Load TensorRT VAE Encoder

3. Node Integration

Connect both loader nodes directly to the SeedVR2 Video Upscaler node:

  • Connect the SEEDVR2_VAE output of SeedVR2 Load TensorRT VAE Encoder to the vae_encode input.
  • Connect the SEEDVR2_VAE output of SeedVR2 Load TensorRT VAE Decoder to the vae_decode input.
+------------------------------------+
| SeedVR2 Load TensorRT VAE Encoder  |
|  [engine_frames: auto / 81f / ...] |--(SEEDVR2_VAE)---+
+------------------------------------+                 |
                                                       v [vae_encode]
                                           +----------------------------+
                                           |  SeedVR2 Video Upscaler    |
                                           +----------------------------+
                                                       ^ [vae_decode]
+------------------------------------+                 |
| SeedVR2 Load TensorRT VAE Decoder  |                 |
|  [engine_frames: auto / 65f / ...] |--(SEEDVR2_VAE)---+
+------------------------------------+

4. Workflow Examples

Complete Video Upscaler Workflow (TensorRT VAE & Quantized Models)

Workflow JSON: example_workflows/SeedVR2_tensorrt_decode.json

Usage Example - Full Workflow

TensorRT VAE Decoder Node Integration

Connect the TRT_VAE output from SeedVR2 Load TensorRT VAE Decoder directly into the SeedVR2 Video Upscaler node:

Usage Example - TensorRT VAE Decoder


πŸ”§ Building Custom Frame Engines on Demand

The custom node repository also includes a dedicated SeedVR2 Build TensorRT VAE Engines node to compile custom frame sizes directly inside ComfyUI:

  • Supported Frames: Any custom frame count following the $4n+1$ format (e.g. 101f, 185f, 205f).
  • Supported Tile Sizes: 256 (mandatory for Decoder; optimal VRAM for long Encoder sequences) and 512 (recommended for standard Encoder sequences).
  • Supported Targets: decoder, encoder, or both.

Built engines are saved directly into tensorrt_backend/artifacts/ and become immediately selectable upon restarting ComfyUI.


πŸ“œ Credits & References

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support