Skip to content

Official Release: AMD Strix Point Native Optimization and Director St… - #796

Open
Cyb3rLab5 wants to merge 23 commits into
lllyasviel:mainfrom
Cyb3rLab5:main
Open

Official Release: AMD Strix Point Native Optimization and Director St…#796
Cyb3rLab5 wants to merge 23 commits into
lllyasviel:mainfrom
Cyb3rLab5:main

Conversation

@Cyb3rLab5

Copy link
Copy Markdown

…udio

Bobby Jackson and others added 23 commits December 18, 2025 01:07
Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
Cached the weight and bias tensors in `vae_decode_fake` based on device and dtype. This prevents re-allocation and conversion on every call, significantly reducing overhead in hot loops.

Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
Refactored `HunyuanVideoRotaryPosEmbed` to calculate Rotary Positional Embeddings (RoPE) iteratively over 1D components (T, H, W). PyTorch's zero-copy `.expand()` and views are then used to map these vectors into 3D. This eliminates massive redundant 3D grid creations (via `torch.meshgrid`) and sequential looping over batches.

Performance benchmark shows ~4.3x forward pass speedup for batch coordinate generations.

Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
…6079084948

⚡ Bolt: [performance improvement] 1D RoPE Expand
…858351843

⚡ Bolt: [performance improvement] Cache tensors in vae_decode_fake
Batch the torch.cuda.memory_stats queries to check only every 25 modules, significantly reducing CPU-GPU synchronization stalls.

Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
…head-55594135257129418

⚡ Bolt: Optimize VRAM polling overhead
…27942520

⚡ Bolt: [performance improvement]
Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
…09869604236974047

⚡ Bolt: Replace np.poly1d with unrolled native Python evaluation
Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
…792659738826133

⚡ Bolt: Vectorize cu_seqlens loop in framepack_transformer
- Replaced `torch.tensor([val] * bs).to(device)` with `torch.full((bs,), val, device=device)` in framepack pipeline.
- Replaced `torch.tensor([0.5], device=device)` with `torch.full((1,), 0.5, device=device)` in UniPC solver.
- Creating a tensor from a Python list forces a CPU allocation and a blocking CPU-GPU synchronization, which introduces a ~1.5x-2x overhead for small tensors compared to natively creating them on the device via `torch.full`.

Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
…-15565401181004026508

⚡ Bolt: Optimize implicit CPU-GPU synchronization during tensor initialization
- Shifted sum operation for `attention_mask` to happen *before* the tensor is transferred to the GPU in `framepack_helpers.py`.
- Added learning to `.jules/bolt.md` reflecting that using `.sum()` on a GPU tensor in PyTorch blocks host execution by syncing data over the PCIe bus, identical to `.item()`.

Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
…10562887828396246823

⚡ Bolt: [Remove CPU-GPU Sync on Device Tensors]
…C by using torch.stack

Co-authored-by: Cyb3rLab5 <224908985+Cyb3rLab5@users.noreply.github.com>
…198123106505

⚡ Bolt: [performance improvement] Avoid CPU-GPU sync in FlowMatchUniPC by using torch.stack
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant