Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
aryan5v
/
FastVideo
Public
forked from
hao-ai-lab/FastVideo
Notifications
You must be signed in to change notification settings
Fork
0
Star
0
Code
Issues
1
Pull requests
27
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Actions
Projects
Security and quality
Insights
[perf] Enable cached H3 generation on RAM-limited RTX hosts
- #48
#48
Open
aryan5v
wants to merge 69 commits into
fasth3-rtx
aryan5v/FastVideo:fasth3-rtx
from
fasth3-rtx-4090-launch
aryan5v/FastVideo:fasth3-rtx-4090-launch
Copy head branch name to clipboard
Conversation
Commits
69
(69)
Checks
Files changed
Open
[perf] Enable cached H3 generation on RAM-limited RTX hosts
#48
aryan5v
wants to merge 69 commits into
fasth3-rtx
aryan5v/FastVideo:fasth3-rtx
from
fasth3-rtx-4090-launch
aryan5v/FastVideo:fasth3-rtx-4090-launch
Copy head branch name to clipboard
Commits
Commits on Oct 3, 2026
[feat]: load packed FastH3 NVFP4 and Comfy int8 VAE on Blackwell
Show description for 2d0ebf1
aryan5v
committed
2d0ebf1
View commit details
Copy full SHA for 2d0ebf1
Browse repository at this point
[feat]: add CompactH3 NVFP4 cookbook paths for Blackwell GPUs
Show description for c38b4e2
aryan5v
committed
c38b4e2
View commit details
Copy full SHA for c38b4e2
Browse repository at this point
[feat]: document CompactH3 NVFP4 on RTX 5090 and RTX PRO 6000
Show description for a068234
aryan5v
committed
a068234
View commit details
Copy full SHA for a068234
Browse repository at this point
[bugfix]: refuse FSDP NVFP4 retain and drop restating comments
Show description for d224e52
aryan5v
committed
d224e52
View commit details
Copy full SHA for d224e52
Browse repository at this point
[feat]: enable CompactH3 VAE compile on RTX PRO 6000
Show description for d078283
aryan5v
committed
d078283
View commit details
Copy full SHA for d078283
Browse repository at this point
[bugfix]: refuse zero-gate CompactH3 VSA
Show description for a766348
aryan5v
committed
a766348
View commit details
Copy full SHA for a766348
Browse repository at this point
[wip]: sm_120 H3 FP4 experiments: SP FP8 exchange path and Modal drivers
Show description for 1fb455d
aryan5v
committed
1fb455d
View commit details
Copy full SHA for 1fb455d
Browse repository at this point
[wip]: H3 single-GPU 5090/4090 shipping path
Show description for 0a4737c
aryan5v
committed
0a4737c
View commit details
Copy full SHA for 0a4737c
Browse repository at this point
[wip]: H3 4090 groundwork, step splice, encoder fallback on sm89
Show description for 4012c37
aryan5v
committed
4012c37
View commit details
Copy full SHA for 4012c37
Browse repository at this point
[feat]: H3 headline benchmark (480p 5 s, 768p 10 s) for GPU clusters and Modal
Show description for 4a72920
aryan5v
committed
4a72920
View commit details
Copy full SHA for 4a72920
Browse repository at this point
[bugfix]: ship app.py into the headline Modal image
aryan5v
committed
f0688b0
View commit details
Copy full SHA for f0688b0
Browse repository at this point
[bugfix]: ship app.py into the headline Modal image as local python source
aryan5v
committed
0a23563
View commit details
Copy full SHA for 0a23563
Browse repository at this point
[bugfix]: run the headline benchmark from app.py (one Modal app, no cross-module import)
aryan5v
committed
a87344c
View commit details
Copy full SHA for a87344c
Browse repository at this point
[bugfix]: import os in gpu_worker; apply pre-commit formatting; ignore local headline results
aryan5v
committed
edbe18b
View commit details
Copy full SHA for edbe18b
Browse repository at this point
[bugfix]: H3 FP4 attention: new _build_block_mask signature; share Q/K/V quantization only at unit scale
Show description for dd10fc7
aryan5v
committed
dd10fc7
View commit details
Copy full SHA for dd10fc7
Browse repository at this point
[bugfix]: address #45 review: int64 FP8 epilogue offsets, order-independent NVFP4/FP8 conversion, splice scope
Show description for a6faddd
aryan5v
committed
a6faddd
View commit details
Copy full SHA for a6faddd
Browse repository at this point
[misc]: headline Modal runs keep the full log on the volume and report error lines
aryan5v
committed
ccebd84
View commit details
Copy full SHA for ccebd84
Browse repository at this point
[bugfix]: restore import os in fsdp_load (dropped in the rebase onto main)
aryan5v
committed
54f335c
View commit details
Copy full SHA for 54f335c
Browse repository at this point
[feat]: headline benchmark: HEADLINE_VAE_PARALLEL=1 decodes VAE tiles on every GPU
aryan5v
committed
4d33717
View commit details
Copy full SHA for 4d33717
Browse repository at this point
[feat]: headline benchmark: engine/pipeline placement overrides for memory-limited GPUs
aryan5v
committed
5386585
View commit details
Copy full SHA for 5386585
Browse repository at this point
[feat]: headline Modal entrypoint takes extra env and a run tag
aryan5v
committed
cb79431
View commit details
Copy full SHA for cb79431
Browse repository at this point
[bugfix]: H3 pinned swaps use exact-size cudaHostRegister arenas
Show description for a97d23f
aryan5v
committed
a97d23f
View commit details
Copy full SHA for a97d23f
Browse repository at this point
[wip]: exact-size pinned arenas for H3 offload and 4090 benchmarks
aryan5v
committed
a216768
View commit details
Copy full SHA for a216768
Browse repository at this point
[wip]: document 4090 setup validation and reproducible baselines
aryan5v
committed
a491f15
View commit details
Copy full SHA for a491f15
Browse repository at this point
[wip]: record checkpoint revision and GPU driver in 4090 benchmark results
aryan5v
committed
4ae791c
View commit details
Copy full SHA for 4ae791c
Browse repository at this point
[wip]: record completed 480p baseline and stage summary tooling
aryan5v
committed
21f9899
View commit details
Copy full SHA for 21f9899
Browse repository at this point
[perf]: reduce H3 attention activation copies and share FP8 input quantization
aryan5v
committed
c716408
View commit details
Copy full SHA for c716408
Browse repository at this point
[wip]: prototype tile-64 INT8 QK and FP8 PV attention on sm89
aryan5v
committed
bd713dd
View commit details
Copy full SHA for bd713dd
Browse repository at this point
[perf]: retain offloaded H3 VAEs on host until their consuming stages
aryan5v
committed
5bf9804
View commit details
Copy full SHA for 5bf9804
Browse repository at this point
[perf]: add opt-in sm89 tile-64 BF16 and INT8 QK attention
aryan5v
committed
59e6946
View commit details
Copy full SHA for 59e6946
Browse repository at this point
[docs]: record cached 4090 timings and sm89 precision validation
aryan5v
committed
b3ab6e1
View commit details
Copy full SHA for b3ab6e1
Browse repository at this point
[wip]: validate tilewise FP8 values and dynamic probability scales
aryan5v
committed
994c220
View commit details
Copy full SHA for 994c220
Browse repository at this point
[docs]: record resident-block timings and encoder constraints
aryan5v
committed
6b18535
View commit details
Copy full SHA for 6b18535
Browse repository at this point
[docs]: record five-second 4090 timing and FP8 encoder footprint
aryan5v
committed
92fc26c
View commit details
Copy full SHA for 92fc26c
Browse repository at this point
[feat]: stream the H3 encoder for consumer VRAM limits
aryan5v
committed
d986589
View commit details
Copy full SHA for d986589
Browse repository at this point
[perf]: fuse serialized NVFP4 encoder weight dequantization
aryan5v
committed
887deaa
View commit details
Copy full SHA for 887deaa
Browse repository at this point
[perf]: release H3 fine-attention copies before the gated merge
aryan5v
committed
2aa19c4
View commit details
Copy full SHA for 2aa19c4
Browse repository at this point
[perf]: share VAE INT8 input preparation and avoid weight copies
aryan5v
committed
9c9f1ed
View commit details
Copy full SHA for 9c9f1ed
Browse repository at this point
[docs]: record 12 GiB and streamed 4090 results
aryan5v
committed
e457b68
View commit details
Copy full SHA for e457b68
Browse repository at this point
[perf]: read H3 INT8 sparse attention from existing layouts
aryan5v
committed
7a0d7d3
View commit details
Copy full SHA for 7a0d7d3
Browse repository at this point
[perf]: fuse H3 INT8 VAE scaling and bias without intermediates
aryan5v
committed
3c0668f
View commit details
Copy full SHA for 3c0668f
Browse repository at this point
[bugfix]: retain consumer attention environment parsing after core rebase
aryan5v
committed
a7f7ede
View commit details
Copy full SHA for a7f7ede
Browse repository at this point
[test]: adapt consumer VSA validation to packed segment metadata
aryan5v
committed
87b22a5
View commit details
Copy full SHA for 87b22a5
Browse repository at this point
[bench]: sample total GPU memory for consumer budget checks
aryan5v
committed
fb92af1
View commit details
Copy full SHA for fb92af1
Browse repository at this point
[docs]: record 42-second 4090 clips and consumer memory validation
aryan5v
committed
5ac85ad
View commit details
Copy full SHA for 5ac85ad
Browse repository at this point
[docs]: record Track B PR and strict 8 GiB budget limit
aryan5v
committed
98cdc7d
View commit details
Copy full SHA for 98cdc7d
Browse repository at this point
Commits on Oct 5, 2026
[docs]: FastH3 NVFP4 on RTX PRO 6000
Show description for 5a377ca
aryan5v
committed
5a377ca
View commit details
Copy full SHA for 5a377ca
Browse repository at this point
[bugfix]: validate sparse FP4 block lists before launch; restore the #44 sparse kernel tests
Show description for 3d727ff
aryan5v
committed
3d727ff
View commit details
Copy full SHA for 3d727ff
Browse repository at this point
[bugfix]: keep ModelOpt-quantized projections off the dense re-quantize path; converter parity test
Show description for 5efdd8b
aryan5v
committed
5efdd8b
View commit details
Copy full SHA for 5efdd8b
Browse repository at this point
[bugfix]: headline benchmarks record their decode mode; 4090 bench never reports unmeasured memory as 0 GiB
Show description for 74a190f
aryan5v
committed
74a190f
View commit details
Copy full SHA for 74a190f
Browse repository at this point
[misc]: merge upstream main into the FastH3 RTX release branch
Show description for 5db0cfa
aryan5v
committed
5db0cfa
View commit details
Copy full SHA for 5db0cfa
Browse repository at this point
[bugfix]: streamed NVFP4 encoder layers pass their FP4 scalars on the GEMM's device
Show description for 4784019
aryan5v
committed
4784019
View commit details
Copy full SHA for 4784019
Browse repository at this point
[bugfix]: preserve ModelOpt H3 activation calibration in NVFP4 export
Aryan Kumar
authored and
aryan5v
committed
7863e08
View commit details
Copy full SHA for 7863e08
Browse repository at this point
[bugfix]: declare NVFP4 static activation-scale cache attributes for mypy
aryan5v
committed
9ca4eba
View commit details
Copy full SHA for 9ca4eba
Browse repository at this point
[bugfix]: clone batched VAE tile output before the next CUDA-graph replay
Show description for 3c7e89c
aryan5v
committed
3c7e89c
View commit details
Copy full SHA for 3c7e89c
Browse repository at this point
[bugfix]: VSA zero-gate guard checks packed NVFP4 gates
Show description for f4df26c
aryan5v
committed
f4df26c
View commit details
Copy full SHA for f4df26c
Browse repository at this point
[bugfix]: sparse FP4 entry points validate block lists by default
Show description for 2964f15
aryan5v
committed
2964f15
View commit details
Copy full SHA for 2964f15
Browse repository at this point
[misc]: register FastH3 single-GPU switches in fastvideo.envs
Show description for 826b5d0
aryan5v
committed
826b5d0
View commit details
Copy full SHA for 826b5d0
Browse repository at this point
[test]: NVFP4 H3 encoder tests follow the pre-Blackwell de-quantized fallback
Show description for a59e842
aryan5v
committed
a59e842
View commit details
Copy full SHA for a59e842
Browse repository at this point
[misc]: drop the os imports layerwise offload no longer uses
aryan5v
committed
e9a35b2
View commit details
Copy full SHA for e9a35b2
Browse repository at this point
[bench]: allow the launch benchmark seed to be specified
aryan5v
committed
0453c08
View commit details
Copy full SHA for 0453c08
Browse repository at this point
[offload]: honor pageable host storage for layerwise models
aryan5v
committed
1d5d9c9
View commit details
Copy full SHA for 1d5d9c9
Browse repository at this point
[bench]: measure memory on cgroup v1 pods
aryan5v
committed
00bd0a1
View commit details
Copy full SHA for 00bd0a1
Browse repository at this point
[bugfix]: keep CPU-targeted checkpoint loads off the GPU
aryan5v
committed
659a484
View commit details
Copy full SHA for 659a484
Browse repository at this point
[bench]: preserve peak memory for failed generations
aryan5v
committed
2c0cd2a
View commit details
Copy full SHA for 2c0cd2a
Browse repository at this point
[bench]: build exact V2 modulation tables with provenance
aryan5v
committed
3c011ef
View commit details
Copy full SHA for 3c011ef
Browse repository at this point
[bench]: render a showcase prompt once without extra warmup clips
aryan5v
committed
1299c7d
View commit details
Copy full SHA for 1299c7d
Browse repository at this point
[perf]: retain mapped H3 encoder weights on pageable hosts
aryan5v
committed
73ee8c2
View commit details
Copy full SHA for 73ee8c2
Browse repository at this point
[bugfix]: preserve padded vocabulary checkpoint validation
aryan5v
committed
cda0579
View commit details
Copy full SHA for cda0579
Browse repository at this point
You can’t perform that action at this time.