Size the whole-frame profile for a 6 GB card
TensorRT plans its memory for the profile's largest shape, and up to a whole 6D frame with Best's border (4608 x 6656) it asked for 4.9-5.9 GB and would not build on the RTX 3050. The profile now ends at 4608 x 3328 (15 MP), tuned for 4160 x 3248, and the tiler cuts a 6D frame into two such tiles: 27 MP of work for 20 MP kept, against 49 MP in 1408 tiles. The engine's directory names the profile, so a later range never loads an engine built for this one.
This commit is contained in:
@@ -73,9 +73,11 @@ pub fn openvino_dir(cfg: &Config, bytes: &[u8], fp16: bool) -> PathBuf {
|
||||
/// shape out, and served one export's engine to another of the same graph
|
||||
/// with a different shape when the denoiser was first cut into pieces
|
||||
/// (2026-10-04) — the fixed 1408² denoiser and its any-size sibling are
|
||||
/// exactly that pair.
|
||||
/// exactly that pair. The profile's largest shape is in the name for the
|
||||
/// same reason: an engine built for one range is not the next one's.
|
||||
pub fn tensorrt_whole_dir(cfg: &Config, bytes: &[u8]) -> PathBuf {
|
||||
model_dir(cfg, "tensorrt-whole", bytes)
|
||||
let (h, w) = crate::WHOLE_FRAME_MAX;
|
||||
model_dir(cfg, &format!("tensorrt-whole-{h}x{w}"), bytes)
|
||||
}
|
||||
|
||||
/// `<cache>/<provider>/<runtime version>/<hash of the bytes>`: one per
|
||||
|
||||
Reference in New Issue
Block a user