Size the whole-frame profile for a 6 GB card
TensorRT plans its memory for the profile's largest shape, and up to a whole 6D frame with Best's border (4608 x 6656) it asked for 4.9-5.9 GB and would not build on the RTX 3050. The profile now ends at 4608 x 3328 (15 MP), tuned for 4160 x 3248, and the tiler cuts a 6D frame into two such tiles: 27 MP of work for 20 MP kept, against 49 MP in 1408 tiles. The engine's directory names the profile, so a later range never loads an engine built for this one.
This commit is contained in:
@@ -619,6 +619,21 @@ mod tests {
|
||||
let squares = plan(3648, 5472, 256, Sizes::Square(1408)).unwrap();
|
||||
assert_eq!(squares.grid, (5, 7));
|
||||
assert!(one.work() * 2 < squares.work());
|
||||
// The whole-frame engine's limit on a 6 GB card: two tiles, each
|
||||
// within it, and still under half the work of the 1408 squares.
|
||||
let halves = plan(
|
||||
3648,
|
||||
5472,
|
||||
256,
|
||||
Sizes::Any {
|
||||
align: 16,
|
||||
max: (4608, 3328),
|
||||
},
|
||||
)
|
||||
.unwrap();
|
||||
assert_eq!(halves.grid, (1, 2));
|
||||
assert_eq!((halves.rows, halves.cols), (4160, 3248));
|
||||
assert!(halves.work() * 2 < squares.work());
|
||||
// A limit no tile fits under.
|
||||
assert!(plan(
|
||||
3648,
|
||||
|
||||
@@ -73,9 +73,11 @@ pub fn openvino_dir(cfg: &Config, bytes: &[u8], fp16: bool) -> PathBuf {
|
||||
/// shape out, and served one export's engine to another of the same graph
|
||||
/// with a different shape when the denoiser was first cut into pieces
|
||||
/// (2026-10-04) — the fixed 1408² denoiser and its any-size sibling are
|
||||
/// exactly that pair.
|
||||
/// exactly that pair. The profile's largest shape is in the name for the
|
||||
/// same reason: an engine built for one range is not the next one's.
|
||||
pub fn tensorrt_whole_dir(cfg: &Config, bytes: &[u8]) -> PathBuf {
|
||||
model_dir(cfg, "tensorrt-whole", bytes)
|
||||
let (h, w) = crate::WHOLE_FRAME_MAX;
|
||||
model_dir(cfg, &format!("tensorrt-whole-{h}x{w}"), bytes)
|
||||
}
|
||||
|
||||
/// `<cache>/<provider>/<runtime version>/<hash of the bytes>`: one per
|
||||
|
||||
@@ -65,13 +65,18 @@ pub enum Role {
|
||||
|
||||
/// The largest input, rows × columns, a [`Role::WholeDenoiser`] session
|
||||
/// takes: TensorRT's optimisation profile is built up to it, and the tiler
|
||||
/// cuts a larger frame into tiles no bigger. A 20 MP 6D frame with Best's
|
||||
/// reflected border is 4160 × 5984.
|
||||
pub const WHOLE_FRAME_MAX: (usize, usize) = (4608, 6656);
|
||||
/// cuts a larger frame into the fewest tiles no bigger.
|
||||
///
|
||||
/// Sized for a 6 GB card. TensorRT plans its memory for the profile's
|
||||
/// largest shape, and at 4608 × 6656 (a whole 6D frame with Best's border
|
||||
/// and room to spare) it asked for 4.9–5.9 GB and could not build on the
|
||||
/// RTX 3050. At 15 MP a 6D frame is two tiles of 4160 × 3248: 27 MP of
|
||||
/// work for 20 MP kept, against 49 MP in 1408² tiles.
|
||||
pub const WHOLE_FRAME_MAX: (usize, usize) = (4608, 3328);
|
||||
|
||||
/// The input size TensorRT tunes a whole-frame engine for: the 6D's frame
|
||||
/// with Best's border, the frame the reference measurements are of.
|
||||
pub const WHOLE_FRAME_OPT: (usize, usize) = (4160, 5984);
|
||||
/// The input size TensorRT tunes a whole-frame engine for: half a 6D frame
|
||||
/// with Best's border, the tile the reference measurements run.
|
||||
pub const WHOLE_FRAME_OPT: (usize, usize) = (4160, 3248);
|
||||
|
||||
/// Whether the selected rung runs [`Role::WholeDenoiser`], and if so the
|
||||
/// largest input it takes. `None` means run the fixed tiles.
|
||||
|
||||
Reference in New Issue
Block a user