Merge: answer Android's memory warnings, and stop reporting a lost root as an empty library
FR-PLAT-AND-5 in full, FR-PLAT-AND-2 in part -- the recovery is built and live for Nextcloud roots, the SAF cause it names does not exist yet. FR-PLAT-AND-4 and FR-PLAT-AND-6 are not here, both blocked behind the same gap: assemble-apk.sh compiles no Java, so the APK cannot carry a Service or a FileProvider. The container has JDK 17 and build-tools 36; the build step is what is missing. Verified: fmt, clippy --workspace --all-targets -D warnings, and 1043 tests across dr-catalog, dr-sync, dr-sync-folder, dr-sync-nextcloud, dr-plat and dr-ui. The aarch64 target was checked before the branch was finished but not after; no device was available. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -3066,6 +3066,33 @@ impl DevelopSession {
|
||||
Ok((image, rw, rh))
|
||||
}
|
||||
|
||||
/// TRACES: FR-PLAT-AND-5 | NFR-RES-1
|
||||
/// Give back the GPU memory this session is holding only to be fast.
|
||||
///
|
||||
/// The edit is untouched: the graph and its history are CPU-side by
|
||||
/// design (ARCH §6.1), so the photograph, the undo stack and the viewport
|
||||
/// all survive and the next frame simply costs what the first one did.
|
||||
///
|
||||
/// # What is not released, and what it is waiting on
|
||||
///
|
||||
/// The demosaiced source is the largest single allocation a session holds
|
||||
/// — a 24 MP frame is about 190 MB of `Rgba16Float` — and it is
|
||||
/// deliberately kept. Dropping it would need the session to be able to
|
||||
/// rebuild itself from the file, and rebuilding a session from a durable
|
||||
/// record is FR-PLAT-AND-3, which is not built. Freeing it now would not
|
||||
/// be an eviction; it would be closing the photograph without telling
|
||||
/// anyone. Likewise the subject distance fields and the segmentation map:
|
||||
/// each is guarded by a key recording what it was built from, and freeing
|
||||
/// one without invalidating its key is the failure `AdjustPass` documents
|
||||
/// under `colour_key`.
|
||||
///
|
||||
/// So this is the part of the GPU tier that can be given back and asked
|
||||
/// for again with no other machinery, which is exactly as far as an
|
||||
/// eviction should go.
|
||||
pub fn release_gpu_caches(&mut self) {
|
||||
self.adjust.release_caches();
|
||||
}
|
||||
|
||||
/// The displayed size, for sizing the viewport.
|
||||
///
|
||||
/// The *framed* size, not the sensor's: cropping and quarter turns change
|
||||
|
||||
@@ -105,6 +105,21 @@ impl IdentityController {
|
||||
fn clear_picks(&self) {
|
||||
self.picked.borrow_mut().clear();
|
||||
}
|
||||
|
||||
/// TRACES: FR-PLAT-AND-5 | NFR-RES-1
|
||||
/// Drop the decoded rail portraits.
|
||||
///
|
||||
/// The one in-memory image cache in this crate that is unbounded by
|
||||
/// anything but the library: one decoded portrait per person, kept for as
|
||||
/// long as the person exists. On a library with a few hundred named people
|
||||
/// that is worth tens of megabytes of nothing but a saved decode.
|
||||
///
|
||||
/// Costless to lose. `refresh` rebuilds any portrait it does not find, so
|
||||
/// the only consequence is the JPEG decode this cache exists to skip, and
|
||||
/// only for the people the rail is actually showing at the time.
|
||||
pub fn clear_covers(&self) {
|
||||
self.covers.borrow_mut().clear();
|
||||
}
|
||||
}
|
||||
|
||||
/// Push the people rail and the face grid into the window.
|
||||
|
||||
@@ -39,6 +39,7 @@ mod library_ui;
|
||||
#[cfg(live_style)]
|
||||
mod live_style;
|
||||
mod masks_ui;
|
||||
pub mod memory;
|
||||
mod net_runtime;
|
||||
mod peaking;
|
||||
mod preset_store;
|
||||
@@ -1012,6 +1013,22 @@ pub fn run(paths: Vec<PathBuf>) -> Result<()> {
|
||||
// faces are ticked for a split.
|
||||
let identity = std::rc::Rc::new(identity_ui::IdentityController::new());
|
||||
|
||||
// TRACES: FR-PLAT-AND-5
|
||||
// The thumbnail tier. Registered here, beside the thing it frees, so that
|
||||
// a controller which grows another cache is one line from offering it up.
|
||||
//
|
||||
// Weak, not strong: `run` returns when the window closes, and a registry
|
||||
// holding the last reference to a controller would keep it — and every
|
||||
// decoded portrait in it — alive past the interface it belonged to.
|
||||
{
|
||||
let identity = std::rc::Rc::downgrade(&identity);
|
||||
memory::evict_at(memory::Tier::Thumbnails, move || {
|
||||
if let Some(ctl) = identity.upgrade() {
|
||||
ctl.clear_covers();
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
// Launch screen: shown when there is nothing to display — no local paths
|
||||
// and no configured library. A user who has already signed in and chosen
|
||||
// a folder goes straight to their images (FR-NC-1).
|
||||
@@ -1434,6 +1451,32 @@ pub fn run(paths: Vec<PathBuf>) -> Result<()> {
|
||||
// The current develop session, if the file yielded sensor data.
|
||||
let session: Rc<RefCell<Option<DevelopSession>>> = Rc::new(RefCell::new(None));
|
||||
|
||||
// TRACES: FR-PLAT-AND-5
|
||||
// The GPU tier — the first thing given back under memory pressure, and on
|
||||
// Android the only thing given back merely for going into the background.
|
||||
//
|
||||
// `try_borrow_mut` rather than `borrow_mut`, and the miss is not an error
|
||||
// worth reporting. A memory warning can land in the middle of a render, at
|
||||
// which point the slot is already borrowed and freeing its textures under
|
||||
// the code drawing with them is not something to do politely — skipping is
|
||||
// correct, because the pass that is running will have finished by the time
|
||||
// the platform asks again, and a warning that has not been acted on is
|
||||
// always followed by another one.
|
||||
{
|
||||
let session = Rc::downgrade(&session);
|
||||
memory::evict_at(memory::Tier::Gpu, move || {
|
||||
let Some(session) = session.upgrade() else {
|
||||
return;
|
||||
};
|
||||
let Ok(mut slot) = session.try_borrow_mut() else {
|
||||
return;
|
||||
};
|
||||
if let Some(open) = slot.as_mut() {
|
||||
open.release_gpu_caches();
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
// TRACES: FR-DEV-6 | FR-CAT-8
|
||||
// The settings clipboard, and where the open image's edit is stored.
|
||||
//
|
||||
|
||||
+101
-5
@@ -80,7 +80,20 @@ pub enum ScanMessage {
|
||||
/// amount of string matching on the far side can reliably recover it.
|
||||
/// Without the flag a dead connection and a bad password produce the same
|
||||
/// banner, which sends the user to re-enter a credential that was fine.
|
||||
Failed { message: String, offline: bool },
|
||||
///
|
||||
/// `lost_root` is the same idea one step further out, and it is carried
|
||||
/// separately from `offline` rather than folded into it because the two
|
||||
/// end differently. An offline library comes back when the network does,
|
||||
/// with nothing asked of anyone; a library whose root cannot be opened
|
||||
/// comes back only when someone restores access to it — a share put back
|
||||
/// on the server, a drive plugged in, and in time a document tree granted
|
||||
/// again once one can be (FR-PLAT-AND-2). Both show the same grid of what
|
||||
/// is stored locally, and they must not offer the same explanation.
|
||||
Failed {
|
||||
message: String,
|
||||
offline: bool,
|
||||
lost_root: bool,
|
||||
},
|
||||
}
|
||||
|
||||
/// One decoded thumbnail, ready for the grid.
|
||||
@@ -1166,6 +1179,7 @@ pub fn spawn_scan(
|
||||
let _ = tx.send(ScanMessage::Failed {
|
||||
message: e.message,
|
||||
offline: e.offline,
|
||||
lost_root: e.lost_root,
|
||||
});
|
||||
}
|
||||
});
|
||||
@@ -1180,6 +1194,7 @@ pub fn spawn_scan(
|
||||
struct ScanFailure {
|
||||
message: String,
|
||||
offline: bool,
|
||||
lost_root: bool,
|
||||
}
|
||||
|
||||
impl ScanFailure {
|
||||
@@ -1189,6 +1204,7 @@ impl ScanFailure {
|
||||
Self {
|
||||
message: message.to_string(),
|
||||
offline: false,
|
||||
lost_root: false,
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1197,11 +1213,48 @@ impl From<dr_sync::RemoteError> for ScanFailure {
|
||||
fn from(e: dr_sync::RemoteError) -> Self {
|
||||
Self {
|
||||
offline: e.indicates_offline(),
|
||||
lost_root: e.indicates_lost_root(),
|
||||
message: e.to_string(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// TRACES: FR-PLAT-AND-2 | FR-CAT-9
|
||||
/// Record that a library can no longer be opened, without losing it.
|
||||
///
|
||||
/// Called on the worker, before the failure crosses the channel, because this
|
||||
/// is where the catalog handle is — and because the marking must be durable
|
||||
/// whether or not anyone is left to draw a banner. A process killed between
|
||||
/// the failure and the next launch must still come back knowing what it could
|
||||
/// not reach.
|
||||
///
|
||||
/// Nothing is deleted. Every rating, every edit and every row stays exactly
|
||||
/// where it was; what changes is that the images now say they are offline, so
|
||||
/// the grid can show them as held-not-here rather than as ordinary
|
||||
/// photographs whose thumbnails happen to be failing one at a time.
|
||||
///
|
||||
/// A root with no row yet is the first scan of a library that has never
|
||||
/// succeeded, and there is nothing to mark — the failure alone is the whole
|
||||
/// story, and the launch screen is where it is told.
|
||||
fn mark_library_offline(catalog: &Catalog, root: &str) {
|
||||
let conn = catalog.connection();
|
||||
let root_id: Option<i64> = conn
|
||||
.query_row(
|
||||
"SELECT id FROM roots WHERE label = ?1 AND kind = 'remote'",
|
||||
[root],
|
||||
|r| r.get(0),
|
||||
)
|
||||
.ok();
|
||||
let Some(root_id) = root_id else {
|
||||
log::info!("library {root} has no catalog root yet; nothing to mark offline");
|
||||
return;
|
||||
};
|
||||
match dr_catalog::mark_root_offline(conn, dr_types::RootId(root_id as u64)) {
|
||||
Ok(()) => log::warn!("library {root} is unreachable; its images are marked offline"),
|
||||
Err(e) => log::error!("could not mark {root} offline: {e}"),
|
||||
}
|
||||
}
|
||||
|
||||
fn run_scan(
|
||||
tx: &Sender<ScanMessage>,
|
||||
conn: Connection,
|
||||
@@ -1219,21 +1272,50 @@ fn run_scan(
|
||||
let rt = crate::net_runtime::build().map_err(ScanFailure::local)?;
|
||||
|
||||
rt.block_on(async {
|
||||
let backend = crate::remote::connect(&conn).map_err(ScanFailure::local)?;
|
||||
// TRACES: FR-PLAT-AND-2 | FR-CAT-9
|
||||
// Classified rather than flattened to a local failure, because the
|
||||
// removed-card case never gets as far as a request: the folder
|
||||
// connector checks its root when it is constructed, so a library on an
|
||||
// ejected card fails here and not in the walk. Reported as an ordinary
|
||||
// error it left the grid showing a healthy library of images that
|
||||
// could no longer be opened, one silent thumbnail failure at a time.
|
||||
let backend = match crate::remote::connect(&conn) {
|
||||
Ok(b) => b,
|
||||
Err(e) => {
|
||||
if e.indicates_lost_root() {
|
||||
mark_library_offline(&catalog, &root);
|
||||
}
|
||||
return Err(e.into());
|
||||
}
|
||||
};
|
||||
|
||||
// Stored folder ETags, so an unchanged subtree is skipped whole. On a
|
||||
// first run this is empty and the walk is complete; on every run after
|
||||
// it is what keeps cost proportional to what changed (ARCH §8.4).
|
||||
let known = load_folder_etags(&catalog, &root);
|
||||
|
||||
let result = dr_sync::scan(&*backend, &RemotePath::new(&root), &filter, &known, |p| {
|
||||
let scanned = dr_sync::scan(&*backend, &RemotePath::new(&root), &filter, &known, |p| {
|
||||
let _ = tx.send(ScanMessage::Progress {
|
||||
directories: p.directories_listed,
|
||||
pruned: p.directories_pruned,
|
||||
images: p.images_found,
|
||||
});
|
||||
})
|
||||
.await?;
|
||||
.await;
|
||||
|
||||
// TRACES: FR-PLAT-AND-2 | FR-CAT-9
|
||||
// Written before the failure is reported, not after: the banner is a
|
||||
// consequence of the catalog state and not the other way round, and a
|
||||
// process that dies between the two must come back knowing.
|
||||
let result = match scanned {
|
||||
Ok(r) => r,
|
||||
Err(e) => {
|
||||
if e.indicates_lost_root() {
|
||||
mark_library_offline(&catalog, &root);
|
||||
}
|
||||
return Err(e.into());
|
||||
}
|
||||
};
|
||||
|
||||
persist(&catalog, &root, &result).map_err(ScanFailure::local)?;
|
||||
|
||||
@@ -1326,13 +1408,27 @@ fn persist(
|
||||
.ok()
|
||||
});
|
||||
|
||||
// TRACES: FR-PLAT-AND-2 | FR-CAT-9
|
||||
// The `availability` arm is what ends an offline library, and it does
|
||||
// it one photograph at a time. 3 is `Availability::Offline` and 0 is
|
||||
// `MetadataOnly`, the same code this statement inserts new rows with —
|
||||
// so a row that was marked offline when the root became unreachable is
|
||||
// returned to exactly the state a fresh scan would have given it, and
|
||||
// a row that was never marked is not touched at all.
|
||||
//
|
||||
// Conditional rather than a blanket reset for the same reason
|
||||
// `dr_catalog::walk` restores per file rather than per root: the only
|
||||
// thing that may clear "I could not reach this" is having reached it,
|
||||
// and this statement runs precisely once per file the scan listed.
|
||||
tx.execute(
|
||||
"INSERT INTO images(root_id, folder_id, source_ref, format, file_size,
|
||||
availability, metadata_state, added_at)
|
||||
VALUES (?1, ?2, ?3, ?4, ?5, 0, 1, ?6)
|
||||
ON CONFLICT(root_id, source_ref) DO UPDATE SET
|
||||
file_size = excluded.file_size,
|
||||
folder_id = excluded.folder_id",
|
||||
folder_id = excluded.folder_id,
|
||||
availability = CASE WHEN images.availability = 3
|
||||
THEN 0 ELSE images.availability END",
|
||||
rusqlite::params![
|
||||
root_id,
|
||||
folder_id,
|
||||
|
||||
+86
-12
@@ -270,6 +270,22 @@ pub struct LibraryController {
|
||||
/// judgement, and carrying the old one over would report a server down
|
||||
/// that was never contacted.
|
||||
reachability: RefCell<dr_sync::Reachability>,
|
||||
/// TRACES: FR-PLAT-AND-2 | FR-CAT-9
|
||||
/// Why the library folder itself could not be opened, if it could not.
|
||||
///
|
||||
/// Beside [`Self::reachability`] rather than inside it, because
|
||||
/// `dr_sync::Reachability` models *the server*, and it is deliberately
|
||||
/// unmoved by a refusal — a forbidden file must not report the network as
|
||||
/// down (see its own tests). A revoked tree grant is a refusal, so folding
|
||||
/// it in would either break that rule or need an exception carved through
|
||||
/// it.
|
||||
///
|
||||
/// Set only by a scan that failed at the root, and cleared only by one
|
||||
/// that succeeded. Both states drive the same banner as being offline
|
||||
/// does, because what the user can do is the same — carry on with what is
|
||||
/// stored on the device — but the sentence under it is different, and so
|
||||
/// is what will end it.
|
||||
root_lost: RefCell<Option<String>>,
|
||||
/// TRACES: FR-NC-6a
|
||||
/// Drains the pin downloader. Held so a second pin replaces the timer
|
||||
/// rather than leaving two draining the same finished channel.
|
||||
@@ -372,6 +388,7 @@ impl LibraryController {
|
||||
sidecar_timer: RefCell::new(None),
|
||||
generation: std::cell::Cell::new(0),
|
||||
reachability: RefCell::new(dr_sync::Reachability::new()),
|
||||
root_lost: RefCell::new(None),
|
||||
outbox_timer: RefCell::new(None),
|
||||
outbox_maybe_dirty: std::cell::Cell::new(true),
|
||||
geometry_timer: RefCell::new(None),
|
||||
@@ -443,10 +460,15 @@ impl LibraryController {
|
||||
}
|
||||
}
|
||||
|
||||
/// TRACES: FR-CAT-9
|
||||
/// Whether the app currently believes the server is unreachable.
|
||||
/// TRACES: FR-CAT-9 | FR-PLAT-AND-2
|
||||
/// Whether the library cannot be reached, for either of the two reasons.
|
||||
///
|
||||
/// One answer rather than two because every caller asks it for the same
|
||||
/// purpose: to decide whether starting a transfer is worth attempting.
|
||||
/// A revoked grant fails that question exactly as a dead network does, and
|
||||
/// a sync started against it would spend its retries proving it.
|
||||
pub fn is_offline(&self) -> bool {
|
||||
self.reachability.borrow().is_offline()
|
||||
self.reachability.borrow().is_offline() || self.root_lost.borrow().is_some()
|
||||
}
|
||||
|
||||
/// Whether the grid is narrowed to locally-stored originals.
|
||||
@@ -941,6 +963,17 @@ fn drain_scan(
|
||||
{
|
||||
log::info!("back online");
|
||||
}
|
||||
// TRACES: FR-PLAT-AND-2
|
||||
// And it is the only evidence that clears a lost root,
|
||||
// for the same reason: the walk began by listing the
|
||||
// root, so a scan that finished is a root that opened.
|
||||
// The rows it marked offline are restored one at a
|
||||
// time by `library::persist`, as each file is listed
|
||||
// again — this only stops the banner claiming what is
|
||||
// no longer true.
|
||||
if ctl.root_lost.borrow_mut().take().is_some() {
|
||||
log::info!("library folder is readable again");
|
||||
}
|
||||
refresh_offline(&w, ctl);
|
||||
|
||||
// An incremental rescan lists almost nothing, so
|
||||
@@ -995,7 +1028,11 @@ fn drain_scan(
|
||||
stop(&ctl.scan_timer);
|
||||
return;
|
||||
}
|
||||
ScanMessage::Failed { message, offline } => {
|
||||
ScanMessage::Failed {
|
||||
message,
|
||||
offline,
|
||||
lost_root,
|
||||
} => {
|
||||
log::warn!("scan failed: {message}");
|
||||
w.set_library_scanning(false);
|
||||
// Recorded as a failure even where it is only the
|
||||
@@ -1004,7 +1041,26 @@ fn drain_scan(
|
||||
// stopped because of it.
|
||||
job.fail(message.clone());
|
||||
|
||||
if offline {
|
||||
if lost_root {
|
||||
// TRACES: FR-PLAT-AND-2 | FR-CAT-9
|
||||
// The library folder itself could not be opened —
|
||||
// a share withdrawn, an unplugged drive, and in
|
||||
// time a revoked document-tree grant. The worker
|
||||
// has already marked every row under this root
|
||||
// offline and deleted none of them; this is the
|
||||
// half the user sees.
|
||||
//
|
||||
// Tested first because it is also true that the
|
||||
// library is unreachable, and the generic answer
|
||||
// would be reached first and be less useful.
|
||||
*ctl.root_lost.borrow_mut() = Some(message);
|
||||
refresh_offline(&w, ctl);
|
||||
// Same reason as the offline arm below: without
|
||||
// this a launch that began with a revoked grant
|
||||
// shows an empty grid, which is the one impression
|
||||
// this whole path exists to avoid.
|
||||
open_catalog_for_offline(&w, ctl, &catalog_path, &coll_ctl);
|
||||
} else if offline {
|
||||
// Not an error state. The catalog from the last
|
||||
// successful scan is still on disk and still
|
||||
// accurate for everything already indexed, so the
|
||||
@@ -1585,17 +1641,35 @@ fn scope_is_pinned(catalog: &Catalog, images: &[dr_types::ImageId]) -> bool {
|
||||
/// which is what keeps it testable without a display server.
|
||||
fn refresh_offline(window: &AppWindow, ctl: &Rc<LibraryController>) {
|
||||
let reach = ctl.reachability.borrow();
|
||||
let offline = reach.is_offline();
|
||||
// TRACES: FR-PLAT-AND-2
|
||||
// A lost root wins over a dead network, and does so even when both are
|
||||
// true — which is the ordinary case, since the scan that discovered the
|
||||
// grant was gone was also the last request the app made. Reported the
|
||||
// other way round the user is told to wait for a connection that is
|
||||
// working, and the thing that would actually fix it is never mentioned.
|
||||
let lost = ctl.root_lost.borrow();
|
||||
let offline = reach.is_offline() || lost.is_some();
|
||||
|
||||
window.set_library_offline(offline);
|
||||
window.set_library_offline_reason(reach.reason().unwrap_or_default().into());
|
||||
window.set_library_offline_reason(match lost.as_deref() {
|
||||
Some(why) => why.into(),
|
||||
None => reach.reason().unwrap_or_default().into(),
|
||||
});
|
||||
window.set_library_offline_since(
|
||||
reach
|
||||
.offline_for(std::time::Instant::now())
|
||||
.map(describe_duration)
|
||||
.unwrap_or_default()
|
||||
.into(),
|
||||
// A duration is what a network outage has and a revoked permission
|
||||
// does not: "for 4 minutes" invites waiting, and waiting is precisely
|
||||
// what will not help here.
|
||||
if lost.is_some() {
|
||||
slint::SharedString::default()
|
||||
} else {
|
||||
reach
|
||||
.offline_for(std::time::Instant::now())
|
||||
.map(describe_duration)
|
||||
.unwrap_or_default()
|
||||
.into()
|
||||
},
|
||||
);
|
||||
drop(lost);
|
||||
|
||||
// A stale scan error under an offline banner reports one problem twice.
|
||||
if offline {
|
||||
|
||||
@@ -0,0 +1,276 @@
|
||||
//! TRACES: FR-PLAT-AND-5 | NFR-RES-1 | FR-NC-6b
|
||||
//! Giving memory back when the platform asks for it.
|
||||
//!
|
||||
//! Android kills the process that will not shrink. It does not negotiate and
|
||||
//! it does not warn twice, and the app it kills is the one holding the most —
|
||||
//! which, on a photo editor, is always this one. So the question this module
|
||||
//! answers is not "how much can be freed" but "in what order", because the
|
||||
//! caches differ enormously in what losing them costs.
|
||||
//!
|
||||
//! # The order, and why it is that order
|
||||
//!
|
||||
//! FR-PLAT-AND-5 states it: GPU tiles first, then proxies, then thumbnails.
|
||||
//! Read as a rule rather than a list, it is *cheapest to rebuild goes first* —
|
||||
//! a GPU allocation is remade from data already in memory, a proxy is remade
|
||||
//! from a file already on disk, and a thumbnail may cost a network fetch.
|
||||
//! [`Tier`] is that order written down where the code can be held to it, so
|
||||
//! adding a cache means choosing its tier rather than choosing its position in
|
||||
//! a hand-maintained sequence.
|
||||
//!
|
||||
//! What each tier actually reaches in this build is documented on the variant,
|
||||
//! including where it reaches nothing yet. An empty tier is worth keeping
|
||||
//! visible: it says the order is complete and the coverage is not.
|
||||
//!
|
||||
//! # Why this is a registry rather than a function that frees things
|
||||
//!
|
||||
//! Every cache worth evicting lives behind an `Rc<RefCell<…>>` owned by a
|
||||
//! local in [`crate::run`], which is a two-thousand-line function whose
|
||||
//! callbacks each hold their own handle. There is no central object to reach
|
||||
//! them through, and inventing one to serve eviction alone would be a large
|
||||
//! change to how the interface is wired for a small change in what it does.
|
||||
//!
|
||||
//! So `run` hands this module a closure per cache as it builds each one, and
|
||||
//! this module owns only the ordering. The registration is next to the thing
|
||||
//! being registered, which is also the property that keeps it honest: a cache
|
||||
//! added later is one line away from being evictable, and a cache removed
|
||||
//! takes its sink with it.
|
||||
//!
|
||||
//! # Everything here is single-threaded, and that is not a limitation
|
||||
//!
|
||||
//! The registry is a `thread_local`, holding `Fn()` rather than `Fn() + Send`,
|
||||
//! because the pressure signal already arrives on the thread that owns the
|
||||
//! caches. Slint's Android backend calls the event listener from inside
|
||||
//! `poll_events`, which runs on the same thread as the event loop, which is
|
||||
//! the thread `run` built everything on. Marshalling through
|
||||
//! `invoke_from_event_loop` would add a hop and a lifetime question to solve a
|
||||
//! problem that does not exist — and would arrive *after* the moment the
|
||||
//! system asked, which for a memory warning is the one thing that matters.
|
||||
//!
|
||||
//! Anything reached from a worker thread — the thumbnail store, the original
|
||||
//! cache — is on disk and bounded by its own budget (NFR-RES-4), and is not
|
||||
//! what a memory warning is about.
|
||||
|
||||
use std::cell::RefCell;
|
||||
|
||||
/// How hard the platform is asking.
|
||||
///
|
||||
/// Two levels rather than Android's eight, because two is what the platform
|
||||
/// actually delivers to this app. `ComponentCallbacks2.onTrimMemory` and its
|
||||
/// `TRIM_MEMORY_*` grades are a Java callback on an `Activity` or
|
||||
/// `Application`; a `NativeActivity` receives only `ANativeActivityCallbacks`,
|
||||
/// whose memory callback is the ungraded `onLowMemory` — which is what
|
||||
/// android-activity surfaces as `MainEvent::LowMemory`. Modelling grades the
|
||||
/// entry point cannot observe would be modelling a wish.
|
||||
///
|
||||
/// [`Self::UiHidden`] recovers the one distinction that *is* observable and is
|
||||
/// worth acting on, because it is the cheapest moment to give memory back:
|
||||
/// nothing is on screen, so nothing that is freed has to be drawn again before
|
||||
/// the user notices. It corresponds to `TRIM_MEMORY_UI_HIDDEN` in intent and
|
||||
/// is derived from the activity being stopped rather than from a memory
|
||||
/// warning at all.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum Level {
|
||||
/// The app is no longer on screen. Free what only a visible window needs.
|
||||
UiHidden,
|
||||
/// The system says it is short of memory. Free everything that can be
|
||||
/// rebuilt.
|
||||
Critical,
|
||||
}
|
||||
|
||||
/// What a cache costs to lose, as an order.
|
||||
///
|
||||
/// Declared in eviction order and iterated in declaration order by
|
||||
/// [`Tier::ORDER`], so the sequence FR-PLAT-AND-5 specifies is a property of
|
||||
/// this type rather than of each call site.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub enum Tier {
|
||||
/// GPU allocations that are rebuilt from data the process still holds.
|
||||
///
|
||||
/// The develop session's compiled pipelines, its detail intermediates and
|
||||
/// its output textures. Rebuilt by the next render from the demosaiced
|
||||
/// source, which is still resident — see
|
||||
/// [`DevelopSession::release_gpu_caches`](crate::DevelopSession::release_gpu_caches)
|
||||
/// for what is deliberately kept and what that is waiting on.
|
||||
///
|
||||
/// First because it is both the largest evictable pool on a mobile GPU and
|
||||
/// the cheapest to refill: no I/O, no network, one frame's work.
|
||||
Gpu,
|
||||
/// Decoded image data rebuilt by reading a file again.
|
||||
///
|
||||
/// **Nothing registers here in this build, and the tier is kept anyway.**
|
||||
/// There is no in-memory proxy cache: the only decoded full-size frame in
|
||||
/// the process is the open develop session's, which belongs to
|
||||
/// [`Tier::Gpu`] and cannot be dropped until a session can be rebuilt from
|
||||
/// a durable record (FR-PLAT-AND-3). The on-disk original cache is a
|
||||
/// different thing wearing the same word — freeing disk relieves no memory
|
||||
/// pressure, and it already has a budget and an LRU of its own
|
||||
/// (`dr_catalog::Cache`, NFR-RES-4).
|
||||
Proxies,
|
||||
/// Decoded thumbnails, rebuilt by decoding a stored JPEG again — or, at
|
||||
/// worst, by fetching one.
|
||||
///
|
||||
/// Last because this is the tier a user sees losing: an evicted portrait
|
||||
/// is a rail that redraws, and an evicted grid cell is a photograph that
|
||||
/// greys out and comes back.
|
||||
Thumbnails,
|
||||
}
|
||||
|
||||
impl Tier {
|
||||
/// The eviction order, in one place.
|
||||
pub const ORDER: [Tier; 3] = [Tier::Gpu, Tier::Proxies, Tier::Thumbnails];
|
||||
|
||||
/// Whether this tier is given up at this level of pressure.
|
||||
///
|
||||
/// Hiding the window frees the GPU tier and nothing else. That is not
|
||||
/// caution about the rest — it is that a backgrounded app has no window to
|
||||
/// draw and therefore no use at all for a render pipeline, while its
|
||||
/// thumbnails are exactly what the user will be looking at half a second
|
||||
/// after they come back. Under [`Level::Critical`] the process is being
|
||||
/// measured against being killed, and a slow return beats no return.
|
||||
fn evicted_at(self, level: Level) -> bool {
|
||||
match level {
|
||||
Level::UiHidden => matches!(self, Tier::Gpu),
|
||||
Level::Critical => true,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A cache that has offered itself up, and the tier it goes in.
|
||||
///
|
||||
/// Named because the registry is a `Vec` of these and the nested type is hard
|
||||
/// to read at the use site rather than because either half means anything on
|
||||
/// its own.
|
||||
type Sink = (Tier, Box<dyn Fn()>);
|
||||
|
||||
thread_local! {
|
||||
/// Registered sinks, in the order they were registered within a tier.
|
||||
///
|
||||
/// Within a tier the order is registration order and nothing depends on
|
||||
/// it; between tiers it is [`Tier::ORDER`], which everything depends on.
|
||||
static SINKS: RefCell<Vec<Sink>> = const { RefCell::new(Vec::new()) };
|
||||
}
|
||||
|
||||
/// Offer a cache up for eviction at `tier`.
|
||||
///
|
||||
/// Called as each cache is built, so that the registration reads next to the
|
||||
/// thing it is about. The closure is kept for the life of the thread; it must
|
||||
/// therefore hold weak or shared handles rather than borrow anything, which is
|
||||
/// the natural shape here because everything it can reach is already an `Rc`.
|
||||
pub(crate) fn evict_at(tier: Tier, sink: impl Fn() + 'static) {
|
||||
SINKS.with_borrow_mut(|sinks| sinks.push((tier, Box::new(sink))));
|
||||
}
|
||||
|
||||
/// TRACES: FR-PLAT-AND-5
|
||||
/// Give memory back, in [`Tier::ORDER`], as far down as `level` calls for.
|
||||
///
|
||||
/// Safe to call when nothing is registered — before the window is built, or on
|
||||
/// a platform that never asks — in which case it does nothing at all.
|
||||
///
|
||||
/// The registry is taken out of the cell for the duration rather than borrowed
|
||||
/// across the calls. A sink runs arbitrary interface code, and interface code
|
||||
/// that registered another cache, or called this again, would otherwise meet a
|
||||
/// `RefCell` it had already borrowed and abort the process. Freeing memory is
|
||||
/// the wrong moment to be brittle about re-entry.
|
||||
pub fn relieve(level: Level) {
|
||||
let taken: Vec<(Tier, Box<dyn Fn()>)> = SINKS.with_borrow_mut(std::mem::take);
|
||||
let mut run = 0usize;
|
||||
for tier in Tier::ORDER {
|
||||
if !tier.evicted_at(level) {
|
||||
continue;
|
||||
}
|
||||
for (t, sink) in &taken {
|
||||
if *t == tier {
|
||||
sink();
|
||||
run += 1;
|
||||
}
|
||||
}
|
||||
}
|
||||
// Put them back, keeping anything a sink registered while it ran — after,
|
||||
// so the order within a tier stays registration order.
|
||||
SINKS.with_borrow_mut(|sinks| {
|
||||
let added = std::mem::replace(sinks, taken);
|
||||
sinks.extend(added);
|
||||
});
|
||||
log::info!("memory pressure ({level:?}): ran {run} eviction(s)");
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use std::rc::Rc;
|
||||
|
||||
/// Registers one sink per tier, backwards, and hands back what they saw.
|
||||
fn recorder() -> Rc<RefCell<Vec<Tier>>> {
|
||||
let seen = Rc::new(RefCell::new(Vec::new()));
|
||||
for tier in [Tier::Thumbnails, Tier::Proxies, Tier::Gpu] {
|
||||
let seen = seen.clone();
|
||||
evict_at(tier, move || seen.borrow_mut().push(tier));
|
||||
}
|
||||
seen
|
||||
}
|
||||
|
||||
fn reset() {
|
||||
SINKS.with_borrow_mut(|s| s.clear());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn eviction_runs_cheapest_to_rebuild_first() {
|
||||
// Registered deliberately backwards, because the guarantee is about
|
||||
// the tier and not about who registered first. A handler that simply
|
||||
// ran its list would pass every other assertion here and fail this
|
||||
// one — and on a device it would throw away thumbnails to keep a
|
||||
// render pipeline that nothing was going to draw.
|
||||
reset();
|
||||
let seen = recorder();
|
||||
relieve(Level::Critical);
|
||||
assert_eq!(
|
||||
*seen.borrow(),
|
||||
vec![Tier::Gpu, Tier::Proxies, Tier::Thumbnails]
|
||||
);
|
||||
reset();
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn hiding_the_window_costs_only_the_gpu() {
|
||||
// The cheap moment: give back what a window that is not on screen
|
||||
// cannot use, and keep what the user will be looking at when they come
|
||||
// back. Widening this to everything would make every task switch a
|
||||
// reload of the grid.
|
||||
reset();
|
||||
let seen = recorder();
|
||||
relieve(Level::UiHidden);
|
||||
assert_eq!(*seen.borrow(), vec![Tier::Gpu]);
|
||||
reset();
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn pressure_before_anything_is_registered_is_not_a_failure() {
|
||||
// The launch window: `android_main` installs the listener before
|
||||
// `run` builds a single cache, so the first minutes of a cold start
|
||||
// can deliver a warning to an empty registry.
|
||||
reset();
|
||||
relieve(Level::Critical);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_sink_may_register_another_without_deadlocking() {
|
||||
// Guards the re-entry the take-and-restore exists for: a sink is
|
||||
// interface code, and interface code that reached this module again
|
||||
// would otherwise meet a borrow it already held.
|
||||
reset();
|
||||
let seen = Rc::new(RefCell::new(0usize));
|
||||
{
|
||||
let seen = seen.clone();
|
||||
evict_at(Tier::Gpu, move || {
|
||||
*seen.borrow_mut() += 1;
|
||||
evict_at(Tier::Thumbnails, || {});
|
||||
});
|
||||
}
|
||||
relieve(Level::Critical);
|
||||
assert_eq!(*seen.borrow(), 1);
|
||||
// And the one it added survived, rather than being dropped with the
|
||||
// temporary list.
|
||||
assert_eq!(SINKS.with_borrow(|sinks| sinks.len()), 2);
|
||||
reset();
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user