Write a panic down where it can still be read, with nothing in it that names the user

NFR-OPS-2 is two sentences — local crash capture always, upload only on
explicit opt-in — and what existed was one `log::error!` in the Android entry
point and nothing at all on desktop. So a panic on desktop went to stderr and
died with the terminal, and a panic on Android went to a logcat ring buffer
that is gone long before anyone reports anything. What the user saw either way
was a job that stopped or a control that went dead, with nothing to send.

That matters more here than it would in most applications, because
NFR-ARCH-4 says no worker error may panic the process and the code is written
that way: errors are typed and attached to the image or job they belong to. A
panic is therefore by construction a bug — an invariant this codebase believed
and got wrong — and it was the one class of failure with no trace.

`dr_plat::crash` writes a record to the XDG state directory: version, time,
os, arch, thread, panic location, message, backtrace. In dr-plat rather than
in either entry point because "where does this platform let an application
keep state" is a platform question, and Android's answer is neither XDG nor
`temp_dir` — `set_state_dir` takes it from `internal_data_path`, the same
place `dr_sync::account::set_data_dir` gets its answer. The hook resolves the
directory when it fires rather than when it is installed, which is what lets
it go in before everything else and cover the startup it would otherwise miss.

**The content rule is the substance of this, not the plumbing.** NFR-SEC-2
forbids credentials in logs and plain files; the same reasoning applies with
more force to what this application is actually about, because a user's
library is private and so is its shape. `/home/anna/Photos/2019 Divorce/` says
something about a person, and a crash record is exactly the file someone
attaches to a bug report while trying to be helpful. So `redact` runs over the
message *and* the backtrace, and is deliberately blunt: anything containing a
slash goes, `content://` and `primary:DCIM/...` included, since SAF names a
library just as precisely as a path does; anything beside a word like
`password` goes; a long opaque run with letters and digits in it goes, which
is the shape of an app password nobody labelled. The one exception is a `.rs`
path, which keeps its basename — a backtrace with no filenames is close to
useless and `library.rs:1270` says nothing about anybody.

Over-redaction costs legibility. Under-redaction costs a user something they
cannot take back. Those are not comparable, so the boundary is not the place
to be clever.

NFR-SEC-5 — face data never in a crash report, under any configuration — is
met structurally rather than by filtering: this module reads no catalog, opens
no image, touches no account. A record is assembled from the panic hook's own
arguments and `std::env::consts`, and there is no code path from here to an
embedding. The message length cap is the backstop for a payload some other
module formatted something large into.

stderr is the one surface that still sees the message unredacted, deliberately:
the previous hook is chained rather than replaced, so a developer watching a
terminal does not lose the panic because the application started writing files.
It is ephemeral, local, and never attached to a report.

**No upload path, and not half of one** — no endpoint, no queue, no "send this
later" flag. Opt-in upload needs a server to receive it and a consent flow
stating what leaves the device (NFR-SEC-4, and the preview-and-consent step
NFR-OPS-1 requires of the diagnostics bundle). Neither exists, and a transport
built ahead of its consent is the shape of thing that later gets switched on
by default.

Leaves NFR-OPS-1 cheaper by three things it will want unchanged: `state_dir`
(the log belongs at `state_dir()/log` beside `crash/`, so the diagnostics
bundle has one directory to collect), `redact` (NFR-OPS-1's "automatic
redaction of credentials and tokens" is this function), and `prune` (a
size-capped rotation is this, counting bytes instead of files).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-30 10:34:12 +02:00
co-authored by Claude Opus 5
parent 8eeb9ba0f6
commit 35954dfa1d
6 changed files with 659 additions and 3 deletions
+4
View File
@@ -19,6 +19,10 @@ dr-ui.workspace = true
# For `account::set_data_dir`: only the platform entry point knows where Android
# lets this app keep files, and it must be set before any store is opened.
dr-sync.workspace = true
# For the panic hook, and for `crash::set_state_dir` — Android has no XDG
# directories, so the entry point is the only place that knows where a crash
# record may be written.
dr-plat.workspace = true
# Directly, not just through dr-ui: `android_main` takes an `AndroidApp` and
# calls `slint::android::init`, both of which come from this crate. The backend
# feature comes from dr-ui's target-specific dependency.
+16 -3
View File
@@ -39,9 +39,18 @@ fn android_main(app: slint::android::AndroidApp) {
// worker thread that panics is invisible: the process survives, the
// channel it was writing to closes, and the UI reports only that
// something "failed unexpectedly" with no way to find out what.
std::panic::set_hook(Box::new(|info| {
log::error!("panic: {info}");
}));
//
// This used to be one `log::error!` of the raw panic, which had two
// problems: logcat is a ring buffer that is gone by the time a user
// reports anything, and the raw message can carry a document URI naming
// their library or a credential a library interpolated into an error
// (NFR-SEC-2). `dr_plat::crash` writes a redacted record to disk and logs
// the redacted form. Nothing uploads it.
//
// Before `set_state_dir` on purpose: the hook resolves the directory when
// it fires, so installing it first covers the startup below rather than
// leaving it uncovered.
dr_plat::crash::install(env!("CARGO_PKG_VERSION"));
log::info!("DarkRoom v{}", env!("CARGO_PKG_VERSION"));
@@ -54,6 +63,10 @@ fn android_main(app: slint::android::AndroidApp) {
match app.internal_data_path() {
Some(dir) => {
log::info!("data dir: {}", dir.display());
// Crash records go beside the account data rather than under it:
// both are app-private and neither is a cache, which is the whole
// distinction that matters here (see `dr_plat::crash::state_dir`).
dr_plat::crash::set_state_dir(dir.join("state"));
dr_sync::account::set_data_dir(dir);
}
None => log::error!("no internal data path; settings will not persist"),
+4
View File
@@ -7,6 +7,10 @@ license.workspace = true
[dependencies]
dr-ui.workspace = true
# For the panic hook alone. Directly rather than through dr-ui, because it has
# to be installed before `dr_ui::run` — a panic during startup is exactly the
# one this exists to catch.
dr-plat.workspace = true
anyhow.workspace = true
env_logger.workspace = true
log.workspace = true
+8
View File
@@ -10,6 +10,14 @@ fn main() -> anyhow::Result<()> {
))
.init();
// Immediately after the logger and before anything that could fail. Until
// now a panic on desktop went to stderr and died with the terminal, which
// means every panic a user has ever hit was unreportable: the process
// survives (the panicking worker does not), a control goes dead, and there
// is nothing on disk to say why. The record is local and stays local —
// there is no upload path, by design; see `dr_plat::crash`.
dr_plat::crash::install(env!("CARGO_PKG_VERSION"));
log::info!("DarkRoom v{}", env!("CARGO_PKG_VERSION"));
let paths: Vec<PathBuf> = std::env::args().skip(1).map(PathBuf::from).collect();