Make the manual's EPUB from the same JPEG pictures as its PDF
Build and test / Publish the release (push) Blocked by required conditions
Benchmarks / CPU and I/O (per commit) (push) Successful in 3m35s
Benchmarks / Frame budget (on demand) (push) Skipped
Build and test / Desktop (Linux) (push) Failing after 1h38m56s
Build and test / Layer separation (push) Successful in 27s
🐳 Android image / Build and push (push) Successful in 5s
Build and test / android-image (push) Successful in 5s
🐳 Windows image / Build and push (push) Successful in 2s
Build and test / windows-image (push) Successful in 3s
Build and test / Windows (x86_64, cross) (push) Waiting to run
Manual pages / Publish the manual (push) Successful in 44s
Traceability / Requirement traces (push) Successful in 1m3s
Build and test / Android (aarch64) (push) In progress

pandoc embedded every recording whole, so the EPUB came out at 63 MB for
readers that mostly show a GIF's first frame anyway. tools/manual/pdf.py
now makes both downloads from its JPEG copies of the pictures: 7.5 MB.
This commit is contained in:
2026-10-09 03:03:31 -04:00
parent b53d4dc391
commit 880915e061
2 changed files with 33 additions and 21 deletions
+31 -16
View File
@@ -1,14 +1,17 @@
#!/usr/bin/env python3
"""The manual as a PDF: docs/manual/index.html through its print stylesheet.
"""The manual as a PDF, and as an EPUB: the downloads the page links to.
python3 tools/manual/pdf.py docs/manual out.pdf
python3 tools/manual/pdf.py docs/manual out.pdf [out.epub]
Needs WeasyPrint (and so Pillow): `pip install weasyprint`. The Manual pages
workflow runs this and publishes the result beside the page.
The PDF is docs/manual/index.html through its print stylesheet; the EPUB is
README.md through pandoc, which reads Markdown better than a page. Needs
WeasyPrint (and so Pillow), and pandoc for the EPUB. The Manual pages
workflow runs this and publishes both beside the page.
The page's pictures are PNG screenshots and GIF recordings, 16 MB and 46 MB
of them, and WeasyPrint embeds a picture as it finds it: losslessly, a GIF
whole. So the PDF is made from a copy of the page whose pictures are JPEGs
of them, and WeasyPrint and pandoc embed a picture as they find it:
losslessly, a GIF whole (63 MB of EPUB, for e-readers that mostly do not
animate). So both are made from copies whose pictures are JPEGs
(a recording's first frame, which is all a page can show) at the size they
were recorded. Full chroma, because a 4:2:0 JPEG smears the coloured noise
the AI denoise close-ups exist to show.
@@ -16,6 +19,7 @@ the AI denoise close-ups exist to show.
import os
import re
import shutil
import subprocess
import sys
import tempfile
@@ -31,7 +35,7 @@ def as_jpeg(src, dst):
im.convert('RGB').save(dst, 'JPEG', quality=QUALITY, subsampling=0, optimize=True)
def main(manual_dir, out):
def main(manual_dir, pdf, epub=None):
with tempfile.TemporaryDirectory() as tmp:
os.mkdir(os.path.join(tmp, 'media'))
renamed = {}
@@ -42,16 +46,27 @@ def main(manual_dir, out):
jpg = f'{stem}.{ext[1:].lower()}.jpg' # two pictures may share a stem
as_jpeg(os.path.join(manual_dir, 'media', name), os.path.join(tmp, 'media', jpg))
renamed[name] = jpg
page = open(os.path.join(manual_dir, 'index.html'), encoding='utf-8').read()
page = re.sub(r'src="media/([^"]+)"',
lambda m: f'src="media/{renamed.get(m.group(1), m.group(1))}"', page)
with open(os.path.join(tmp, 'index.html'), 'w', encoding='utf-8') as f:
f.write(page)
HTML(os.path.join(tmp, 'index.html')).write_pdf(os.path.join(tmp, 'manual.pdf'))
shutil.move(os.path.join(tmp, 'manual.pdf'), out)
def copy(source, pattern, out):
text = open(os.path.join(manual_dir, source), encoding='utf-8').read()
text = re.sub(pattern, lambda m: m.group(1) + renamed.get(m.group(2), m.group(2)), text)
with open(os.path.join(tmp, out), 'w', encoding='utf-8') as f:
f.write(text)
return os.path.join(tmp, out)
page = copy('index.html', r'(src="media/)([^"]+)', 'index.html')
HTML(page).write_pdf(os.path.join(tmp, 'manual.pdf'))
shutil.move(os.path.join(tmp, 'manual.pdf'), pdf)
if epub:
md = copy('README.md', r'(\]\(media/)([^)\s]+)', 'README.md')
title = next(l[2:].strip() for l in open(md, encoding='utf-8') if l.startswith('# '))
subprocess.run(['pandoc', 'README.md', '-o', 'manual.epub', '--metadata', f'title={title}',
'--toc', '--toc-depth=2', '--epub-chapter-level=2'], cwd=tmp, check=True)
shutil.move(os.path.join(tmp, 'manual.epub'), epub)
if __name__ == '__main__':
if len(sys.argv) != 3:
if len(sys.argv) not in (3, 4):
sys.exit(__doc__)
main(sys.argv[1], sys.argv[2])
main(*sys.argv[1:])