Compare commits

...
3 Commits
Author SHA1 Message Date
dtourolle c5c61a3503 fix(functional): centre text vertically in buttons and form fields (S14)
Python CI / test (3.10) (push) Has been cancelled
Python CI / test (3.12) (push) Has been cancelled
Python CI / test (3.13) (push) Has been cancelled
Both renderers placed the baseline at box_top + height/2 + descent/2. Centring
glyphs of visual height ascent+descent in a box of height H puts the baseline at
box_top + H/2 + (ascent-descent)/2; the two agree only when ascent is exactly
twice descent. DejaVu is nearer 4:1, so labels sat high against the top edge -
measured at 5px above and 11px below for a 14px button.

ButtonText also sized itself from the nominal font size, which is smaller than
the text's visual height (17px of ink for a 14px DejaVu font), leaving the
button too short to centre its label in. It now measures ascent+descent, with a
fallback for font objects that cannot report metrics.

docs/images/example_07_pressed_state.png was stale - no example writes it, the
demo emits demo_07_pressed.png at the repo root and the docs copy had been
placed by hand in November. Refreshed here; the demo should write straight to
docs/images/ so it cannot drift again.
2026-08-06 22:59:27 +02:00
dtourolle 202dacf350 fix(page): separate measurement context from the render canvas (S3)
add_child invalidated the canvas but left _draw bound to it, and the draw
property only rebuilt when _draw was None. Callers therefore got a context
pointing at a discarded image while page._canvas stayed None. table_layouter
reads page._canvas directly, so every image inside a table cell laid out after
any other content silently degraded to a grey [Image: WxH] placeholder.

The property now rebuilds when either half is missing. On its own that would
make layout allocate a full-page RGBA canvas per line, because layout measures
text through the page - so measurement moves to page.measurement_draw, a 1x1
scratch context that is never invalidated. Its mode matches the render canvas
so that Text's width cache does not hold two entries per word.

Children built against the scratch context are re-bound to the live canvas by
render_children, which already synchronised _draw and _canvas; that behaviour
was incidental and is now load-bearing and documented as such.

Regenerating the examples shows table images rendering as images rather than
placeholders. The empty header row in the second table of example 05 is
unrelated and pre-existing - row height ignores cell padding, so text is clipped
as padding grows - recorded as evidence under S6.
2026-08-06 22:37:24 +02:00
dtourolle 284d521125 fix(html): keep inline content inside block containers (S1)
Inline tags map to ignore_handler because they are meant to be consumed by
extract_text_content, but only paragraph_handler and heading_handler ever called
it. Every other container walked its children calling process_element, so inline
tags returned None and their text was dropped:

  <p>hello <b>world</b> again</p>      -> hello world again   (correct)
  <div>hello <b>world</b> again</div>  -> nothing at all
  <li>hello <b>world</b> again</li>    -> nothing at all
  <td>hello <b>world</b> again</td>    -> nothing at all
  <td><a href=u>link</a> text</td>     -> text   (link discarded)

div_handler ignored bare text nodes outright, so a div containing text produced
no blocks whatsoever - which for real HTML and EPUB is most of the document.
Where text did survive, in cells and list items, each text node became its own
paragraph, so "a <b>b</b> c" fragmented onto separate lines.

process_block_children now walks a container's children once, gathering runs of
inline content into a single paragraph and letting block children through to
their own handlers, preserving document order. div, li, td, th and blockquote
all delegate to it, so they gain nested blocks, links and mixed content
together. <br> ends the current run rather than being a no-op.

extract_text_content is split so the run-level logic can be reused without
building a synthetic element: extract_words_from_nodes takes the nodes directly,
and skips comments, which previously had their text extracted as content.

paragraph_handler keeps its own image-splitting path for now; folding it into
process_block_children would also fix the ordering of text around images in a
paragraph, but it carries the EPUB cover-detection behaviour and is left alone.
2026-08-06 22:31:34 +02:00
12 changed files with 686 additions and 105 deletions
+69
View File
@@ -28,6 +28,7 @@ It is independent of every other spec here.
| [S11](#s11--partial-block-progress-is-discarded) | Partial-block progress is discarded | 0 | | [S11](#s11--partial-block-progress-is-discarded) | Partial-block progress is discarded | 0 |
| [S12](#s12--background-rendering) | Background rendering | 4 | | [S12](#s12--background-rendering) | Background rendering | 4 |
| [S13](#s13--word-spacing-and-alignment) | Word spacing and alignment | 0 | | [S13](#s13--word-spacing-and-alignment) | Word spacing and alignment | 0 |
| [S14](#s14--vertical-centring-in-buttons-and-fields) | Vertical centring in buttons and fields | 0 |
## Design invariants ## Design invariants
@@ -560,6 +561,22 @@ Four separate geometry defects:
drops every cell past `len(column_widths)` — two of three cells never render. drops every cell past `len(column_widths)` — two of three cells never render.
4. **rowspan is parsed and stored but never read** by any renderer or measurer; 4. **rowspan is parsed and stored but never read** by any renderer or measurer;
spanned rows just shift left. spanned rows just shift left.
5. **Row height ignores the cell padding it must contain.** The 40px minimum in
`_calculate_row_height_for_section` is a constant, so a larger `cell_padding`
eats into the content box rather than growing the row, and
`_render_cell_content` then clips the text against `available_height`.
Rendering the same header at two paddings:
```
padding=(8,10,8,10) border=1: header h=40, ink=593
padding=(10,12,10,12) border=2: header h=40, ink=288
```
Both rows are 40px tall; the second silently loses half its text. This is
visible in `docs/images/example_05_html_table_with_images.png`, whose second
table renders an empty header row. It is the same measure/render disagreement
as defect 1, and S5 removes it by construction: the cell page's content box
*is* the box its padding leaves.
### Design ### Design
@@ -1142,6 +1159,58 @@ Three defects, all visible as a right edge that wobbles from line to line.
--- ---
## S14 — Vertical centring in buttons and fields
### Problem
`ButtonText.render` and `FormFieldText.render` both placed the text baseline at
`box_top + box_height / 2 + descent / 2`. Centring glyphs whose visual height is
`ascent + descent` inside a box of height `H` puts the baseline at
`box_top + H/2 + (ascent - descent)/2`. The two agree only when
`ascent == 2 * descent`; DejaVu is nearer 4:1, so labels rode high against the
top edge of the control.
`ButtonText` also sized itself as `font_size + padding`, but the text's visual
height exceeds the nominal size — DejaVu at 14px measures 17 — so the button was
too short to centre its own label in.
### Evidence
A 14px "Save Document" button with 6px vertical padding, measuring the label's
ink against the button rectangle:
```
gap above text: 5px
gap below text: 11px
```
### Design
- `baseline = area_top + (area_height - (ascent + descent)) / 2 + ascent` in both
renderers.
- `ButtonText._padded_height` derives from `ascent + descent`, guarded so a mock
or unusual font object falls back to the nominal size.
### Acceptance criteria
- Label ink is centred within ±2px at font sizes 10, 14 and 20.
- Label ink stays inside the button rectangle.
- Button height is at least `ascent + descent + vertical padding`.
- A form field's value is centred within its input box (±3px).
### Files
`pyWebLayout/concrete/functional.py`
### Note
`docs/images/example_07_pressed_state.png` was stale — no example regenerates it;
`07_pressed_state_demo.py` writes `demo_07_pressed.png` at the repository root
and the docs copy had been placed by hand. It has been refreshed. Worth wiring
the demo to write straight to `docs/images/` so it cannot drift again.
---
## Test plan ## Test plan
Findings were reproduced with four probe scripts; each becomes a regression test Findings were reproduced with four probe scripts; each becomes a regression test
Binary file not shown.

Before

Width:  |  Height:  |  Size: 95 KiB

After

Width:  |  Height:  |  Size: 98 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 20 KiB

After

Width:  |  Height:  |  Size: 20 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 16 KiB

After

Width:  |  Height:  |  Size: 16 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 24 KiB

After

Width:  |  Height:  |  Size: 24 KiB

+33 -10
View File
@@ -153,7 +153,22 @@ class ButtonText(Text, Interactable, Queriable):
self, '_width', 0) if not hasattr( self, '_width', 0) if not hasattr(
self._width, '__call__') else 0 self._width, '__call__') else 0
self._padded_width = text_width + padding[1] + padding[3] self._padded_width = text_width + padding[1] + padding[3]
self._padded_height = self._style.font_size + padding[0] + padding[2]
# Size the button from the text's visual height (ascent + descent), not
# from the nominal font size. The two differ by several pixels - DejaVu at
# 14px measures 17 - so sizing by font_size leaves the button too short to
# centre its own label in.
self._text_height = self._visual_text_height()
self._padded_height = self._text_height + padding[0] + padding[2]
def _visual_text_height(self) -> int:
"""Height of the rendered text, ascender to descender."""
try:
ascent, descent = self._style.font.getmetrics()
return int(ascent + descent)
except (AttributeError, TypeError, ValueError):
# Mock or unusual font object; the nominal size is the best guess.
return int(getattr(self._style, 'font_size', 0) or 0)
@property @property
def button(self) -> Button: def button(self) -> Button:
@@ -237,11 +252,18 @@ class ButtonText(Text, Interactable, Queriable):
# Total button height minus top and bottom padding gives us text area height # Total button height minus top and bottom padding gives us text area height
text_area_height = self._padded_height - self._padding[0] - self._padding[2] text_area_height = self._padded_height - self._padding[0] - self._padding[2]
# Center the text visual height (ascent + descent) within the text area # Centre the text's visual height (ascent + descent) within the text area.
# The y position is where the baseline sits # text_y is the baseline, since Text renders with anchor "ls".
# Visual center = area_height/2, baseline should be at center + descent/2 #
vertical_center = text_area_height / 2 # top of glyphs = area_top + (area_height - (ascent + descent)) / 2
text_y = self._origin[1] + self._padding[0] + vertical_center + (descent / 2) # baseline = top of glyphs + ascent
#
# The previous form, area_top + area_height/2 + descent/2, is only
# equivalent when ascent == 2 * descent. Real fonts sit nearer 4:1, so the
# label rendered several pixels above centre, against the top edge.
text_top = self._origin[1] + self._padding[0] \
+ (text_area_height - (ascent + descent)) / 2
text_y = text_top + ascent
# Temporarily set origin for text rendering # Temporarily set origin for text rendering
original_origin = self._origin.copy() original_origin = self._origin.copy()
@@ -360,11 +382,12 @@ class FormFieldText(Text, Interactable, Queriable):
# Get font metrics to properly center the baseline # Get font metrics to properly center the baseline
ascent, descent = value_font.font.getmetrics() ascent, descent = value_font.font.getmetrics()
# Center the text vertically within the field # Centre the value within the input box. As in ButtonText, the
# The y coordinate is where the baseline sits (anchor="ls") # baseline sits at the top of the glyphs plus the ascent; centring on
vertical_center = self._field_height / 2 # half the box height plus half the descent only works for a 2:1
# ascent/descent ratio and otherwise rides high.
value_x = field_x + 5 value_x = field_x + 5
value_y = field_y + vertical_center + (descent / 2) value_y = field_y + (self._field_height - (ascent + descent)) / 2 + ascent
# Draw the value text # Draw the value text
self._draw.text((value_x, value_y), value_text, self._draw.text((value_x, value_y), value_text,
+33 -3
View File
@@ -15,6 +15,10 @@ class Page(Renderable, Queriable):
contains a given point. contains a given point.
""" """
# Mode of the render canvas. The measurement context matches it so that text
# width caching keys stay consistent between layout and rendering.
_CANVAS_MODE = 'RGBA'
def __init__(self, size: Tuple[int, int], style: Optional[PageStyle] = None, def __init__(self, size: Tuple[int, int], style: Optional[PageStyle] = None,
origin: Tuple[int, int] = (0, 0)): origin: Tuple[int, int] = (0, 0)):
""" """
@@ -32,6 +36,7 @@ class Page(Renderable, Queriable):
self._children: List[Renderable] = [] self._children: List[Renderable] = []
self._canvas: Optional[Image.Image] = None self._canvas: Optional[Image.Image] = None
self._draw: Optional[ImageDraw.Draw] = None self._draw: Optional[ImageDraw.Draw] = None
self._measurement_draw: Optional[ImageDraw.ImageDraw] = None
# Initialize y_offset to start of content area # Initialize y_offset to start of content area
# Position the first line so its baseline is close to the top boundary # Position the first line so its baseline is close to the top boundary
# For subsequent lines, baseline-to-baseline spacing is used # For subsequent lines, baseline-to-baseline spacing is used
@@ -168,13 +173,38 @@ class Page(Renderable, Queriable):
@property @property
def draw(self) -> Optional[ImageDraw.Draw]: def draw(self) -> Optional[ImageDraw.Draw]:
"""Get the ImageDraw object for drawing on this page's canvas""" """
if self._draw is None: Get the ImageDraw object bound to this page's render canvas.
Rebuilt whenever the canvas has been invalidated: a draw context
outlives the image it was created from, so checking only _draw would
hand back a context pointing at a discarded canvas.
"""
if self._draw is None or self._canvas is None:
# Initialize canvas and draw context if not already done # Initialize canvas and draw context if not already done
self._canvas = self._create_canvas() self._canvas = self._create_canvas()
self._draw = ImageDraw.Draw(self._canvas) self._draw = ImageDraw.Draw(self._canvas)
return self._draw return self._draw
@property
def measurement_draw(self) -> ImageDraw.ImageDraw:
"""
A scratch draw context for text metrics during layout.
Layout asks for text widths constantly, but has no reason to touch the
render canvas - and the canvas is invalidated on every add_child, so
measuring through `draw` would allocate a full-page image per line.
This context is 1x1 and never invalidated.
Its mode matches the render canvas because Text keys its width cache on
the draw mode; a mismatch would double every cache entry. Children built
against it are re-bound to the real canvas by render_children.
"""
if self._measurement_draw is None:
scratch = Image.new(self._CANVAS_MODE, (1, 1))
self._measurement_draw = ImageDraw.Draw(scratch)
return self._measurement_draw
def add_child(self, child: Renderable) -> 'Page': def add_child(self, child: Renderable) -> 'Page':
""" """
Add a child renderable object to this page. Add a child renderable object to this page.
@@ -333,7 +363,7 @@ class Page(Renderable, Queriable):
PIL Image with background and borders applied PIL Image with background and borders applied
""" """
# Create base image # Create base image
canvas = Image.new('RGBA', self._size, (*self._style.background_color, 255)) canvas = Image.new(self._CANVAS_MODE, self._size, (*self._style.background_color, 255))
# Draw borders if needed # Draw borders if needed
if self._style.border_width > 0: if self._style.border_width > 0:
+121 -85
View File
@@ -8,6 +8,7 @@ Each handler function has a robust signature that handles style hints, CSS class
from typing import List, Dict, Any, Optional, Union, Callable, Tuple, NamedTuple from typing import List, Dict, Any, Optional, Union, Callable, Tuple, NamedTuple
from bs4 import BeautifulSoup, Tag, NavigableString from bs4 import BeautifulSoup, Tag, NavigableString
from bs4.element import CData, Comment, Doctype, ProcessingInstruction
from pyWebLayout.abstract.inline import Word from pyWebLayout.abstract.inline import Word
from pyWebLayout.abstract.block import ( from pyWebLayout.abstract.block import (
Block, Block,
@@ -369,6 +370,24 @@ def extract_text_content(element: Tag, context: StyleContext) -> List[Word]:
element: BeautifulSoup Tag object element: BeautifulSoup Tag object
context: Current style context context: Current style context
Returns:
List of Word objects (including LinkedWord for hyperlinks)
"""
return extract_words_from_nodes(list(element.children), context)
def extract_words_from_nodes(nodes: List, context: StyleContext) -> List[Word]:
"""
Extract words from a sequence of sibling nodes.
Separated from extract_text_content so that a container holding a mix of
inline and block children can hand over just the inline runs, without
building a synthetic element to wrap them in.
Args:
nodes: BeautifulSoup nodes (Tags and NavigableStrings) in document order
context: Current style context
Returns: Returns:
List of Word objects (including LinkedWord for hyperlinks) List of Word objects (including LinkedWord for hyperlinks)
""" """
@@ -377,7 +396,12 @@ def extract_text_content(element: Tag, context: StyleContext) -> List[Word]:
words = [] words = []
for child in element.children: for child in nodes:
# Comments and processing instructions are NavigableString subclasses;
# their text is markup, not content.
if isinstance(child, (Comment, Doctype, CData, ProcessingInstruction)):
continue
if isinstance(child, NavigableString): if isinstance(child, NavigableString):
# Plain text - split into words # Plain text - split into words
text = str(child).strip() text = str(child).strip()
@@ -466,6 +490,93 @@ def extract_text_content(element: Tag, context: StyleContext) -> List[Word]:
return words return words
# Tags that flow within a line of text rather than forming a block of their own.
# They carry no handler of their own: extract_words_from_nodes consumes them,
# applying their styling to the words they contain.
INLINE_TAGS = frozenset({
"a", "span", "strong", "b", "em", "i", "u", "s", "del", "ins", "mark",
"small", "sub", "sup", "code", "q", "cite", "abbr", "time",
})
def is_inline(node) -> bool:
"""
Whether a node belongs to a run of text rather than standing as its own block.
Args:
node: A BeautifulSoup Tag or NavigableString
Returns:
True for text and inline tags, False for block-level tags
"""
if isinstance(node, Tag):
return node.name.lower() in INLINE_TAGS
if isinstance(node, (Comment, Doctype, CData, ProcessingInstruction)):
return False
return isinstance(node, NavigableString)
def process_block_children(element: Tag, context: StyleContext) -> List[Block]:
"""
Process a container's children into a list of blocks.
Containers may hold a mix of inline and block content. Consecutive inline
children are gathered into a run and become one Paragraph; a block child ends
the current run and is processed by its own handler. This is the single entry
point for every container that is not itself a paragraph - div, li, td, th,
blockquote and the semantic containers.
Without this, inline tags reach process_element, whose handler for them is
ignore_handler, and their text is silently dropped.
Args:
element: The container element
context: Current style context
Returns:
Blocks in document order
"""
blocks: List[Block] = []
run: List = []
def flush_run():
"""Turn the pending inline run into a paragraph, if it holds any words."""
if not run:
return
words = extract_words_from_nodes(run, context)
run.clear()
if words:
paragraph = Paragraph(context.font)
for word in words:
paragraph.add_word(word)
blocks.append(paragraph)
for child in element.children:
# <br> ends the current line of text and starts a new one.
if isinstance(child, Tag) and child.name.lower() == "br":
flush_run()
continue
if is_inline(child):
run.append(child)
continue
if not isinstance(child, Tag):
continue # comments and similar
flush_run()
child_context = apply_element_styling(context, child)
result = process_element(child, child_context)
if result:
if isinstance(result, list):
blocks.extend(result)
else:
blocks.append(result)
flush_run()
return blocks
def process_element( def process_element(
element: Tag, context: StyleContext element: Tag, context: StyleContext
) -> Union[Block, List[Block], None]: ) -> Union[Block, List[Block], None]:
@@ -557,17 +668,7 @@ def paragraph_handler(element: Tag, context: StyleContext) -> Union[Paragraph, L
def div_handler(element: Tag, context: StyleContext) -> List[Block]: def div_handler(element: Tag, context: StyleContext) -> List[Block]:
"""Handle <div> elements - treat as generic container.""" """Handle <div> elements - treat as generic container."""
blocks = [] return process_block_children(element, context)
for child in element.children:
if isinstance(child, Tag):
child_context = apply_element_styling(context, child)
result = process_element(child, child_context)
if result:
if isinstance(result, list):
blocks.extend(result)
else:
blocks.append(result)
return blocks
def heading_handler(element: Tag, context: StyleContext) -> Heading: def heading_handler(element: Tag, context: StyleContext) -> Heading:
@@ -592,16 +693,8 @@ def heading_handler(element: Tag, context: StyleContext) -> Heading:
def blockquote_handler(element: Tag, context: StyleContext) -> Quote: def blockquote_handler(element: Tag, context: StyleContext) -> Quote:
"""Handle <blockquote> elements.""" """Handle <blockquote> elements."""
quote = Quote(context.font) quote = Quote(context.font)
for child in element.children: for block in process_block_children(element, context):
if isinstance(child, Tag): quote.add_block(block)
child_context = apply_element_styling(context, child)
result = process_element(child, child_context)
if result:
if isinstance(result, list):
for block in result:
quote.add_block(block)
else:
quote.add_block(result)
return quote return quote
@@ -655,28 +748,8 @@ def ordered_list_handler(element: Tag, context: StyleContext) -> HList:
def list_item_handler(element: Tag, context: StyleContext) -> ListItem: def list_item_handler(element: Tag, context: StyleContext) -> ListItem:
"""Handle <li> elements.""" """Handle <li> elements."""
list_item = ListItem(None, context.font) list_item = ListItem(None, context.font)
for block in process_block_children(element, context):
for child in element.children: list_item.add_block(block)
if isinstance(child, Tag):
child_context = apply_element_styling(context, child)
result = process_element(child, child_context)
if result:
if isinstance(result, list):
for block in result:
list_item.add_block(block)
else:
list_item.add_block(result)
elif isinstance(child, NavigableString):
# Direct text in list item - create paragraph
text = str(child).strip()
if text:
paragraph = Paragraph(context.font)
words = text.split()
for word_text in words:
if word_text:
paragraph.add_word(Word(word_text, context.font))
list_item.add_block(paragraph)
return list_item return list_item
@@ -728,27 +801,8 @@ def table_cell_handler(element: Tag, context: StyleContext) -> TableCell:
rowspan = int(context.element_attributes.get("rowspan", 1)) rowspan = int(context.element_attributes.get("rowspan", 1))
cell = TableCell(False, colspan, rowspan, context.font) cell = TableCell(False, colspan, rowspan, context.font)
# Process cell content for block in process_block_children(element, context):
for child in element.children: cell.add_block(block)
if isinstance(child, Tag):
child_context = apply_element_styling(context, child)
result = process_element(child, child_context)
if result:
if isinstance(result, list):
for block in result:
cell.add_block(block)
else:
cell.add_block(result)
elif isinstance(child, NavigableString):
# Direct text in cell - create paragraph
text = str(child).strip()
if text:
paragraph = Paragraph(context.font)
words = text.split()
for word_text in words:
if word_text:
paragraph.add_word(Word(word_text, context.font))
cell.add_block(paragraph)
return cell return cell
@@ -759,26 +813,8 @@ def table_header_cell_handler(element: Tag, context: StyleContext) -> TableCell:
rowspan = int(context.element_attributes.get("rowspan", 1)) rowspan = int(context.element_attributes.get("rowspan", 1))
cell = TableCell(True, colspan, rowspan, context.font) cell = TableCell(True, colspan, rowspan, context.font)
# Process cell content (same as td) for block in process_block_children(element, context):
for child in element.children: cell.add_block(block)
if isinstance(child, Tag):
child_context = apply_element_styling(context, child)
result = process_element(child, child_context)
if result:
if isinstance(result, list):
for block in result:
cell.add_block(block)
else:
cell.add_block(result)
elif isinstance(child, NavigableString):
text = str(child).strip()
if text:
paragraph = Paragraph(context.font)
words = text.split()
for word_text in words:
if word_text:
paragraph.add_word(Word(word_text, context.font))
cell.add_block(paragraph)
return cell return cell
+8 -7
View File
@@ -165,7 +165,7 @@ def paragraph_layouter(paragraph: Paragraph,
# Create a temporary Text object to calculate word width # Create a temporary Text object to calculate word width
if word: if word:
temp_text = Text.from_word(word, page.draw) temp_text = Text.from_word(word, page.measurement_draw)
temp_text.width temp_text.width
else: else:
pass pass
@@ -174,7 +174,7 @@ def paragraph_layouter(paragraph: Paragraph,
spacing=word_spacing_constraints, spacing=word_spacing_constraints,
origin=(x_cursor, y_cursor), origin=(x_cursor, y_cursor),
size=(page.available_width, baseline_spacing), size=(page.available_width, baseline_spacing),
draw=page.draw, draw=page.measurement_draw,
font=font, font=font,
halign=text_align halign=text_align
) )
@@ -225,7 +225,7 @@ def paragraph_layouter(paragraph: Paragraph,
return False, i, overflow_text return False, i, overflow_text
# Check if the word will fit on the new line before adding it # Check if the word will fit on the new line before adding it
temp_text = Text.from_word(word, page.draw) temp_text = Text.from_word(word, page.measurement_draw)
if temp_text.width > current_line.size[0]: if temp_text.width > current_line.size[0]:
# Word is too wide for the line, we need to hyphenate it # Word is too wide for the line, we need to hyphenate it
if len(word.text) >= 6: if len(word.text) >= 6:
@@ -234,13 +234,13 @@ def paragraph_layouter(paragraph: Paragraph,
(Text( (Text(
pair[0], pair[0],
word.style, word.style,
page.draw, page.measurement_draw,
line=current_line, line=current_line,
source=word), source=word),
Text( Text(
pair[1], pair[1],
word.style, word.style,
page.draw, page.measurement_draw,
line=current_line, line=current_line,
source=word)) for pair in word.possible_hyphenation()] source=word)) for pair in word.possible_hyphenation()]
if len(splits) > 0: if len(splits) > 0:
@@ -455,7 +455,7 @@ def button_layouter(button: Button,
available_height = page.remaining_height available_height = page.remaining_height
# Create ButtonText renderable # Create ButtonText renderable
button_text = ButtonText(button, font, page.draw, padding=padding) button_text = ButtonText(button, font, page.measurement_draw, padding=padding)
# Check if button fits on current page # Check if button fits on current page
button_height = button_text.size[1] button_height = button_text.size[1]
@@ -505,7 +505,8 @@ def form_field_layouter(field: FormField, page: Page, font: Optional[Font] = Non
available_height = page.remaining_height available_height = page.remaining_height
# Create FormFieldText renderable # Create FormFieldText renderable
field_text = FormFieldText(field, font, page.draw, field_height=field_height) field_text = FormFieldText(field, font, page.measurement_draw,
field_height=field_height)
# Check if field fits on current page # Check if field fits on current page
total_field_height = field_text.size[1] total_field_height = field_text.size[1]
+129
View File
@@ -0,0 +1,129 @@
"""
Regression tests for the page draw/canvas lifecycle (spec S3).
add_child invalidates the canvas but left _draw pointing at it, and the draw
property only rebuilt when _draw was None. Callers therefore received a context
bound to a discarded image while page._canvas stayed None - which is how images
inside table cells ended up as grey placeholders: table_layouter passed
canvas=None through to the cell renderer.
Fixing that alone would make layout allocate a full-page canvas per line, since
layout measures text through the page. Measurement now goes through a dedicated
scratch context.
"""
import pytest
from PIL import Image
from pyWebLayout.abstract.block import Image as AbstractImage, Paragraph
from pyWebLayout.abstract.inline import Word
from pyWebLayout.concrete.page import Page
from pyWebLayout.layout.document_layouter import DocumentLayouter
from pyWebLayout.style import Font
from pyWebLayout.style.page_style import PageStyle
@pytest.fixture
def font():
return Font(font_size=12)
@pytest.fixture
def page():
return Page(size=(400, 600), style=PageStyle())
def paragraph_of(font, count=40):
paragraph = Paragraph(font)
for i in range(count):
paragraph.add_word(Word(f"word{i}", font))
return paragraph
class TestDrawIsNeverStale:
def test_draw_matches_canvas_after_add_child(self, page, font):
page.draw # force canvas creation
DocumentLayouter(page).layout_paragraph(paragraph_of(font))
assert page.draw.im is page._canvas.im, \
"draw must be bound to the page's current canvas"
def test_canvas_is_present_after_layout(self, page, font):
DocumentLayouter(page).layout_paragraph(paragraph_of(font))
page.draw
assert page._canvas is not None
def test_repeated_draw_access_is_stable(self, page):
first = page.draw
assert page.draw is first, "draw must not be rebuilt while the canvas stands"
class TestMeasurementDoesNotAllocateCanvases:
def test_layout_allocates_no_page_canvas(self, page, font, monkeypatch):
calls = []
original = Page._create_canvas
def counting(self):
calls.append(1)
return original(self)
monkeypatch.setattr(Page, "_create_canvas", counting)
DocumentLayouter(page).layout_paragraph(paragraph_of(font, 400))
assert calls == [], \
f"layout allocated {len(calls)} full-page canvases; it should allocate none"
def test_measurement_context_is_tiny_and_matches_canvas_mode(self, page):
scratch = page.measurement_draw
assert scratch.im.size == (1, 1)
assert scratch.mode == Page._CANVAS_MODE
def test_measurement_context_is_stable(self, page, font):
first = page.measurement_draw
DocumentLayouter(page).layout_paragraph(paragraph_of(font))
assert page.measurement_draw is first, \
"the scratch context must survive canvas invalidation"
class TestRenderIsRepeatable:
def test_two_renders_are_identical(self, page, font):
DocumentLayouter(page).layout_paragraph(paragraph_of(font))
first = page.render().copy()
second = page.render().copy()
assert first.tobytes() == second.tobytes()
class TestImageInCellGetsARealCanvas:
"""The concrete symptom: table images degraded to placeholders."""
@pytest.fixture
def image_path(self, tmp_path):
path = tmp_path / "swatch.png"
Image.new("RGB", (40, 30), (10, 200, 10)).save(path)
return str(path)
def test_table_after_paragraph_receives_a_canvas(self, page, font, image_path):
from pyWebLayout.abstract.block import Table, TableCell, TableRow
from pyWebLayout.layout.document_layouter import table_layouter
layouter = DocumentLayouter(page)
layouter.layout_paragraph(paragraph_of(font, 10))
table = Table()
row = TableRow()
cell = TableCell()
cell.add_block(AbstractImage(image_path))
row.add_cell(cell)
table.add_row(row)
# The canvas is invalidated by the preceding add_child; the table must
# still be handed a real one.
assert table_layouter(table, page) or True # placement may fail on space
assert page._canvas is not None, \
"table layout must not run against a None canvas"
@@ -0,0 +1,124 @@
"""
Regression tests for vertical centring of text in buttons and form fields.
Both placed the baseline at `top + height/2 + descent/2`. Centring text whose
visual height is ascent+descent inside a box of height H puts the baseline at
`top + H/2 + (ascent-descent)/2`; the two agree only when ascent == 2*descent.
Real fonts have a much larger ratio - DejaVu is nearer 4:1 - so the text sat
several pixels high, hugging the top edge of the button.
The button was also sized from the nominal font size rather than the text's
actual visual height, leaving it too short to centre anything in.
"""
import numpy as np
import pytest
from PIL import Image, ImageDraw
from pyWebLayout.abstract.functional import Button, FormField, FormFieldType
from pyWebLayout.concrete.functional import ButtonText, FormFieldText
from pyWebLayout.style import Font
CANVAS = (300, 120)
PADDING = (6, 10, 6, 10) # top, right, bottom, left
@pytest.fixture
def draw_ctx():
image = Image.new("RGB", CANVAS, (255, 255, 255))
return image, ImageDraw.Draw(image)
def ink_rows(image, box):
"""
Rows within box that carry text ink.
Only the central columns are sampled: the button has rounded corners, so the
page background shows through at the extremes of every row and would read as
white text on all of them.
"""
x0, y0, x1, y1 = box
inset = (x1 - x0) // 4
pixels = image.convert("RGB").load()
rows = []
for y in range(y0, y1):
for x in range(x0 + inset, x1 - inset):
r, g, b = pixels[x, y]
# Button text is white on a blue fill; look for near-white ink.
if r > 240 and g > 240 and b > 240:
rows.append(y)
break
return rows
class TestButtonTextCentring:
@pytest.mark.parametrize("font_size", [10, 14, 20])
def test_text_is_vertically_centred(self, draw_ctx, font_size):
image, draw = draw_ctx
font = Font(font_size=font_size, colour=(255, 255, 255))
button = ButtonText(Button(label="Save Document", callback=lambda p: None),
font, draw, padding=PADDING)
button.set_origin(np.array([20, 20]))
button.render()
x0, y0 = 20, 20
x1 = x0 + int(button.size[0])
y1 = y0 + int(button.size[1])
rows = ink_rows(image, (x0, y0, x1, y1))
assert rows, "the button should have visible text"
gap_above = min(rows) - y0
gap_below = y1 - max(rows) - 1
assert abs(gap_above - gap_below) <= 2, (
f"text not centred at size {font_size}: "
f"{gap_above}px above, {gap_below}px below")
def test_button_is_tall_enough_for_its_text(self):
font = Font(font_size=14, colour=(255, 255, 255))
image = Image.new("RGB", CANVAS, (255, 255, 255))
draw = ImageDraw.Draw(image)
button = ButtonText(Button(label="Cancel", callback=lambda p: None),
font, draw, padding=PADDING)
ascent, descent = font.font.getmetrics()
assert int(button.size[1]) >= ascent + descent + PADDING[0] + PADDING[2], \
"button height must accommodate the text's visual height, not the nominal size"
def test_text_stays_inside_the_button(self, draw_ctx):
image, draw = draw_ctx
font = Font(font_size=14, colour=(255, 255, 255))
button = ButtonText(Button(label="Save Document", callback=lambda p: None),
font, draw, padding=PADDING)
button.set_origin(np.array([20, 20]))
button.render()
y0, y1 = 20, 20 + int(button.size[1])
rows = ink_rows(image, (20, y0, 20 + int(button.size[0]), y1))
assert min(rows) >= y0, "text escaped above the button"
assert max(rows) < y1, "text escaped below the button"
class TestFormFieldValueCentring:
def test_value_is_centred_in_the_input_box(self):
image = Image.new("RGB", (300, 120), (0, 0, 0))
draw = ImageDraw.Draw(image)
font = Font(font_size=12, colour=(0, 0, 0))
field = FormField(name="who", field_type=FormFieldType.TEXT, value="Hello")
renderable = FormFieldText(field, font, draw, field_height=28)
renderable.set_origin(np.array([10, 10]))
renderable.render()
field_y = 10 + font.font_size + 5
pixels = image.convert("RGB").load()
rows = [y for y in range(field_y, field_y + 28)
if any(pixels[x, y] == (0, 0, 0) for x in range(12, 200))]
assert rows, "the field value should be visible"
gap_above = min(rows) - field_y
gap_below = (field_y + 28) - max(rows) - 1
assert abs(gap_above - gap_below) <= 3, (
f"field value not centred: {gap_above}px above, {gap_below}px below")
@@ -0,0 +1,169 @@
"""
Regression tests for inline content inside block containers (spec S1).
Inline tags are registered to ignore_handler because they are meant to be
consumed by extract_text_content. Only <p> and <h1>-<h6> ever called it, so
every other container - div, li, td, th, blockquote - iterated its children as
blocks, and inline tags returned None. Their text was silently discarded, and
bare text nodes each became a separate paragraph.
"""
import pytest
from pyWebLayout.abstract.block import (
HList,
Paragraph,
Quote,
Table,
)
from pyWebLayout.abstract.inline import LinkedWord, Word
from pyWebLayout.io.readers.html_extraction import parse_html_string
def words_of(block):
return [w.text for w in getattr(block, 'words', [])]
def all_words(blocks):
out = []
for block in blocks:
out.extend(words_of(block))
return out
def cell_blocks(table):
for _, row in table.all_rows():
for cell in row.cells():
yield list(cell.blocks())
EXPECTED = ["hello", "world", "again"]
class TestInlineContentIsKept:
"""The same markup must survive in every container."""
def test_paragraph_control(self):
"""<p> already worked - this is the reference behaviour."""
blocks = parse_html_string("<p>hello <b>world</b> again</p>")
assert all_words(blocks) == EXPECTED
def test_div(self):
blocks = parse_html_string("<div>hello <b>world</b> again</div>")
assert all_words(blocks) == EXPECTED
def test_list_item(self):
blocks = parse_html_string("<ul><li>hello <b>world</b> again</li></ul>")
hlist = next(b for b in blocks if isinstance(b, HList))
item = list(hlist.items())[0]
assert all_words(item.blocks()) == EXPECTED
def test_table_cell(self):
blocks = parse_html_string(
"<table><tr><td>hello <b>world</b> again</td></tr></table>")
table = next(b for b in blocks if isinstance(b, Table))
assert all_words(next(cell_blocks(table))) == EXPECTED
def test_table_header_cell(self):
blocks = parse_html_string(
"<table><tr><th>hello <b>world</b> again</th></tr></table>")
table = next(b for b in blocks if isinstance(b, Table))
assert all_words(next(cell_blocks(table))) == EXPECTED
def test_blockquote(self):
blocks = parse_html_string("<blockquote>hello <b>world</b> again</blockquote>")
quote = next(b for b in blocks if isinstance(b, Quote))
assert all_words(quote.blocks()) == EXPECTED
class TestInlineRunsCoalesce:
"""A run of inline content is one paragraph, not one per text node."""
def test_div_yields_a_single_paragraph(self):
blocks = parse_html_string("<div>a <b>b</b> c</div>")
paragraphs = [b for b in blocks if isinstance(b, Paragraph)]
assert len(paragraphs) == 1, f"expected one paragraph, got {len(blocks)} blocks"
assert words_of(paragraphs[0]) == ["a", "b", "c"]
def test_cell_yields_a_single_paragraph(self):
blocks = parse_html_string("<table><tr><td>a <b>b</b> c</td></tr></table>")
table = next(b for b in blocks if isinstance(b, Table))
cell = next(cell_blocks(table))
assert len(cell) == 1
assert words_of(cell[0]) == ["a", "b", "c"]
def test_block_child_splits_the_run(self):
"""Inline runs either side of a block child stay separate, in order."""
blocks = parse_html_string(
"<table><tr><td>before<p>middle</p>after</td></tr></table>")
table = next(b for b in blocks if isinstance(b, Table))
cell = next(cell_blocks(table))
assert [words_of(b) for b in cell] == [["before"], ["middle"], ["after"]]
def test_line_break_splits_the_run(self):
blocks = parse_html_string("<div>first<br>second</div>")
paragraphs = [b for b in blocks if isinstance(b, Paragraph)]
assert [words_of(p) for p in paragraphs] == [["first"], ["second"]]
def test_whitespace_between_blocks_makes_no_paragraph(self):
blocks = parse_html_string("<div>\n <p>one</p>\n <p>two</p>\n</div>")
assert [words_of(b) for b in blocks] == [["one"], ["two"]]
class TestLinksSurvive:
"""<a href> must produce LinkedWord wherever it appears."""
def test_link_in_cell(self):
blocks = parse_html_string(
'<table><tr><td><a href="http://x">link</a> text</td></tr></table>')
table = next(b for b in blocks if isinstance(b, Table))
cell = next(cell_blocks(table))
found = [w for b in cell for w in getattr(b, 'words', [])]
assert [w.text for w in found] == ["link", "text"]
linked = [w for w in found if isinstance(w, LinkedWord)]
assert len(linked) == 1
assert linked[0].location == "http://x"
def test_link_in_div(self):
blocks = parse_html_string('<div>see <a href="#s2">Section 2</a> now</div>')
found = [w for b in blocks for w in getattr(b, 'words', [])]
assert [w.text for w in found] == ["see", "Section", "2", "now"]
assert all(isinstance(w, LinkedWord) for w in found[1:3])
def test_link_in_list_item(self):
blocks = parse_html_string('<ul><li><a href="u">click</a> here</li></ul>')
hlist = next(b for b in blocks if isinstance(b, HList))
item = hlist._items[0]
found = [w for b in item.blocks() for w in getattr(b, 'words', [])]
assert [w.text for w in found] == ["click", "here"]
assert isinstance(found[0], LinkedWord)
class TestNestedContainers:
def test_div_in_div(self):
blocks = parse_html_string("<div>outer <div>inner</div> tail</div>")
assert [words_of(b) for b in blocks] == [["outer"], ["inner"], ["tail"]]
def test_block_children_still_pass_through(self):
blocks = parse_html_string("<div><h1>Title</h1><p>Body</p></div>")
assert len(blocks) == 2
assert words_of(blocks[0]) == ["Title"]
assert words_of(blocks[1]) == ["Body"]
def test_cell_containing_a_list(self):
blocks = parse_html_string(
"<table><tr><td>intro<ul><li>item</li></ul></td></tr></table>")
table = next(b for b in blocks if isinstance(b, Table))
cell = next(cell_blocks(table))
assert isinstance(cell[0], Paragraph)
assert words_of(cell[0]) == ["intro"]
assert isinstance(cell[1], HList)
class TestComments:
def test_comment_text_is_not_content(self):
blocks = parse_html_string("<div>real<!-- hidden note -->text</div>")
assert all_words(blocks) == ["real", "text"]