-
Notifications
You must be signed in to change notification settings - Fork 0
Current Scope
The exhaustive, itemised list of what is implemented — migrated from the README. Functionality is the by-area overview and the place to check first; this page is the fine-grained detail behind it, and Known Gaps is what is genuinely absent.
Implemented: cells merged down the page (w:vMerge), which behave as one tall cell — no rule is
drawn across the run, its shading covers all of it, and its text is placed over the run as a whole,
by its own vertical alignment. The rows a run covers keep the heights their own cells ask for and
the merged text runs down through them, so three lines merged across three one-line rows leave
those rows a line tall each; only what will not fit makes the run taller, and the last row of the
run takes all of it. A run reaching the foot of the page divides there wherever the break falls —
inside the row it begins in as readily as between the rows it covers — and Word rules the merged
cell at a page break although it rules none between a run's own rows, so both halves are closed
boxes. Breaking a table row across a page, at a line boundary inside its cells, with both
halves closed off by a full border box the way Word draws them — and moving the row whole instead
where it says w:cantSplit or where not even a line of it would fit. Vertical page alignment (top, centred, bottom, and justified — which spreads the
spare height between the paragraphs rather than between the lines, so a paragraph that wraps stays
whole), font subsetting for both kinds of outline. A TrueType face is numbered again from
nothing, so a document that used thirty glyphs embeds thirty rather than the three thousand the
face has; a CFF one keeps its numbering — renumbering it means rewriting what its charset says
about every glyph — and has its charstrings and the subroutines nothing reaches emptied instead.
Times New Roman goes from 676KB to 33KB, Hiragino Sans GB from 11.4MB to 497KB. Kerning, read from a font's GPOS table as well as the legacy one — Calibri has only
the former and Times New Roman only the latter, so both are needed to kern either — applied where
w:kern asks for it and from the type size it names upwards, tab stops of every alignment — left, centre, right and decimal, the last three
resolved once the text after the tab has been measured, with a stop the line has already passed
falling through to the next one, and leaders filling the gap a stop opens (dots, hyphens,
underscores and middle dots, set on a grid measured from the edge of the page so that entries of
different lengths line up with each other), and the vertical rule a bar stop asks for, down every
line of its paragraph — widow and orphan control (two lines of a paragraph on each side of a page or column
break, or none — a three-line paragraph moves whole, since it cannot give two to both), keeping a
paragraph with the next one and keeping its own lines together (w:keepNext and w:keepLines,
including chains of headings that move as one), section breaks (next-page, continuous, even- and odd-page, each section with its own
page size, orientation and margins, running heads inherited per kind from the section before where
one says nothing), multiple text columns (evenly divided or individually stated, column breaks,
the rule down the gap where the document asks for one, and the last page of a section evened out
between its columns where a continuous break closes it), footnotes and endnotes (numbered in reference order, arabic for footnotes and roman
for endnotes unless the document says otherwise, and numbered again from the beginning on every
page or in every section where the section asks for it; a footnote goes to the foot of the page its
reference lands on — or under the last line of text on it, where the section asks for that — and
takes that space out of the body above it, dividing between that page and
the next where it is too long for the room left under it, an endnote carries on after the
body like ordinary content — or after each section where the document asks for that — and both are
ruled off by the separator), hyperlinks (external addresses as clickable regions, internal links to bookmarks
anywhere in the document, with the regions placed and padded the way Word places them), headers
and footers (per page, with separate first-page and even-page variants, and
fields evaluated), fields — page numbers (PAGE, NUMPAGES, SECTION,
SECTIONPAGES, PAGEREF, each following a section that begins its numbering again), what a document says about itself (AUTHOR, TITLE, SUBJECT, KEYWORDS,
COMMENTS, LASTSAVEDBY, CREATEDATE, SAVEDATE, PRINTDATE, DOCPROPERTY, FILENAME), counters (SEQ, with
its restart, repeat and format switches), references to a bookmark's text (REF), running heads
(STYLEREF), literal text
(QUOTE), and the clock (DATE, TIME) — each spelled the way its \* switch asks, in arabic, roman,
letters, ordinals, words, hex or dollars, and cased by Upper, Lower, FirstCap or Caps, with a
\@ picture deciding how a date reads, lists and numbering (decimal, letters, roman and bullets, nested levels with
independent counters and multi-level templates, hanging indents), images, inline and floating (PNG — interlaced or not — GIF, BMP, TIFF and EMF all read from scratch, JPEG passed through untouched — and decoded, in every
form including progressive, arithmetic and four channels of ink, where a TIFF holds one; the
four-channel pictures a printing press wants, as either a JPEG or a TIFF; transparency via a soft mask; square, top-and-bottom and no-wrap text flow around anchored
pictures), charts of columns, bars, lines, pies, areas and scatters (the plot area where the chart places it
or, where it does not, worked out from the room the labels need — including the room a label that
has to wrap takes; bars standing or lying, clustered, stacked or
stacked to the whole, sized by the gap and the overlap between them, a line curved through its
points the way Word curves one or straight where it says so, areas filled down to the axis or
banded one on another, markers of nine shapes at the size and on the grid Word draws them, a pie
centred and divided clockwise from the top; gridlines, both axis lines and their marks, and the
labels along each axis in the number format it asks for, a title over the chart and one on each
axis, a legend on any side, and a number at every point,
read from the numbers the chart part caches, with the axis scaled and the plot placed the way Word
does both where the chart leaves them to be worked out), equations (fractions stacked, slanted and
written on one line; superscripts and subscripts, alone and together, before what they belong to as
well as after; radicals with and without a degree; brackets of any character, grown to fit what they
hold out of the shapes the face keeps for the purpose; sums, integrals and the rest of the n-ary
operators with limits beside them or above and below; functions, matrices, aligned arrays, accents
and bars — set from the face's own MATH table, in the mathematical alphabets Unicode keeps for the
purpose, at the size of the text carrying them, on a line of their own where the markup says so), diagrams — SmartArt — drawn from the arrangement the document keeps of them (every shape
with its geometry, its fill and outline in colours named outright, by theme slot or as percentages,
and its text laid out into the rectangle the diagram set aside for it and set at its top, middle or
foot), watermarks of both kinds (a word set across every page of a section, behind the text,
stretched to fill the shape that holds it, turned, and painted see-through — which is a graphics
state of its own, since a PDF carries transparency there rather than in the colour; or a picture,
washed out to the contrast and brightness it carries), text boxes and the shapes they are a kind of, in both spellings (rectangles, rounded rectangles, ellipses
and triangles drawn as themselves and every other preset geometry as the rectangle it is bounded
by; filled and outlined in a colour named outright or taken from the theme; in the line of text or
anchored with the text flowing round them; and holding a document of their own — paragraphs and
tables, laid out into the box, clear of its insets and its outline, and set at its top, middle or
foot as it asks. A shape arrives wrapped in the compatibility element that offers the same drawing
twice over, once for a newer reader and once for an older; the newer is read and the older passed
over, so it is drawn once), table styles (all thirteen conditional formats — the whole table, the banding across
its rows and down its columns, its first and last rows, its first and last columns and its four
corner cells — resolved through the style's basedOn chain, gated by the table's w:tblLook in
either spelling, and giving a cell its borders, its shading, its margins, its alignment and the
formatting of the text inside it, with everything the table declares for itself winning over all of
it; a TableGrid table from a real document is ruled where Word rules it, which before this it was
not, having no borders drawn at all), tables (fixed and autofit column sizing, horizontal spans, borders, shading,
cell margins and vertical alignment, rows kept whole across page breaks), page size and margins
from sectPr, paragraphs and runs, xml:space handling,
line and page breaks, line breaking by the Unicode algorithm (so the scripts written without
spaces — Chinese, Japanese, Thai, Lao, Khmer, Burmese — wrap where Word wraps them), tabs
(left, centre, right, decimal and bar stops, with leaders), font family via theme resolution, size, bold,
italic, underline, strikethrough, colour, highlighting, caps, super/subscript, character spacing
and scaling, the background behind a paragraph or a run (w:shd, patterns included), the box round
a paragraph (w:pBdr, shared between paragraphs bordered alike) and round a run (w:bdr),
emphasis marks over a run's characters (w:em),
alignment including justification, indents including hanging, spacing before/after with
contextual spacing, line spacing (auto/exact/at-least), pagination, real font metrics with
.ttc support, the faces a document carries in its own font table read and deobfuscated so a file
brings its own type rather than falling back to substitutes, and Type0/CIDFontType2 embedding with
a ToUnicode map so text stays selectable.
The PDF carries more than the ink. A document outline hangs off the headings, so a reader opens to a navigation pane rather than a flat file (#66). A structure tree tags the page in reading order — headings, paragraphs, tables and their cells written as marked content, tied back through a parent tree, with the document's language on the catalogue — so the text has an order and a shape to it and is not a wall of glyphs to a screen reader (#67). And where the caller asks for it, the output claims and honours PDF/A-2b: an XMP metadata packet agreeing with the information dictionary, an sRGB output intent, and a file identifier derived from the body itself (#68).
Ten kinds of chart are drawn: columns, bars, lines, pies, areas, scatters, doughnuts, bubbles, radars and stock charts, the bars and areas clustered, stacked or stacked to the whole. What is there: the plot area, the bars or the line or the slices or the filled areas or the rings or the bubbles or the web, the markers at a series' points, the lines a stock chart draws between its series, the gridlines, the two axis lines and the marks along them, and the labels along both axes whichever way round they run — with the numbers read from the cache the part carries rather than from the workbook stored beside it, and with the axes scaled and the plot placed the way Word does both where the chart leaves them to be worked out; a title over the top and one on each axis, the one up the side turned on its end; a legend on any of the four sides; and a number written at every point, in the format the chart asks for and where the kind of chart puts it; and the trendlines a series carries, of all six kinds the format allows — straight, polynomial, exponential, logarithmic, power and a moving average — run forward or back past the data and forced through an intercept where the chart asks for it. What is not: a surface chart. That one is declined by decision rather than pending — a surface is a mesh over a grid, not a series of points, and shares nothing with the family above — so its page draws the room a 3-D chart stands in (frame, walls, gridlines, axes) and no mesh, rather than passing a 3-D line off as one (#104).
Those kinds are drawn in three dimensions where a chart asks for it (c:view3D): the box is
projected the way Word's camera projects it — a right-angle projection where c:rAngAx says so
(#97), and otherwise a perspective one, c:perspective being a field of view in half-degrees and
the eye backed off the scene by a law measured to a tenth of a percent across Word's whole
reachable range of angle and fov (#98, #141). The room is built around it: the back and side walls
and the floor, the gridlines ruled on them, and the depth axis for the series stacked into the page
(#99, #100). The bars are shaded boxes, three faces to a lightness, drawn back to front so the near
cover the far (#101, #110); the pie is that same tilted disc under that same camera — its top an
ellipse and its rim a cylinder wall, the two flattening and deepening together as the perspective
bites, projected rather than fitted family by family, deepened by c:hPercent, lit from the left,
and, where a slice is exploded, shrinking the whole disc to hold that slice's tip on the fill edge
(#102, #166); line and area are ribbons in depth (#103). What a value's height, and the box's own extent, come to inside all that — from the
categories, series, depth and c:hPercent — is each measured against Word rather than guessed
(#109, #113, #114, #116). Beyond the angles Word's dialog can state the camera clamps to its last
good law, which is visible only on a hand-edited file (#247). The pie's one measured-but-unclosed
part is the perspective pinch of its sector angles: its outline is Word's and its slices land in
the right families, but a slice at the back reads a shade larger than Word draws it, the exact
redistribution a two-variable surface (angle against fov) that no closed form yet fits — instrumented
and left affine rather than approximated (#166).
A title or a legend placed by hand is put where the chart asks. A c:manualLayout names a corner
as fractions of the chart's own width and height, and what sits at that corner was measured at two
placements each — which is what says the offsets are constants rather than shares of anything. A
title's first letter is 3pt across and its first baseline an ascent plus 1.43pt down; a
legend's first key is 5pt across, and the pitch and key drop of its entries are the automatic
layout's own arithmetic reused. A hand-placed one still takes its room off the plot, which the
probe's own axis labels confirm: they do not move.
A title says nothing about its weight and is bold all the same: Word's own styling of a chart
gives one to its title and to each axis title, and a document does not carry that styling. Measured
from chart-title-weight-probe, whose first page states no weight at all — Word's fourteen-character
title comes out 86.26pt wide against the 84.35 the same text measures regular, and its axis title
28.65 against 27.93. It is a default rather than an override, which the same probe's third page
settles: a title stating b="0" is left regular, and ours agrees with Word's there to five
hundredths of a point. Since the chart's own title is centred, getting the weight wrong moved it as
well as thinning it — 0.96pt on that title, and more on a longer one.
Two things were wrong in the automatic placement and were found by the same probe. A legend up a
side was centred on the middle of the frame rather than of the room below the title, which put
it half a title's height too high — 12.96pt on the probe, now exact. And c:overlay was read as
"do not draw" where it means "do not take room", so an overlaid legend was being dropped from the
picture altogether; Word draws one and now so does this.
A stated c:w and c:h alongside the corner is a different thing again: the legend is laid into
that box rather than run down from the corner. For a legend of one entry that is settled and
exact — the entry is centred in the box across and down, at two widths, two heights and two
corners, within six hundredths of a point. What is centred is the entry less its key: Word
centres a block 23.72 wide where the entry's own extent is 29.09, and the difference is a swatch.
Down, the baseline sits at the middle of the box plus the same key drop used everywhere else here.
A legend up a side does not simply grow to hold the longest name it carries: past 0.3635 of the chart's width, counting its key and the gap after it, Word wraps the name instead, and it is the wrapped width the plot gives way to. That share cannot be read off one chart — it was measured on five, from 240 points wide to 480, each saying the limit held the line Word kept and refused the next letter; together they put it in [0.3618, 0.3652) and the middle is used. Nothing had caught this because no fixture here carries a legend name anywhere near long enough to reach the limit, and the alphabet at ten point only just does.
The lines of one entry sit a label's line height apart, consecutive entries a legend pitch apart, and the baselines as a whole are centred on the middle of the frame plus the key drop — which with one line to an entry is exactly what a side legend did before, so no ordinary chart moves.
A name too long for the box it is given is wrapped, not cut short. It breaks at a space where
there is one and inside a word only where a single word will not fit — alpha beta gamma delta epsilon in a box of 64.8 points comes back as three lines and the alphabet in the same box as
three of its own, broken wherever they had to be. The lines are centred across by the widest of
them, and down it is the baselines that are centred rather than the line boxes: the first sits
half the distance between the outermost baselines above the middle of the box, plus the same key
drop used everywhere else, which for a single line is exactly the rule above. Where the box is too
short to hold every line, the last one Word draws ends in an ellipsis. chart-legend-cut-probe
settles all of it over nine pages.
What a line is broken against is the box's width less 12.28, which cannot be measured directly — where a line starts depends on the wrapping already done, since Word breaks first and centres what it broke. It is measured by its consequences instead: each page says the room held the line Word kept and refused the next thing it would have taken, and together they put the figure in (12.059, 12.5]. The middle is used, the way the cell margin Word puts in a table declaring none is settled at ten twips out of a possible seven to twelve.
A legend of several entries in a sized box is settled too, and it took a probe that moved one dimension at a time to do it — the earlier attempts moved both, so a rule fitted to one page kept moving another. They go in a single row where the row fits across the box and one to a row where it does not, all or none: three entries coming to 144.27 across make one row in boxes 180, 198 and 306 wide, and three rows in boxes 54 and 108. Height has no part in that choice, which is what four pages holding the width at 180 and taking the height from 21.6 to 172.8 are there to say.
Along a row, what the box has left over is shared out between the two margins and the gaps, all equal: the same three entries sit 14.20 apart in a box 198 wide and 41.20 apart in one 306 wide, so a constant gap fits neither. The one wrinkle is that the left margin is measured to the key and the right to the end of the words, which leaves them a swatch apart on the page; where that swatch splits is the only fitted number here, and a legend of one entry — where the split has no room to hide — pins its total.
Down a column the entries share one left edge, centred by the widest of them rather than each centring itself. And every row's first baseline sits the same distance into its share, set by the tallest entry in the legend rather than by its own: three entries in a box 129.6 tall get shares of 43.2 and first baselines 12.24 into each — including the two that are a single line, where centring those on themselves would put them at 24.24. A box whose entries are all one line does put them at 24.24. So the offset is a property of the legend, not of the entry.
A box too short for what it holds neither overflows nor shrinks anything: it stops drawing the
last entry, and asks again of what is left, so dropping one makes the remaining shares larger.
A box 54 wide gives the third entry three lines, and taking the height down through 108, 86.4, 64.8
and 36.72 makes Word draw three, two, two and one. What each entry needs is the tallest entry's
block and half a line besides — bounded by those pages to above two pitches and a bit over a third
and at most two and two thirds, with the middle used. The dropped entry still counts: two survivors
in a box 86.4 tall are spaced around the three-line entry that is not there. Yet across, the
survivors are centred on themselves alone — the box that keeps only the first entry centres that
entry on its own words. The asymmetry is Word's, and reproducing Word is the point.
chart-legend-box-probe settles all of it over twelve pages.
The lines a chart hangs from its points are drawn too — a drop line from each point down to the
axis, and a high-low line from the lowest of a category's series to the highest. The second was
always being read, and from the plot element whatever kind of chart it was; it was only ever
drawn for a stock chart, so a line chart asking for one got nothing. Three things about them were
measured. A drop line stops at the category axis rather than at the floor of the plot, which
only shows once the scale runs below nought. Where a category holds several series the line hangs
from the point furthest from the axis, not one line per series. And both are stroked with a
round cap, which is what Word writes (1 J) and what makes a short line reach half its own
width past each end.
Error bars are drawn too, of all five kinds the format allows — a fixed amount, a share of each
point, a multiple of the series' deviation, its standard error, and a distance stated for every
point and each side of it separately — reaching both ways or one, capped or not. Three things
about them had to be measured, and chart-error-bar-probe measures each on a page of its own.
The end cap is 4.5pt wide, read out of Word's own path geometry rather than guessed from the picture: it spans 57,150 EMU. A page that halves the plot and near enough doubles the type around it draws the same cap, so it follows neither.
The deviation is the sample one, dividing by n−1. The format does not say which, and the two are 15% apart over the probe's four points — Word reaches 15.55 of the value axis where the whole series' deviation would reach 13.46.
And the strangest of the three: a deviation's bars are not drawn about the points they belong to at all. Word draws every one of them about the series' mean, so four points get four identical bars at four places along the foot, saying where the middle of the data is and how far it scatters rather than anything about the point each stands on. All four cover the same 144..203 of the page where drawing them about their points would spread them over 111..236. A standard error does not do this — its bars sit on their own points, which the next page confirms to the pixel. The two are inconsistent with each other, and this reproduces the inconsistency.
Two things about a trendline had to be measured rather than read, and neither is in the format's
description of one. A line asked to run forward does not merely draw further: Word widens the
category axis to hold it and everything placed by category compresses to fit, so four categories
across a 252pt plot that put the first label centre 31.5pt in put it 21pt in once a trendline runs
two categories on — six slots where there were four. And a category's number for the purpose of
fitting is counted from one, not from nought. That second one is invisible almost everywhere:
shifting the origin of a free fit moves its constant term and nothing else, so every kind of
trendline draws identically either way. It shows up only where the chart forces an intercept,
which pins the origin — a line through nought spans 40 of the value axis in Word and 53.57 counted
from nought, and chart-trendline-probe's sixth page is the one that says so.
The fitting itself is the one part of a chart with no Word in it. A least-squares curve through
given points is not a matter of anybody's opinion, so it is tested against coefficients worked out
independently rather than against Word's export, and the three curved fits are the straight one in
disguise: taking logs turns y = a·e^(bx) and y = a·x^b into straight lines, which is what Excel
does and why those fits minimise the error in the logarithm rather than in the value. A fit that
cannot be taken — a logarithm of a nought or a negative — draws nothing rather than a NaN, which
Word also does.
Both spellings of a shape are read: the w:drawing Word writes today and the w:pict it wrote
before 2007 and still writes for a watermark. The older one says in a CSS-like style attribute
what the newer says in elements — its size, its position, what it is anchored to — names its
geometry by the element rather than by an attribute, and defaults its fill to white and its
outline to three quarters of a point of black. Where a document offers both, in a compatibility
wrapper, the newer is read and the older passed over so the shape is drawn once.
A shape turns and mirrors (a:xfrm's rot, flipH, flipV), and the wrap region beside it
stays the stated extent, unturned — measured: Word lets the turned corner overhang the text
beside it rather than pushing it out to the turned bounds; it fills with a linear gradient of any number of stops or with a picture
kept inside its path, and carries an outer shadow, drawn as the same path offset where the shadow
falls at the shadow's own solidity. A box told to size itself to its text grows at render, and one
told to shrink its text draws full size where the text fits — the stored scale is Word's cache of
its own computation, applied only where full-size content overflows, which is what Word was
measured doing. What a shape does not do yet: rotate the paragraphs laid out inside it (the
geometry turns; flowed text within stays upright), tighten its line spacing to fit
(lnSpcReduction — the font shrinks, the leading does not), fill with a pattern (declined —
Word's own gallery no longer offers one), or blur its shadow's edge.
Watermarks are drawn, of both kinds: the word, in the face and colour and half-solidity it asks for,
turned the way it asks to be turned; or the picture, washed out to the gain and black level it
carries. Both go behind everything else on every page of their section. A picture in a running head
is read from that part's own relationships, so a header and a body may number their pictures alike
— which they routinely do, both calling their first one rId1 — without either drawing the other's.
Table autofit is the one piece here that approximates rather than reproduces. Word's algorithm is
undocumented; ours measures each column's minimum (widest word) and maximum (unwrapped) width and
shares out the available space between them. It reproduces both behaviours that were measured —
content-width columns when the table fits, and a table filling the text area exactly when it does
not — and agrees with Word to 0.16pt on table-autofit-probe, but it is not derived from the real
algorithm the way the paragraph rules are.
A declared cell width (w:tcW) enters it as the width the column would like, which
table-width-probe measures five ways and this now follows exactly: widths that fit are taken as
they stand (72, 108 and 144 points come out as those); a column whose content will not fit the
width it asks for grows to hold it while its neighbours keep theirs (36/36/36 with an unbreakable
word in the middle comes out 36/142.56/36); widths adding to more than the measure are scaled down
together (three of 200 come out three of 156); a column asking for nothing is sized by its own
content beside ones that ask; and where two rows ask for different widths of one column, the wider
wins. Before this the declaration was ignored outright, which put a table of three declared columns
300 points from Word's.
A width asked for as a share of the table (w:type="pct", in fiftieths of a percent) is a
share of whatever the table's own width came to, which cell-percent-width-probe settles seven
ways. Of a table stating its width in points, the share is of that; of a table stating its width as
a share of the measure, it is of the measure through it. Of a table stating nothing there is
nothing to take a share of but the contents, and Word makes such a table as narrow as the shares
allow — a quarter, a half and a quarter round a letter each come out 5.28, 10.8 and 5.28, the
narrowest table at which a quarter still holds its letter — capping that at the measure, so the
same table with a column of text in its middle cell fills the 468 points instead. Shares falling
short of the whole are stretched to fill it and shares beyond it are taken in order until it is
spent; a share beside a stated 72 points and a column asking for nothing comes out 162, 72 and 90,
the share taken first, the statement kept, and the remainder going to the column that asked for
nothing.
Every column edge is on the grid, which is what makes those numbers come out as they do:
Word's quarter of 324 points is 81.12 and the next quarter 80.88, because 81 and 243 land either
side of a step. It is the edges that are snapped and not the widths, so three columns of one
declared width need not be equal — column-grid-probe's three fifty-point columns, fifty points
being 208 steps and a third, come out 49.92, 50.16 and 49.92 in Word and now here. Five of that
probe's six pages are identical to Word's: declared widths, awkward ones, a stated grid under a
fixed layout, widths scaled down to fit the measure, and halves falling the other side of a step.
All six pages are Word's exactly, and so is every column of the three other table probes. Getting the sixth — the one whose columns are sized by their contents — took two things measured elsewhere: that a cell's content width is rounded up to a whole twip before anything is shared out, and that what a cell's text is broken against is the width the arithmetic gave rather than the width that was drawn. A column is drawn on the grid and its text broken against the exact width, which is why a word can end a fraction past the column it sits in.
Not at all. break-tolerance-probe moves the measure a twip at a time — a twentieth of a point,
five times finer than the grid — past the width of a word with nowhere to break: ten capital Ms of
Times at twelve point, which are 106.6992pt wide. A measure of 106.7 holds them; 106.65 breaks
them, nine and one. A table cell is no different, and neither is a page.
That answers a question two earlier probes had raised the other way. A word had seemed to survive in a column a tenth of a point too narrow for it, which looked like tolerance and was not: the column was drawn on the grid while its text was broken against the width the arithmetic gave. Carrying both — the snapped widths for the drawing, the exact ones for the breaking — is what makes every column of the table probes come out Word's.
The font's own advances at the font's own resolution, and nothing else. text-measure-probe sets
every line against the right margin, so where a line begins is the margin less the width Word
measured, and repeats the same letter up to forty times so a single rounding is divided by forty.
Over eighty lines — Times at eleven, twelve and thirteen and a half points, and Arial at twelve —
every one of ours begins exactly where Word's does, the worst a ten-thousandth of a point.
The probe also lays a trap this repository walked into. A PDF records the widths it draws with in thousandths of an em, so reading Word's own export back gives 444 thousandths for Times 'a' — 5.328 points at twelve. Word did not measure it as 5.328: the font says 909 units of 2048, which is 5.32618. Two hundredths of a point a letter is nothing on a line and a whole step of the grid across a table column, and two commits here blamed a column that was a step out on "our measure running a hair above Word's". It runs exactly with Word's. What is left over in a table column is the column, not the text: Word's own page says it rounds what a cell wants up to a whole twip before dividing, and that it will keep a word in a column a fiftieth of a point too narrow for it rather than break it. Neither is implemented here, and both are written down in the backlog with their numbers rather than guessed at.
A table's own stated width (w:tblW) is met exactly, and table-preferred-width-probe settles
five things about it: the width is taken whether it is wider than the contents want or narrower; a
share is a share of the measure, so half of a 468 point column comes out 234; a width narrower
than the contents wraps them; a width wider than the page is not brought back to it — Word
writes such a table straight off the paper's edge, and so does this; and the width is divided in
proportion to what each column wants, each want being its content rounded up to a whole twip.
Five of the probe's seven pages come out exactly Word's. The two that do not are the same shape —
three columns of nearly equal content — and two-column-sweep-probe was written to settle whether
a better rule exists. It settles it the other way.
Two columns leave one edge between them, so where that edge falls is the whole of what Word decided. The sweep puts an 'i' and a 'b' either side of it and moves the table's stated width through thirty-six values, each of which pins the share the first column got to a window a grid step wide divided by that width. No share satisfies them all: 3450 twips of table needs at least 0.358261 and 3250 needs below 0.358154. Sweeping every ratio between 0.3555 and 0.3605 against every rounding of the column — exact, whole twip, half twip, up, down — the best any of them manages is 35 of the 36, and that one wants a ratio matching no measurement of the two letters. Word's division is not a fixed proportion of the table applied to a fixed pair of wants.
Ours is proportional to wants of 67 and 120 twips, each letter rounded up to a whole twip, which
lands 34 of the 36 — the best any natural ratio does — and misses only where its edge falls a
hundredth of a point the wrong side of a rounding boundary. The four tables at the end of the probe
put the same edge in from the right and agree exactly. TwoColumnSweepTests holds both halves of
that, including the crossing bounds, so nobody re-opens this expecting a tidy ratio.
How far inside its edge a cell starts its text was settled by measuring rather than by reading, and
is stranger than it sounds. table-inset-weights-probe holds the same one-cell table fifteen times
over — border weights from nothing to six points against no margin, then margins against a fixed
border, then no margin element at all — each on a page of its own so that no table's height carries
into the next one's position. Word's export of it gives three rules, and none of them is the
addition that would be guessed:
- Across, the inset is the greater of the cell margin and half the border, not their sum. Half a border falls inside the cell and half outside — Word's own border rectangles straddle the margin at every weight — so text starts at the border's inner edge unless a margin reaches further in.
- Down, the whole border is cleared rather than half of it. A six point border pushes the first line six points down and three points across.
- Declaring no cell margin is not the same as declaring one of zero: Word puts half a point in a
table that says nothing about the matter. The familiar 108 twips comes from the built-in
TableNormalstyle, which a hand-written document does not have, but what is left is not nothing. Word rounds every position to 1/300 inch, so the true value is somewhere between seven and twelve twips; ten is used.
A field that depends on where it falls cannot be worked out while the page it is on is still being filled, so a document holding one is laid out twice: the first pass records the page each field landed on and how the pages divide between sections, and the second uses it. Word settles its own page numbers the same way, and like Word this converges rather than being exact — a field whose text changes length between the two passes could in principle move to another page and be a page out. The second pass is only run for a document that needs one.
Word itself recalculates only these page-dependent fields when it exports, and leaves every other
field showing whatever it last computed, which is what the fields fixture's reference had to work
around: it is exported with its fields updated first, and tools/make-reference-pdfs.sh names the
fixtures that are. Anything this converter cannot work out keeps its cached result, which is the
honest answer — showing nothing would lose text the document has, and guessing would show something
it never said. COMPANY and MANAGER are among those: the Word this was measured against does not
evaluate them either.
A running head is the one field whose answer depends on where the pages fell rather than only on
what the document says, and Word's rules for it were read off its export of the styleref fixture:
a header shows the first paragraph of the named style on its page, a footer looks down its page in
the same way rather than up it, \l asks for the last one on the page instead, and a page holding
none of that style carries the last one before it — which is what walks a chapter title through the
pages under it. In the body the field looks backwards to the nearest one above it, and only where
there is none does it look forward. The style is named rather than identified: Word answers
STYLEREF Heading1 with an error telling the reader to apply the style, so an id is not a name
here even where it looks like one.
A table of contents is the one field whose answer is a run of paragraphs rather than a few words,
and it is worked out again rather than read back: a stale table is as wrong as a stale page number,
and a document that has never had one built has nothing to read back at all. The headings are
gathered by the outline level their style stands at — \o "1-3" says which levels, \t "Style,Level" names styles outright, \n leaves the page numbers off — and each entry is set in
the TOCn style named for its level, with a tab out to that style's right stop and the page the
heading landed on. A document defining no such styles gets an indent and a dotted leader of its own
so that the numbers still line up.
Two details of Word's own come from measuring its export of the toc fixture rather than from
reasoning. Tab stops are measured from the margin and not from the paragraph's indent, so the page
numbers of a second-level entry line up with a first-level one's rather than sitting eleven points
further out — which was wrong here until this fixture showed it. And the paragraph the field closes
in outlives it, empty: Word leaves the mark of it on a line of its own below the entries, set in the
document's default rather than in a table-of-contents style, which is the extra line between a table
of contents and the first heading under it.
An index is written in two halves, and both are implemented. Where a term belongs, the document
carries an XE field that draws nothing at all — it is there to be found, not read — and where the
index goes, an INDEX field gathers every one of them, sorts them, and lists each term against the
pages it was marked on. A term written Engine:analytical is a subentry and reads as analytical
under a heading of Engine, indented by its Index2 style; a page marked twice over is one page
number; \h asks for a line holding the letter each group begins with, \e and \l say what goes
between a term and its pages and between one page and the next, \t shows something else in place
of the pages ("see Engine"), and \f keeps two indexes in one document apart.
A document written for a mail merge does not carry its own data: it names a source — a
spreadsheet, an address book — that only the machine it was written on can reach. Converted as it
stands, its fields show what Word shows for the same document, each field's own name in guillemets
with whatever \b and \f ask to be printed around it, so a letter reading Dear «Title», reads
that way. Converted with a MergeRecord given to it, the same document reads as the letter itself
— and the text around a field then prints only where the field has something to print, so an empty
middle name takes its brackets with it. MERGEREC and MERGESEQ number the record, and the fields
that carry a merge from one record to the next (NEXT, NEXTIF, SKIPIF) draw nothing, which is what
Word draws for them.
The two fields that work something out rather than look it up are the last of them. IF compares two
things and chooses between two pieces of text — numbers as numbers, anything else as text without
regard to case, and the text an equality is asked against may hold * and ? wildcards. A formula
field is an equals sign and an expression: the five operators and their precedence, brackets,
comparisons, percentages, and the functions Word knows (SUM, PRODUCT, AVERAGE, COUNT, MIN, MAX,
ABS, INT, ROUND, SIGN, MOD, AND, OR, NOT, IF, TRUE, FALSE, DEFINED). In a table it reads the cells
around it, by direction or by name.
Three of its answers were measured rather than reasoned about, and none is what would be guessed. A
picture's # reserves a space where it has no digit to show, so five against $#,##0.00 comes
out as $ 5.00. A direction reads only as far as the numbers go — a column of 10, "n/a" and 3
sums to 3 from below it, not to 13 — while a range named outright reads all of it and passes over
what is not a number, so the same column averaged as A1:A3 is 6.5 rather than 4.33. And a formula
with no picture reads to two decimal places with the zeros at the end dropped: 10/3 is 3.33, 10/4
stays 2.5.
Pictures come in as GIF, BMP, TIFF, PNG or JPEG. Only JPEG passes through untouched — PDF's own image filter is JPEG, so decoding and re-encoding it would cost quality for nothing — and the rest are unpacked to samples by decoders written here: a bitmap of one to thirty-two bits a pixel, written from the foot up or the top down and run-length packed or not; a GIF through its colour table, interlaced or not, with the colour it treats as transparent becoming the mask a PDF carries separately; a TIFF at either end, in grey, colour or a palette, written in strips of rows or in tiles, with its channels together or each kept apart, packed with nothing, LZW, PackBits, Deflate or one of the fax encodings.
The two layouts are the format's other way of dividing a picture up. Tiles are rectangles rather
than rows, each written at the full size the tags declare however little of the picture it covers,
so the ones along the right and the foot carry padding that has to be left behind. Kept apart, the
channels are three pictures of one sample each rather than one picture of three, laid over one
another at the end. Both were written and read back through sips as well as here, which is what
said a tile's sides must be multiples of sixteen: it refused a file of eight-by-four tiles that
this read quite happily.
A fax is not written as pixels at all. A page of black on white is sent as the lengths of the runs its lines fall into, in a Huffman code the standard fixes rather than one built from the page — and a line may be written against the line above it instead, saying how far each change of colour has moved rather than where it is, which on a page of text is far shorter. All three encodings are read: lines written on their own, lines written either way with a bit each saying which, and lines written against one another throughout.
Getting that right took the two-way check further than anywhere else here. The tables are large and
mechanical, so a file is written with the library's own tables and handed to macOS's sips to
read: if a single code were wrong it could not read it. Then the same file is read back here. The
first version of the writer took the easy way and wrote every group 4 line in full, spelling out
its runs — legal, and passed. Rewriting it to write lines the way a fax actually does, against one
another, immediately produced a file sips read perfectly and this one got 171 pixels of 480
wrong: the reader was starting each line at its first pixel where the standard starts one pixel
before it, so a change at the very start of a line could never be found. No amount of round-tripping
against the easy encoding would have shown it.
An interlaced PNG is not one picture but seven, each a coarser or finer sieve of the whole — the first every eighth pixel of every eighth row, the last every pixel of every other row — and each is written as an image in its own right, with its own rows and its own filters over them. So each pass is unfiltered on its own and its pixels are then put where they belong, a few bits at a time where four of them share a byte.
A picture written with sixteen bits a sample keeps them, in a PNG or in a TIFF. A PDF carries either precision, so reducing one to eight would throw away exactly what it was written to keep. A PNG and a PDF both write the bigger half of a sample first, so those two bytes go through as they lie; a TIFF writes them the way round the rest of its numbers go, which has to be read from the file rather than assumed — a sample read from the wrong end is still a picture, only the wrong one. The single thing still reduced to eight bits is the colour table of a TIFF written through a palette, where the index is what the depth describes and the table is a few hundred entries whose lower halves no document has ever needed. What says the halves have not been transposed or the precision misdeclared is the page itself: a picture written at sixteen bits and described as eight comes out as noise, so the check is that a reader which shares nothing with this one draws the colour that went in.
Reading a format is easy to do nearly right, so these are checked two ways. Files whose every pixel
is known are built by the tests and read back — which is the only way to reach a bitmap written
upside down or a GIF written in four passes — and then the same picture is turned into each format
by macOS's own sips, and by tiffutil for each of the TIFF packings, and read back again: what
comes out has to be what the PNG it was made from holds, to the sample. Both found real faults. A
bitmap's height is at a different offset in the modern header than in the old one, and a TIFF
written big end first keeps a small number in the high half of the four bytes its tag reserves.
A metafile is not a picture at all but the record of one being drawn — move here, line to there, fill with this, write that — so reading one is an interpreter rather than a decoder, and what comes out stays a drawing all the way to the PDF, whose own operators write it out again. That keeps a chart sharp at any size a reader looks at it, and it keeps the text inside one selectable. What is handled is what a picture in a document is made of: paths and the shapes that are shorthand for them, the pens and brushes that colour them, the fonts and the text, and the bitmaps a drawing can carry, and the clipping a drawing keeps its ink inside — the intersect and exclude rectangles and the regions built of rectangles, honoured as PDF clip paths (#69). What is not is the rest of an interface built to drive a screen: raster operations and palettes, declined by decision — the one is a screen idiom a document's picture does not use, and the other matters only to old paletted bitmaps no Office export writes into a metafile.
A metafile written by anything modern carries the same drawing twice — once in those records, and once in the newer GDI+ ones that travel inside their comments, a format smuggled through a format. Both are read, and where a file has the newer ones they are what draws it. That is what they are for: the old records beside them are a copy left for readers that have never heard of the new, and where the two differ at all it is the new half that is the fuller. They are read in the one order the file puts them in, because the old records are not always only a copy — a file may hand the drawing back to them part way through, for something the newer interface had no way to record, and says so where it does. From there the old records draw until the newer ones resume, which is what the specification asks for and what the file means.
Preferring them used to be impossible to justify, and the reason is worth keeping. Word for Mac renders classic metafile records and draws nothing whatever for EMF+ ones — its export of a file holding only those is a blank space — so there was no second implementation here to read EMF+ against, and everything else in this project is measured against Word. What settles it is that a file carrying both carries one picture twice: Word draws the old half of the metafile fixture and this draws the new, so comparing the two pages compares this reading of EMF+ against Word after all. The drawing's text lands within 0.07pt across and 0.12pt down of where Word puts it, and the two pages agree on 97% of what is covered and what is left as paper — the rest being the edges along which no two renderers ever agree. That is the same standard the rest of the metafile work is held to, and it is now the newer records being held to it.
Its text needed one measurement. A drawing says where the top of its text goes and a PDF says where the baseline goes, and the distance between them is the height of the characters themselves: the em less what hangs below the line, with the leading above them left out. Word's own rendering of the fixture puts the baseline 10.98pt below the point the record names, at 14pt Times New Roman, and the em less the descent is 10.97pt of it.
A drawing is also the one thing here that cannot be checked by reading text positions out of a PDF:
a chart could be drawn upside down and the comparison would not notice. So tools/rasterize.swift
draws the page with macOS's own PDF reader and the pixels are looked at — which is how the drawing
in the fixture is known to be where it was put, in the colours it was given. It found a real fault
in the fixture rather than the reader: a metafile says how big it is twice over, in the frame it
declares and in the resolution of the device it was recorded for, and the test's writer had been
writing the two at odds. Word believed one and this believed the other, and both were right.
A TIFF may also hold a JPEG rather than pixels, and that one is not decoded at all: a PDF carries a
JPEG as the file it already is, so the work is putting the file back together rather than taking it
apart. A TIFF divides one in two ways — the older keeps the whole file in a tag of its own, and the
newer keeps the tables every scan shares apart from the scan itself, so that a picture in many
strips need not repeat them, which makes the file the tables without their end followed by the scan
without its beginning. What comes out is handed to sips to read, because a file put back together
wrongly still parses as far as its header: reading its size back would prove nothing.
A picture divided into several JPEGs, one to a strip, is the one case that has to be decoded: the strips are separate files with nothing in common but the picture they are parts of, and a PDF has no way to be handed several of them as one image. So there is a baseline JPEG decoder here after all, used for that alone — a JPEG holds not the picture but a description of it, each block of eight by eight pixels written as how much of each of sixty-four waves it is made of, and reading one is that in reverse. Two decoders never agree to the sample, since they round the same sums differently, but ours and the one macOS uses agree to within six levels of 255 on the same file, and a strip put in the wrong place would be out by far more than that.
A JPEG's numbers may be written all at once or a little at a time, and both are read. A sequential file gives every wave of a block before moving to the next; a progressive one gives the coarsest waves of the whole picture first and returns for the rest, and may send the high bits of a number in one pass and its low bits in another — so the numbers are gathered and turned into pixels only once the last pass has been read. A progressive file is where a JPEG is most easily read nearly right: a pass misread leaves a picture that is still a picture, only softer or blockier than it should be. So these are tested against real ones rather than any this could write — macOS ships several, written by encoders that had no idea this existed — and read against its decoder they agree to three levels of 255, mean under a tenth, at sizes up to 2048 square.
A JPEG may also be coded arithmetically rather than by code tables, which almost nothing writes — the method was patented for most of the format's life — but files exist and there is no reading round one. Nothing here recovers from a mistake: Huffman codes resynchronise at the next symbol, whereas an interval narrowed by the wrong probability is wrong for ever after, so a decoder of this is either right or produces noise. That makes it testable in a way little else is. Recoding a JPEG from one entropy coding to the other changes not one number in it, so the same picture is written both ways by an encoder that shares nothing with this, and the arithmetic file has to decode to what the ordinary one decodes to — not close to it, equal to it, sample for sample. It does, for the sequential and progressive forms, with and without restarts. Reading a whole picture correctly is not evidence that the hundred and thirteen probability states are right, it is proof of it, which is worth having: two columns of that table transposed reads eleven decisions correctly and then quietly falls apart.
A picture bound for a printing press holds four channels rather than three — not the light a screen adds up to a colour but the ink a press lays down to take light away, and the fourth is black because the first three together make a muddy brown rather than black. Both a JPEG and a TIFF may hold one, and both are read. The catch is which way up: Adobe's tools write such a JPEG with nought standing for all of an ink rather than none of it, and every one in practice is one of theirs. Nothing warns a reader but a marker beside the picture. Read without noticing, the page comes out in exactly the wrong colours, which is easy to do and — in a document of photographs — not always easy to see. What is decoded here is turned back as it is read; a JPEG passing through untouched cannot be, so the PDF is told instead, and told only where that marker says so.
That is the half no amount of reading samples can settle, so it is tested by drawing the page: a picture of known inks goes into a document and macOS's own PDF reader draws it, and each quarter has to come out the colour its ink stands for. The inks were chosen so that no two channels hold the same value, which is what says they arrive in the right order as well as the right way up. A TIFF of a press's own inks rather than the four is reported rather than drawn in colours that are not its own.
Two things have no second opinion behind them. The older way of holding a JPEG inside a TIFF:
sips will not read a file written that way at all, so what is tested is that the JPEG comes back
and no more. And the four-channel files whose colours were turned into brightness before coding —
nothing installed here writes one, so that path is transcribed from libjpeg's own conversion rather
than checked against a file, and it is the one thing in these formats that has never met a real
example.
A note too long for the room left under the page its reference falls on is divided rather than
moved: a note belongs to the page its mark is on, so what will not fit goes to the foot of the page
after. Where it divides was read off Word's export of footnote-split-probe, and it is simpler
than it looks — the note takes everything left under the line that refers to it, the body stops
there, and the rest carries over. Nineteen of that note's twenty lines fit, which is what this
produces, with every line of both pages within half a point of Word's.
Two things follow from it that are worth stating. The rule above a carried note is drawn right
across the measure rather than the two inches drawn above a note that begins where it stands, which
is Word's way of saying without words that what follows is the end of something; the document keeps
that second rule as a note of its own, in the same way as the first. And a note may outlast the
document it belongs to — one referenced near the end and long enough to fill several pages has no
body text left to carry it — so pages are made for the rest of it, holding nothing else. Word does
the same, which footnote-overrun-probe is there to have asked: its second page has no body at all
and the last thirty-seven lines of the note at the foot of it.
A section of columns closed by a continuous break has its last page evened out: the columns come to
much the same depth rather than the first being full and the last empty, which is what a continuous
break is usually inserted to do. A section closed by a break to a new page is not evened out, and
neither is the last section of a document — columns-balanced holds all three cases and Word's
export of it says so.
Where the columns divide had to be measured too, since the obvious rules disagree. Word divides thirty-five lines eighteen and seventeen, and ten lines five and five: a column takes lines while it is still short of the depth the content is to be divided at, so the line that reaches that depth is the last one in. Rounding either way gives one of those two answers and not the other.
In a section of columns a note goes under the column its reference is in, set to that column's
measure and ruled off by a separator of its own. Each column keeps its own area, so what one column
gives up for its notes is not taken out of the next: in footnote-columns, whose first column
carries two notes and second one, the columns stop 13.4pt apart, and that difference is what says
the space comes out of the column rather than the page.
A section may also ask for its notes under the last line of text rather than at the foot of the
page. On a page whose text reaches the bottom margin the two are the same place, which is most
pages; on one whose text stops early — the last page of nearly every document — the notes come up
with the text, and footnote-beneath-text is a document of exactly those two pages.
That fixture answered a question nobody had asked, and corrected this. A line carrying a reference
never moves to make room for its own note. footnote-carry-probe puts a reference on the very last
line a page has room for and Word keeps the line there, squeezing the whole note in beneath it;
the other fixture puts one where there is no room left at all, and Word still keeps the line and
carries the whole note to the next page under the wide rule. This used to move the line instead,
which is the obvious way to keep a note with its reference and moves body text Word leaves alone.
A section may begin the page numbering again, which is what a document with a preface does, and at
a number of its own rather than necessarily at one. What follows the restart and what does not was
read off Word's export of page-numbering-restart, a document of three sections whose footer counts
pages four ways at once: the page number follows it, the total counts the document through
regardless, and a reference to a page names the number that page is printed as rather than where it
stands. The properties on a section break describe the section it closes rather than the one it
opens, which is what makes a fixture for this worth writing rather than reasoning about — the first
number stated belongs to the pages before it.
Endnotes gather at the end of the document, or at the end of each section where the document asks for it, which is what a book of chapters does with them. Each group is written where its section stops, before the break that opens the next one, so the notes of a chapter belong to the pages of that chapter.
Where the instruction lives is the reverse of everywhere else, and cost a fixture to find out. Every other thing about how a note is set is read from the section; this one Word reads from the settings part and nowhere else. A document asking for it in its sections alone — which is what the format's own reading suggests, and what this fixture was written as — comes back from Word with every note at the end regardless. Word's own writer puts it in both places, which is what says which of them it believes: setting the option through Word itself and reading back what it wrote is how that was settled.
Where a document numbers its notes again from the beginning, two things had to be measured rather than read. The first is where the instruction lives: the format allows it in the settings, as a default for the whole document, and in each section's properties. Word reads only the section. A document asking in its settings alone for its notes to begin again on every page comes back from Word numbered straight through, so that is what happens here, and there is a test that says so.
The second is per-page numbering itself, which cannot be settled while the page it depends on is still being filled — the line carrying a mark may yet move to the next page and take its number with it. So it is done the way page numbers are: the first pass records the page each mark landed on and the second numbers from that, converging rather than being exact for the same reason. Per section needs none of that, since a mark is composed inside the section it belongs to. Endnotes restart by section too, and are still gathered at the end of the document, so their numbers repeat down the one list — which looks wrong until you check, and is exactly what Word does with the same document.
Not yet: for pictures, nothing a document holds, and nothing left of these formats at all. What remains elsewhere is Hangul, whose syllables are composed rather than shaped, and the one kind of Apple attachment that names points on the outlines themselves rather than in a table — no face on this machine asks for it.
A line of Hebrew or Arabic is not a line drawn backwards. Text is stored in the order it is read and drawn in the order it appears, and the two part company the moment a line holds both directions, which nearly every real line does: a number is written left to right whatever is around it, and so is a Latin name inside a Hebrew sentence. The Unicode bidirectional algorithm decides which way each character runs and what order the line is drawn in, and it is implemented here in full — the paragraph rules, the explicit embeddings and isolates, the weak and neutral characters, the bracket pairs, the levels and the reordering.
The tables it reads are generated from the Unicode character database rather than written, by
tools/make-bidi-tables.py, because they are a hundred thousand answers and any of them typed by
hand would be a chance to be wrong. The library carries no dependencies, so the output is committed
as source.
This is the one part of the converter with a reference implementation to hand, and it is checked the way that deserves: GNU FriBidi implements the same standard and shares nothing with this, so fifteen hundred lines built at random out of Hebrew, Arabic, Latin, digits, brackets, marks and directional characters go through both and are compared level for level, and another eight hundred are compared on where every character of them ends up. They agree on all of it. The comparison found two real faults while it was being written — the backward searches of two rules end on what a sequence sits after and not only on a character, and the formatting characters count as whitespace for the rule that resets the end of a line — and a third when the drawing order was compared: a mark drawn on a letter has to be put back after it when its run is turned round, which is a rule that matters to anything that draws rather than merely reorders.
Hebrew is laid out with all of that behind it. A paragraph that says w:bidi begins at the right
margin; the words of a line are placed in the order they are drawn rather than the order they are
stored; a word that runs right to left is drawn from its own far end, with the marks kept on the
letters they belong to and the brackets facing the way the reader is going.
A run is stored in the order it is read all the way to the writer, and turned round there, as glyphs. That is not a detail of where the reversal happens. Which letter joins to which, which mark belongs to which letter, and which letters may be written as one shape are all questions about the text, and a shaper handed a word backwards answers all three backwards. What can be turned round safely is the glyphs, because each carries its own advance and its own offset and both are measured from where that glyph itself begins. What has a direction of its own keeps it: a number inside a Hebrew sentence reads as it was written, and so does a Latin name. Hebrew inside an ordinary left-to-right paragraph is turned round too — which way a paragraph runs says where its lines begin, not which way its characters go.
Word's export of the hebrew fixture says the lines begin where Word begins them, to within a
hundredth of a point, in paragraphs running both ways and with numbers, Latin, brackets and
punctuation inside them. What that comparison cannot say is what the lines say: Word writes a
line of Hebrew as many runs, and encodes some of them as pairs whose map back to characters gives
the two the other way round, so the text this reader recovers from Word's file is not quite the
text Word drew. The drawn order is checked instead against the algorithm's own answer, and that is
checked against another implementation of the standard.
A font is chosen per character rather than per run, because most fonts hold very few of the characters there are: Arial Hebrew has no Latin letters at all, Times New Roman has no Japanese, and a document written in Hebrew with an English name in it names one font for both. Asked for a character its face has not got, the run is set in two — the rest of it where it was, and that character in a face that can draw it. A converter that does not do this loses text the document plainly holds, without failing and without saying so.
Which face is borrowed is a matter of taste rather than of correctness, and the taste is the document's: the substitution chain a missing family already walks is walked again, and only where none of those can draw it is everything else tried, in a fixed order so that two machines holding the same fonts do not disagree. Where nothing at all can draw it the run keeps its own face; the document is then short of a glyph, which is the truth, rather than short of a page.
That is also why font-fallback is compared to Word on where its text goes rather than on how wide
it is. The lines begin exactly where Word begins them, and every character a font could not draw is
on the page; the borrowed faces are not the same width as Word's, because Word's choice is its own
and not discoverable from the document.
A mark is drawn where the font says it goes. An accent, a Hebrew vowel point, an Arabic dot: none of them has a place of its own and none can be drawn by advancing the pen. The font gives the mark an anchor and the letter an anchor, and the two are brought together — a movement of the mark alone, which the pen does not know about, so the letter after is set as though the mark were not there. A mark drawn on a mark is placed against that mark rather than the letter, which is how a letter carries two.
Where no movement along the line can express it — a point below a letter, an accent above one — the text is raised for that glyph and put back down after, which is the only thing a PDF has to say it with.
There is a reference implementation for this too, and it is used the same way. HarfBuzz shapes text for nearly everything that draws it and shares nothing with this; asked for the same characters in the same face it gives the same glyphs, the same advances and the same offsets, to the design unit, for Latin accents, Hebrew points and pointed Arabic alike — including the marks written over a shape that stands for four letters. Both are asked for the run in the order it is drawn, so the numbers are compared as they are rather than translated first.
For the Indic and South-East Asian scripts it is used with one thing to be careful about. Several of the faces macOS ships for them carry two complete descriptions of how to shape them: the OpenType tables every other platform reads, and Apple's own state tables. HarfBuzz prefers Apple's wherever a font has them, so asking it about one of those faces answers a question about a table this converter does not implement and Word does not read either. The tests therefore ask it about a copy of the face with those tables taken out. The two mostly agree; where they do not, the difference is worth seeing rather than hiding — Khmer Sangam MN's OpenType tables write a consonant and its vowel as a shape plus a blank, and its state machine deletes the blank instead.
Eighty words across sixty-two scripts are compared that way, glyph for glyph, advance for advance and offset for offset, and agree on all of it. Word's exports of the three fixtures agree about the page: every line begins exactly where Word begins it, and no line is more than a tenth of a point wider or narrower. What those comparisons cannot say is what the lines say — a shaped syllable is one glyph standing for several characters, and Word's file maps them back to whatever code the glyph happens to sit at, so a line of Devanagari comes out of it as "नम#$".
For Apple's tables the same comparison holds: twenty-six words in ten faces for the shaping and thirteen more in eleven for the placing, and Word agrees about the page on every line of the fixture to four hundredths of a point. Three things are kept out of it because Word does something else with them — or nothing at all. Asked for a line of Malayalam in a face whose positioning is Apple's, its export holds nothing where the line should be. Asked for Thai in Thonburi, it draws the line in a font of its own. And for one Devanagari cluster its reading of the same table comes out two points wider than HarfBuzz's. Which of the two Apple's own engine agrees with is not a question this machine can answer, so the difference is recorded rather than resolved.
Reading those faces turned up a fault of ours that had nothing to do with shaping. They carry their family name several times over in several languages, and the first record of the right kind is not the English one: Gujarati MT calls itself ગુજરાતી એચટી. A document naming "Gujarati MT" then matched nothing at all, and its text was drawn in whichever face was borrowed for it. The English record now wins.
For the universal engine the Word comparison covers one script rather than seventy, and that is Word's limit rather than a choice: asked for Tibetan, Javanese or Cham on this machine, Word draws the letters side by side without stacking or reordering anything. Sinhala it draws properly, and agrees with this converter to four hundredths of a point on every line — including the line whose vowel is written on both sides of its letter. For the rest, HarfBuzz is the only reference there is, and it is the same reference Word's own engine was written against.
Arabic joins its letters, which makes the shape of a letter a fact about its neighbours rather than
about itself. Most letters have four — alone, opening a word, inside one, ending one — and a
handful join only on the right, which is why a word can end in the middle of itself: nothing after
alef or dal joins back to them. A mark written over a letter must not break the join, and the four
shapes are four glyphs of the same character, chosen through the font's isol, init, medi and
fina features. Some pairs may not be written as two at all: lam followed by alef is one shape,
and drawing them apart is a spelling mistake rather than an ugly line.
Two things about ligatures took measuring rather than guessing. A font says, per lookup, whether a match may reach across the marks between the letters, and both answers are needed in the same font and the same word: the lookup that writes lam, lam and heh as the one shape for the name of God reaches across the vowels written over them, while the lookup that combines a shadda with the vowel beside it is matching marks and must not skip them. And a mark on such a shape is placed by a further table again — the ligature offers a place for each of the letters it stands for, and a vowel over the second lam is not a vowel over the first. Which letter it belongs to is how many of the shape's letters stand before it in the text.
What is drawn as one shape is still read as four characters. A ligature is written into the PDF's map back to text as everything it stands for, so a word joined on the page can still be searched for and copied out as the word.
Word's export of the arabic fixture agrees on all of it: every line begins within four tenths of
a point of Word's, and along the line each glyph stands where Word stands it to a hundredth. What
that comparison cannot say is what the lines say, for a nearer reason than Hebrew's — Word writes
Arabic as the presentation forms, one glyph to a run, so its file reads back as characters nobody
typed, and the name of God, which it draws as the single glyph the font holds for it, reads back
out of Word's own file as the letter J.
An Indic syllable is drawn neither in the order it is stored nor one shape to a letter. A vowel may be written to the left of the consonant it is pronounced after though it is stored after it; consonants with no vowel between them are written as one stacked shape; and an r at the head of a cluster is written as a small mark at the end of it. A converter that walks the characters and looks each one up does not draw an ugly line — it draws a different word.
None of that can be settled character by character, and the syllable is the unit throughout. It is divided out of the run; the consonant the rest hangs from is found by asking the font what shapes it has — the rules are written in terms of what a font can do, so there is no way round asking; every part is given a place in the visual order and sorted into it; the font's rules for making conjuncts are applied one at a time; and then what those rules managed to make decides where the vowel and the r finally go. Sort, ask the font, sort again. Nine scripts follow it — Devanagari, Bengali, Gurmukhi, Gujarati, Oriya, Tamil, Telugu, Kannada and Malayalam — differing in where a repha ends up, how it is asked for, and which side of the base a below-base form may appear on.
The specification for those scripts was rewritten, and a font says which set of rules it was drawn against by which of two names it files its script under. Both are implemented, because both are shipped: this machine has Shree Devanagari 714 and Arial Unicode MS written to the older rules, and Devanagari Sangam MN to the newer, and the two are shaped differently — under the older rules a joining mark after the base is moved to the end of the syllable, the below-base feature is not applied before the base, and what the font is asked about a pair of letters is asked of the pair in company rather than standing alone.
Most of the writing systems descended from Brahmi are not given rules of their own at all. There are too many of them, they are alike enough, and what differs between them is what the font already describes — so one engine shapes all of them, working from what each character is rather than from which script it belongs to. Something to build on; a vowel drawn above, below, before or after; a consonant written under the one before it; a mark on a mark. Classify the characters, divide the run into clusters on that basis, ask the font for its shapes in a fixed order, and move the two things that are drawn away from where they are stored: an r at the head of a cluster, which is drawn as a mark at its end, and a vowel written to the left of the letters it is pronounced after. That is the whole of it, and it shapes some seventy scripts — Sinhala, Tibetan, Javanese, Balinese, Cham, Newa, Chakma, Adlam, Egyptian hieroglyphs — none of which has a line of code to itself.
The classification is not a property in the database. It is worked out from five that are — what
part of a syllable a character is, which side of its consonant it is drawn, whether it joins like
Arabic, whether it is ignorable, and its general category — by rules Microsoft publishes, together
with the overrides Microsoft publishes for the characters the database has not caught up with.
tools/make-use-tables.py does that once and writes out the answer.
Three things about it were only found by comparing. A joiner does not divide a cluster: the grammar is read over what is visible, or the very character written to ask for two letters to be joined would separate them. A font that files its rules under no script in particular is not written to this engine and must not be shaped by it — Noto Sans Tai Tham is such a font, and draws a left-side vowel by moving the glyph, so reordering the characters first moves it twice. And a vowel written on both sides of its consonant at once is stored as one character and cannot be drawn as one: it is taken apart first, into the two halves the database says it is made of, and only the left half is moved.
Khmer and Myanmar descend from the same writing and reorder for the same reasons, more plainly. Khmer marks a stacked consonant with a character of its own rather than by the absence of a vowel, and moves an r written under a consonant to the front of the whole cluster; Myanmar decides which part of the syllable each character belongs to in one pass and sorts, with the medial r and the left-side vowel going before the consonant they are written on. Thai and Lao reorder nothing at all: they stack their vowels and tone marks above and below, and store every character in the order it is drawn, so what they need is the marks put where the font says.
What each character is to a shaper, and where round its consonant it is drawn, is generated from the
Unicode character database by tools/make-indic-tables.py rather than typed. Not all of it can be:
where a matra ends up depends on the script as well as on the side it is written on, because the
sorted order is a visual one and the scripts stack their parts differently. Those per-script answers
are in the generator, from the OpenType script development specifications.
Not every face describes its shaping in OpenType. A hundred and sixty of the ones on this machine
carry Apple's morx table and no GSUB at all: Devanagari MT, Gujarati MT, Gurmukhi MT, Thonburi,
Geeza Pro, Corsiva Hebrew, and the whole of Helvetica, Palatino and Optima. A converter that reads
only OpenType draws those scripts as rows of unjoined letters, and there is nothing in the file to
fall back on.
It is the older idea and the more general one. Where OpenType says "these glyphs in this company become those glyphs", this says "in this state, a glyph of this class takes you to that state, and on the way you may mark this one, swap that one, or write several as one" — one machine expressing what OpenType needs four kinds of lookup for, and two more besides: rearranging glyphs, and inserting ones the text never held. All five kinds are read here, and the reading is used only where the font has no OpenType tables. A face carrying both carries the same shaping twice, and the OpenType half is the one every other reader of the file will use.
Where those faces go on to say how far apart the glyphs go, they say that in Apple's tables too, and those are read as well. Two quite different things live in the one table. Most of it is kerning — by naming pairs, by naming classes of glyphs, or by a machine that keeps a stack of what it has passed and moves several of them at once — and it is applied only where the document asks for kerning, as everywhere else here. The rest is attachment: a machine that marks a letter and fastens what follows to it by naming a point on each out of a table of anchors, which is how these faces put a vowel sign on a consonant. That is applied always, because a mark that is not fastened is not merely unkerned but in the wrong place.
Its kerning is shared between the two glyphs rather than taken out of the first one's advance: half moves the glyph on the left and half moves the one on the right, which is drawn half a kern along as well. OpenType expresses the same thing the other way, as a shortening of the left glyph alone. The two give a pair the same width and put the second glyph in different places, and each face is drawn against one of them.
Two things about it cost an afternoon each. A machine that writes several glyphs as one leaves the others behind marked as gone, and they have to be swept up before the next machine runs rather than at the end — left in place, they sit between the pair the next machine is looking for and the run comes out with its shapes half made. And which way round a subtable reads the run is decided by a flag saying so, not by comparing the flag that says "logical order" against the direction of the text: getting that wrong shapes Arabic with every letter's join one place out.
Where a line may be broken is not a question about spaces. Chinese and Japanese are written with none at all and break between one character and the next; Thai, Lao, Khmer and Burmese have none between their words either; and even English has spaces that may not be broken at — the one in "10 kg" — and places without a space where a break is allowed, as after a hyphen. A converter that looks for spaces draws a line of Japanese straight off the edge of the page, which is what this one did.
The Unicode line breaking algorithm decides it from a property of each character and a list of rules
about pairs, applied in order, the first that matches winning. The property is generated from the
database by tools/make-linebreak-tables.py, and the rules are checked against the file Unicode
publishes for exactly that purpose: 7310 of its 7654 cases, every one that does not turn on a script
needing a dictionary, and all of them pass. Half of what that file caught was rules read backwards
or applied to the wrong side of a pair — no break before an opening bracket rather than after one,
glue looked for past the spaces rather than beside them — and the sort of thing no amount of reading
the text again would have shown.
The 344 it does not answer are the scripts written without spaces and without a break between every character. For those the algorithm says to consult a dictionary of the language, which this converter has not got. What it does instead is what Word does with the same paragraph: break between one syllable and the next. The marks are folded into the letters they are written on, along with the handful of Thai and Lao vowels that are letters in their own right but sound after a consonant and cannot begin a line; the four vowels those two scripts write before the consonant they are sounded after keep hold of it, as does the Khmer sign that turns the letter following it into a subscript; and what is left standing is a letter that may begin a syllable, and so a place a line may be broken. It is a coarser answer than a dictionary would give — more places than a Thai reader would choose — but every place it offers is one, and a greedy filler picks the last that fits.
Which is where the measurement is. Given the wrapping fixture — a paragraph of Japanese, one of
Chinese and one of Thai, none with a space in it — Word and this converter break all three into the
same lines, to a hundredth of a point in width, kinsoku and all: no line begins with a full stop, a
comma or a closing bracket, and the Thai breaks where Word breaks it, in the middle of a word.
Text reaches the page as glyphs rather than as characters. Between the two stands a shaper: it takes a run and a face and gives back the glyphs that draw it, each carrying its own advance and the character it came from. Everything downstream reads that — a line's width is the sum of its glyphs' advances, and what is written into the page is those same glyphs.
What the shaper does is decided by the writing in front of it: a character takes the glyph the font's character map gives it; a letter of a script that joins takes the shape its neighbours call for; letters the font says may not be written apart are written as one; a pair is drawn closer together where the face says so; and a mark is moved onto what it belongs to. The point of it is not cleverness but shape. A character is not a glyph — one may need several, several may need one, and which glyph a character takes can depend on its neighbours — and a pipeline that carries characters as far as the page cannot be told the difference. This one carries glyphs, so it can be.
It also puts right something that was true before it: a line's width and the kerning written into the page were worked out by two separate walks over the same text, one in the measurer and one in the writer. Two walks that must agree are two walks that can disagree. There is one now.
Underneath it is a reader for the two tables a font describes its own shaping in. GSUB says what
glyphs a run becomes and GPOS says where they go, and the two are the same machine pointed at
different ends of the problem: both walk a run, both pass over what a lookup says it cannot see,
both have rules that match on what stands before and after, and in both a matched rule does not act
itself but names other lookups to run at places inside the match. So the walking, the skipping, the
matching and the nesting are written once, and only the acting is written twice — substitution of
one glyph for another, of one for several and of several for one; adjustment of a single glyph, of
a pair, and of a mark onto the letter, the ligature or the mark it belongs to.
Two of its parts are what make a complex script work at all. The first is GDEF, which says what
each glyph is: a rule about two letters that must still fire with a vowel written between them says
"ignore marks", and which glyphs are marks is a question only the font can answer. The second is
the script list, which had been passed over on the grounds that a lookup fires only on the glyphs
it covers and those are its own script's. That is true of nearly every feature and false of the one
whose whole purpose is to draw the same glyphs differently: locl, taken from every script at once,
draws Hindi in Marathi's letters. The script a run is in is now asked for by name.
Hinting is kept, because Word keeps it: its own exports carry cvt, fpgm and prep in every
subset they embed. It cannot be subset in any case — control values are reached by index and
function numbers are worked out as the instructions run, so which of them a glyph needs is not a
question that can be answered without running them. ConversionOptions.DropFontHinting takes all
three tables out along with the instructions inside each glyph, which more than halves the file at
no cost to the shapes; it is off because it is a departure from Word rather than a step towards
it.
CFF subsetting is the one thing here with no Word reference behind it: every PostScript-outline
face on this machine is for a script the converter cannot shape, so there is no document Word
could be asked to render. It is checked instead by fontTools, which reads the rebuilt font and
draws its glyphs — a subset that emptied a subroutine something still calls parses perfectly and
draws rubbish, so executing them is the check that matters. Install it with python3 -m pip install fonttools; without it those tests report and skip, and N8PDF_REQUIRE_FONTTOOLS=1 makes absence
a failure.
ContentCoverageTests asserts that every text run and every placeable image in a document reaches
the PDF, so an unimplemented construct fails loudly instead of vanishing from the output.
Fixtures/Real/ holds documents Word wrote. tools/make-real-fixtures.sh takes the seed
documents defined in Fixtures.RealSeeds, opens each in Word and saves it straight back out,
which rewrites the package in Word's own terms — a styles.xml carrying several hundred latent
styles, settings.xml, its theme, docProps. None of that can be produced by hand, and it is
what these fixtures exist to test. They go through the same per-line comparison as everything
else.
Seven of them: smartart and smartart-lines, whose cached drawing is only worth comparing when
Word wrote it; report and memo; newsletter, a running head and a footer that counts the pages
over a body in two columns, with an address that lives only in a relationship; notes, whose
footnotes and endnotes are in parts Word wrote, separators and all; minutes, whose numbering part
Word rewrote wholesale; and brochure, a picture and a text box, where Word writes the drawing
twice over — the modern markup and a VML fallback beside it.
brochure is what the real documents are for. Every fixture written by hand sets its line spacing
to a single line, and Word's own Normal asks for 1.08 — so nothing here had ever put a picture on a
line that asked for a multiple. Word applies the multiple to the line the text would have made
and leaves the picture out of it, where we multiplied the whole line box and left a 96pt picture
6.8 points too low. image-line-probe measures it sixteen ways and ImageLineTests holds it.
tools/make-real-fixtures.sh --list shows what would be generated. Add to RealSeeds to cover
more; third-party templates are best avoided, since their licence terms would come with them.
Using n8PDF
What it does
How it works
Contributing