Releases: Artui/django-data-shape
Release list
v0.21.0
Added
-
values=expressions can name one another, as{values.<name>}. A
projected table's measure columns are usually related to each other -- a
requested amount and an approved one, a quantity and a total, an amount and
the rate derived from it -- and an expression could not say so. The
relationship had to be restated from whatever both columns were computed from,
with coefficients chosen so that it happens to hold: a reader could only find
the invariant by doing the arithmetic, and nothing rechecked it when either
expression was edited.The name is dotted rather than a bare
{requested_amount}because{per}and
{source}already occupy that space and a model is entitled to a column
calledper. Onlyvalues=entries are referenceable -- a copied column is
already reachable as{source}.name, and the primary key is the
row_number()window itself and is reachable nowhere.It is substitution, not sharing, which is what the spelling says: the
referenced expression is written out again, parenthesised, and the database
evaluates it once per reference. That is free for a deterministic expression,
and every expression here already had to be deterministic --template_database
reuses a database keyed on the declaration and nothing else. What a reference
adds is that a volatile expression now disagrees with itself within one
build, so the documentation pairs the feature with anInvariantrather than
leaving the rule stated and unchecked.A subquery computing each expression once was declined: the key is a
row_number()window whose outerORDER BYdecides where rows physically
land, so nesting moves the one thing this package exists to control -- and a
shape using no reference would have had its statement changed to buy a feature
it does not use. Substitution leaves every existing shape's SQL byte-identical.A reference that names nothing, and a cycle, are refused at declaration time
and name the path. A name that is a copied column is refused separately and
hands over{source}.name, because that is a mistake about the spelling
rather than about the column.
Fixed
-
The documented
SqlValueexample was a type error on UUID-keyed models.
({per}.id * 31 + {source}.id * 17) % 5 + 1works on Django's default
BigAutoFieldand has no operator at all on a schema whose models carry
id = UUIDField(primary_key=True)-- which a shared abstract base makes an
ordinary layout rather than an unusual one. So the first thing such a reader
copied out of the documentation did not run, and it failed from inside a
generated statement at build time rather than at declaration.The guide and the docstring now carry a worked example for UUID keys, and it
is executed by the suite rather than asserted: the documentation's Python
blocks are only parsed, and an expression is a string literal that parses
perfectly, which is exactly how this shipped. Two parts of the replacement are
load-bearing and fail rarely enough to reach production --::bigintbefore
abs, becausehashtextreturnsint4andabs(-2147483648)isinteger out of range; andabsat all, because PostgreSQL's%keeps the sign of
the dividend, so a measure column would hold negatives and every plan over it
would still look fine. -
The projections guide implied UUID keys were settled by one sentence about
sql=. That sentence is about the projected table's own primary key. The
tables named byper=andcopying=may be keyed however they like, and the
care an expression over one of them needs is a different subject that the page
did not cover -- so a reader with UUID keys throughout read a paragraph that
appeared to address them and was sent onward into the broken example.
v0.20.0
Added
-
A lone
%in asql=statement is refused at declaration time. The
statement is run with its parameters, and an empty parameter sequence is still
a sequence, so a%that is not a placeholder is read as the start of one:
the failure came frompsycopg/cursor.pyat build time, naming the driver and
nothing about the shape. A caller who passed no parameters at all had no model
that explained it, and the modulo operator is the reason anyone writes a bare
%in the first place.A valid statement can never contain one, so nothing legal is refused -- which
is the whole reason this can be a refusal rather than an escape.sql=takes
params=, so pyformat is its interface and%%stays the spelling for the
operator.SqlValueon the derived path escapes instead, and the two are not
inconsistent: there the caller supplies no parameters and has no reason to
know one exists. Both docstrings and the projections guide now say which is
which, at both call sites, because the asymmetry is exactly what a reader
finding one of them will next be confused by.
v0.19.0
Added
-
Paired(..., parents=): an edge narrowed to a subset of the partner table.Two through tables over one partner model chose independently -- every
mechanism here computes a column from the row index and nothing else -- so
they overlapped by construction, and any rule of the form these two
relationships must not overlap was unsatisfiable however the shape was
written. Declared as an invariant, such a rule correctly fired and rolled
every build back with nothing a declaration could change.The family is larger than it looks: reviewer-is-not-author,
approver-is-not-requester, auditor-is-not-audited. It is every
separation-of-duties constraint there is.It behaves as
FanOut(parents=)does, which gained the same narrowing in
0.13.0 for the same class of problem one primitive over: keys are read through
the database as a predicate rather than filtered afterwards, named partners
are put back in declaration order so the weights do not follow a sort order
nobody wrote down, and a named key matching no row is refused rather than
silently dropped.Narrowing can make a declaration impossible that was fine over the whole
table, because what binds is the busiest group rather than the product. The
existing build-time refusal already says so and now says it about the narrowed
set. -
Projection(..., values={...})andSqlValue: a projected column of the
table's own.A projection copies a column from the source it names or takes the model's
own default, and a projected table's measure column is neither -- the
score on a review, the amount on a generated line, the reading on a sample.Leaving it to a model default was legal and was the wrong answer for this
package specifically: one value across every projected row is
n_distinct = 1, the exact shape a planner cannot use. A library whose whole
purpose is planner realism was building tables it had made unplannable, and
the declaration looked correct.sql=already answered this and answered it expensively -- it replaces the
wholeSELECT, so the join stops being derived from the model graph and can
drift from it afterwards, the copied columns are written by hand, and the key
strategy has to be spelled in SQL.values=gives up none of that:Projection( ReviewScore, per=Review, copying=Criterion, values={"score": SqlValue("({per}.id * 31 + {source}.id * 17) % 5 + 1")}, )
{per}and{source}are substituted with the aliases the derived statement
uses. They are placeholders rather than the aliases themselves because the
aliases are this package's private business.It takes SQL rather than a distribution, and that is a decision. A
Distributioncomputes fromdraw(stream, row), which is SplitMix64 --
expressible in PostgreSQL only throughnumericmodular arithmetic and casts
across the sign boundary, where one mistake gives a declaration two meanings
depending on which statement filled the table. That is the divergence
SqlKeysexists to refuse, and it is not worth buying convenience with.Declaring an expression for a column the source already carries, or for a
column the model does not have, or alongsidesql=, is refused at declaration
time.Two things the expression carries that the declaration should not have to. A
literal%is escaped, because the statement is executed with bound
parameters and an unescaped one is an incomplete placeholder to psycopg and to
Django's SQLite wrapper alike -- the paramstyle is no more the declaration's
business than the join's aliases are. And the expression is the one part of a
shape that is not portable, which is documented rather than papered over:
mod(x, 5)is an integer on PostgreSQL and a REAL on SQLite, so%is the
spelling the examples use.
v0.18.1
Fixed
-
scaled_shapenow multiplies a projection'smax_rowsby the factor.
It passed aProjectionthrough untouched, which is right for its size --
that grows because the tables it reads did -- and wrong for a declared
ceiling, which is a number in the same units and did not move. A ceiling sized
for the world as written therefore refused every growth assertion, because
those build the same declaration at a larger factor.The consumer that asked for
max_rowsin 0.18.0 hit this on the first run:
10,000 declared against a world of 3,188 rows, and 29,078 at factor 10.Multiplying is the arithmetic rather than an approximation of one, and the
reason is the same one that makes the size need no factor: every table scales,
parents included, so a parent has the same number of children at every factor
and the projection is a sum overfactortimes as many parents of an
unchanged per-parent product.A declaration with no ceiling is still passed through by identity.
v0.18.0
Added
-
Projection(..., max_rows=N): a declared ceiling, checked before the insert.A projection is the one declaration with no
rows=, deliberately -- its
cardinality comes from the join, which is what reproduces a correlation a
FanOuton the child would destroy. The consequence is that the largest table
in a database can be the one nobody declared a size for, and it grows as a
product: when both sides of the join fan out over the same parents, the
busy parents multiply, so raising either declared count by four grows the
result by sixteen. A consumer measured 2,413,223 rows against a declaration
whose largest number was 300,000.The count is taken first and compared, so a declaration that has run away
costs a scan of the join rather than the time to write every row of it. It is
exact rather than estimated: the derived form counts the same join the insert
selects from, and asql=projection is counted by wrapping the caller's own
select, because this package cannot know what that statement is one row per.The refusal names the number it would have written and the tables the join is
over, because the surprise is never the ceiling -- a reader told only that a
limit was exceeded still has to work out which of the two counts moved.A declaration that does not ask is not charged for the answer: with no
ceiling, no count is taken. There is no default ceiling and there will not be
one -- how many rows is too many is a judgement about size, which this package
does not make on a caller's behalf anywhere else either.
v0.17.1
Fixed
-
A scaled world left the identity sequence pointing at rows that came back,
a regression introduced by 0.17.0 and found by the consumer that prompted it.0.17.0 made
scaled_worldempty the tables its shape declares, so a session
world could sit under one, and the rows come back because the emptying happens
inside the transaction it rolls back. The sequence does not:setvalis
not transactional, so the counter kept whatever the scaled build moved it to.
A scaled world is usually smaller than the session world it was built over,
which left the counter below the ids that had just returned.The symptom was an
IntegrityErroron a primary key, in a later test, for a
row the failing test never wrote.The sequences are now recomputed once the transaction has ended -- from
max(pk)of whatever actually survived, rather than from a number captured on
the way in, so it is correct in both directions and correct too when the
caller's block raised partway through the build.The Postgres statement-count constant moved from 17 to 19: one reset per
declared table, a property of the declaration rather than the factor, so what
those tests pin is unchanged. The portable constant did not move, because
Django emits no sequence reset for a SQLite table withoutAUTOINCREMENT.
v0.17.0
Added
-
Product,OffsetandCopied: the commonest arithmetic, said as data.
They compute nothingDerivedcould not. They exist because of whatDerived
costs: it takes a callable, a callable cannot be honestly digested, and a
shape holding one is refused bytemplate_database. So a column as ordinary
astotal = quantity * unit_pricetook a whole declaration out of the reuse
that turns a forty-second build into a hundred-millisecond clone -- and
build()kept working, so nothing said what it had cost.The refusal stays; it is right. These three are pure data, implement
Canonical, and hash.Derivedremains the answer for computation that
really is code.Offsetis also the same-row half ofAfter, which is parent-scoped only.
Its gap is fixed rather than spread across a window, because a due date thirty
days after an issue date is a term and not a distribution.
Changed
-
A scaled world can now be built over a session world. The two pytest
surfaces this package ships wanted the same tables in any application with one
model graph, and were mutually exclusive: the session world holds its rows for
the whole run, so the scaled build met a table that was not empty and refused.
The documented answer -- give the two different models -- is not available to a
project whose plan assertions and growth assertions are about the same flow,
because that is what the application is.It was also order-dependent, which is what made it worth fixing rather than
documenting better: a suite whose growth tests happened to be collected first
passed, and the same tests named in the other order failed.scaled_worldnow empties the tables its shape declares before building.
Nothing is snapshotted, and nothing needs to be: this happens inside the
atomic block it already rolls back, so whatever was there comes back when the
block ends. A table whose keys areDisjointis left alone, mirroring the
exemptionbuildalready makes, so the documented hybrid -- parents made by
your code, children made here -- keeps working.build()keeps its refusal. It has no transaction of its own to undo and
would be destroying rows for good.The two statement-count constants moved: one TRUNCATE on PostgreSQL, one
DELETE per declared table elsewhere. Both are properties of the declaration
rather than the factor, so what those tests pin -- fixed on PostgreSQL at
every factor, a curve off it -- is unchanged.
Fixed
-
Afteracross a parent column of a different kind is refused.
date + timedeltais adate, so aDateTimeFieldchild whose parent column
was aDateFieldwas filled with dates:COPYaccepted them and they landed
as naive midnights. A world whose orders went on sale in 2024 had shows
starting in 1900, and the only signal was Django's own per-row
RuntimeWarningin the middle of a build that prints thousands of lines.Both field types are on
_metaand both are known before a row is generated,
so this belonged in the class of things already refused at declaration time.
OnlyAfteris checked, deliberately:Afterisparent + offsetwritten
into this column, so the two describe the same quantity by construction, while
aDerivedreading a parent column may legitimately convert -- turning a
timestamp into a count of days is an ordinary thing to declare.
v0.16.0
Fixed
- A
FanOut(parents=[...])partition no longer depends on what the keys are.
Keys are read back ordered by primary key and the sizes are assigned by
position, so the weights followed the sort order of the values rather than
the declaration. With integer keys that is merely surprising; with the UUID
primary keys a factory row has on a modern schema it means the same shape
builds differently every run -- measured, one declaration gave the
first-named parent 5, 11 or 79 rows across twelve builds. That is the promise
the package rests on, and the same reasonUuidKeysderives rather than
draws, so a narrowing that broke it was a defect and not a preference. Named
parents are put back in the order they were named before anything is weighed.- Reversing the list is now a different declaration, which is what a
reader writing one would expect and was the symptom reported. - Weights are still scattered across positions, so the reasoning that
argued for scattering survives: a caller's order decides which key lands
where, and nothing correlates a parent's key with its child count. - Worth knowing for anyone testing near this: a test written over
integer parents passes against the broken behaviour, because the
database's sort order matches insertion order there. Only UUID parents show
it.
- Reversing the list is now a different declaration, which is what a
shape_from_factorytakesdefaults=, and names a factory that cannot be
called. Most factories in a real codebase take arguments and need them; a
TeamFactorywanting apermission_role, called with nothing, left a
required column empty and failed as a rawIntegrityErrornaming a
constraint. That said nothing about what the caller ran, which is the one
thing every other refusal here avoids. Anything the factory raises is now
named with the call number, because "it failed" and "it failed after three"
are different bugs.
v0.15.0
Added
shape_from_factoryruns a factory you already have and returns source, not
aShape. That is the design rather than a limitation of it: a shape this
package builds is declared, which is what makes it reviewable and
assertable, and a shape learned from a sample is neither -- one used directly
would change whenever the factory did, silently. So it returns text, and a
person decides what of it to keep.- What it finds matters more than what it writes, because a faithful
reading of a typical factory produces the flat world this package exists to
argue against. Factories are written for single-object tests, so they fix
values and reach foreign keys in the two most unrealistic ways there are.
Every such column is a finding, and the report leads with them. - A sub-factory is the sharpest case and the reason to run it.
company = SubFactory(CompanyFactory)creates one parent per child, which is
a fan-out of degree one: every parent has exactly one row, the average is the
truth, and a join over it cannot be misestimated. It is invisible in the
factory's own source, so it is detected by watching which other tables grew
and by how much. A round-robin over four parents is the same defect with a
different number, and is reported too. - A relation is rendered as a
FanOutand never as a value distribution,
becauseTablerefuses the second -- emitting it would hand back source that
cannot be built, which is worse than a wrong number. A factory that already
varies everything gets a declaration and no findings: if the quiet case
were not quiet, the loud one would stop meaning anything. - Nothing is left behind. The calls run inside a transaction that is rolled
back, so it can be pointed at a development database without writing to one.
The sample size is stated in the output, because the answer moves with it.
- What it finds matters more than what it writes, because a faithful
Fixed
Md5Keyscould not fill a projection on about half of all tables.
key_sqlwroteto_hex(<stream>::bigint), asking PostgreSQL to re-derive a
number Python produced as unsigned -- andbigintis signed, so any
stream above 2^63 raisedNumericValueOutOfRangebefore a row was written.
The stream is a hash of the table and field names, so it is a coin flip per
table rather than anything about a schema or a seed: measured, 49% of table
names land above the limit. The stream is a constant by the time the statement
is built, so the sixteen hex digits are embedded and nothing is re-derived.- Reported by a consumer, and it had never executed for them. In 0.13.0
the join-ambiguity refusal answered first, sokey_sqlwas unreachable;
fixing that in 0.14.0 withthrough=is what exposed a feature 0.13.0 had
shipped and nobody could run. - The test that was supposed to prove the two halves agree used
stream=12345-- a number chosen for a test rather than one the producer
makes. It now uses a pairfield_streamactually produced, one either side
of the limit, and asserts that the pair straddles it so a later change to
how streams are derived cannot quietly make both tests vacuous.
- Reported by a consumer, and it had never executed for them. In 0.13.0
v0.14.0
Added
-
Pairedfills the second key of a many-to-many, and the edge count stays
exact. Two fan-outs on one through table are refused because a fan-out is a
partition computed from the row index alone, so two of them partition the same
rows without either seeing the other -- nothing enumerates the pairs and a
collision is a matter of the seed.person=Paired("company", Zipf())is the
half that looks:companypartitions the rows as any fan-out does, and within
each of its groupspersontakes that many distinct partners. A duplicate
pair is impossible by construction rather than removed afterwards, so nothing
is deduplicated and the build never reports an achieved count against a
requested one. The design notes for this milestone had given that up; they
were wrong, and the rule that cardinality is declared survives M2M.- What binds is the busiest group, not the product. Every row of one
group needs a different partner, so the constraint is
largest group <= partnersrather thanrows <= groups x partners-- and a
heavy tail puts a large share of every edge on one group,Zipf(1.2)over
five thousand of them putting 21% on the top one. A declaration that looks
sparse against the product can still be impossible, and it cannot be known
until the partition is resolved: this is the one structural refusal that
waits for the build rather than firing where the shape is written. - The second side is derived and approximate, and the numbers are
published. Both marginals plus the edge count is over-determined, and
fixing all three is a constraint satisfaction problem that cannot stream
intoCOPY. Measured against exact weighted sampling over 200,000 edges,
the busiest partner comes out 1.09 times as busy, the 99th percentile about
30% high, and about 7% fewer partners are touched. - Partners are allocated across weight bands and strided within one, so
nothing is ever drawn and asked whether it is taken. Rejection sampling
degenerates exactly where the shape is most interesting -- 245 probes per
row on that world -- and the usual fix, sampling the ones to leave out when
a group wants more than half, is ten times faster and a different sampling
rule, so a group one row over the halfway mark would get a materially
different membership from one row under. The band count is derived from
the declared weights rather than chosen, which systematic allocation is
what makes possible: it converges as bands get finer where largest-remainder
runs away, so "enough" is a limit rather than a number somebody tuned.
- What binds is the busiest group, not the product. Every row of one
-
FanOutPlan.group_ofanswers which parent a row belongs to, which a
pairing needs to choose partners inside this row's group. -
Projection(..., through=Model)says which model a derived join runs on.
An abstract base carryingcreated_by/updated_bytoUser-- a pattern most
Django schemas have -- makes any two models in the schema joinable, so the
derivation had candidates everywhere and could settle nothing: not for one
pair but for every pair, which put the derived form out of reach for a whole
application and leftsql=as the only route to this feature's own motivating
example. The refusal now namesthrough=first, and names a model that can
actually resolve the join rather than the first one alphabetically -- an audit
model is reached by two edges from each side, so it is a real candidate and
never a usable answer. Where every candidate is reached more than once,
through=cannot narrow anything and the refusal says so instead of
suggesting it.
Changed
- The two-fan-out refusal names the declaration that replaces it. Its remedy
used to be a statement of your own, because nothing else could keep the
constraint. It now namesPairedfirst and keepsProjection(columns=, sql=)
as the escape hatch for an edge set that has to be a particular one.
Fixed
-
A rounded
Uniformcan size a fan-out.
FanOut(Uniform(1, 10, places=0))was refused with "needs numeric fan-out
sizes, but producedDecimal('3')" -- which reads as a contradiction, and was
one keyword away fromFanOut(Uniform(1, 10)), the spelling this package's own
docstring recommends and two of its refusals suggest.places=0is the natural
way to say that a fan-out size is a count. The guard was
isinstance(weight, (int, float))while the next line already did
float(weight), so it was stricter than the arithmetic it protected. It is now
the numeric tower plusDecimal, whichnumbers.Realdeliberately excludes,
andFractionand numpy scalars come along with it.boolstays refused. The
message also names a remedy now, which it was alone in not doing. -
columns=accepts the column anINSERTactually lists. The refusal that
sends a reader there says "the select's columns are what the insert lists", and
what an insert lists for a foreign key isevent_id-- which was then refused
as "no field named event_id". That is a contradiction rather than a correction,
and the surrounding documentation reinforced the wrong reading. Both spellings
are accepted and resolve to the same column, the messages say so, and the same
dict is read forstatistics=so the two cannot disagree about what a column
is called.