Skip to content

Releases: Artui/django-data-shape

v0.21.0

Choose a tag to compare

@github-actions github-actions released this 06 Sep 11:44
e6e67a5

Added

  • values= expressions can name one another, as {values.<name>}. A
    projected table's measure columns are usually related to each other -- a
    requested amount and an approved one, a quantity and a total, an amount and
    the rate derived from it -- and an expression could not say so. The
    relationship had to be restated from whatever both columns were computed from,
    with coefficients chosen so that it happens to hold: a reader could only find
    the invariant by doing the arithmetic, and nothing rechecked it when either
    expression was edited.

    The name is dotted rather than a bare {requested_amount} because {per} and
    {source} already occupy that space and a model is entitled to a column
    called per. Only values= entries are referenceable -- a copied column is
    already reachable as {source}.name, and the primary key is the
    row_number() window itself and is reachable nowhere.

    It is substitution, not sharing, which is what the spelling says: the
    referenced expression is written out again, parenthesised, and the database
    evaluates it once per reference. That is free for a deterministic expression,
    and every expression here already had to be deterministic -- template_database
    reuses a database keyed on the declaration and nothing else. What a reference
    adds is that a volatile expression now disagrees with itself within one
    build
    , so the documentation pairs the feature with an Invariant rather than
    leaving the rule stated and unchecked.

    A subquery computing each expression once was declined: the key is a
    row_number() window whose outer ORDER BY decides where rows physically
    land, so nesting moves the one thing this package exists to control -- and a
    shape using no reference would have had its statement changed to buy a feature
    it does not use. Substitution leaves every existing shape's SQL byte-identical.

    A reference that names nothing, and a cycle, are refused at declaration time
    and name the path. A name that is a copied column is refused separately and
    hands over {source}.name, because that is a mistake about the spelling
    rather than about the column.

Fixed

  • The documented SqlValue example was a type error on UUID-keyed models.
    ({per}.id * 31 + {source}.id * 17) % 5 + 1 works on Django's default
    BigAutoField and has no operator at all on a schema whose models carry
    id = UUIDField(primary_key=True) -- which a shared abstract base makes an
    ordinary layout rather than an unusual one. So the first thing such a reader
    copied out of the documentation did not run, and it failed from inside a
    generated statement at build time rather than at declaration.

    The guide and the docstring now carry a worked example for UUID keys, and it
    is executed by the suite rather than asserted: the documentation's Python
    blocks are only parsed, and an expression is a string literal that parses
    perfectly, which is exactly how this shipped. Two parts of the replacement are
    load-bearing and fail rarely enough to reach production -- ::bigint before
    abs, because hashtext returns int4 and abs(-2147483648) is integer out of range; and abs at all, because PostgreSQL's % keeps the sign of
    the dividend, so a measure column would hold negatives and every plan over it
    would still look fine.

  • The projections guide implied UUID keys were settled by one sentence about
    sql=.
    That sentence is about the projected table's own primary key. The
    tables named by per= and copying= may be keyed however they like, and the
    care an expression over one of them needs is a different subject that the page
    did not cover -- so a reader with UUID keys throughout read a paragraph that
    appeared to address them and was sent onward into the broken example.

v0.20.0

Choose a tag to compare

@github-actions github-actions released this 05 Sep 14:47
3c4f470

Added

  • A lone % in a sql= statement is refused at declaration time. The
    statement is run with its parameters, and an empty parameter sequence is still
    a sequence, so a % that is not a placeholder is read as the start of one:
    the failure came from psycopg/cursor.py at build time, naming the driver and
    nothing about the shape. A caller who passed no parameters at all had no model
    that explained it, and the modulo operator is the reason anyone writes a bare
    % in the first place.

    A valid statement can never contain one, so nothing legal is refused -- which
    is the whole reason this can be a refusal rather than an escape. sql= takes
    params=, so pyformat is its interface and %% stays the spelling for the
    operator.

    SqlValue on the derived path escapes instead, and the two are not
    inconsistent: there the caller supplies no parameters and has no reason to
    know one exists. Both docstrings and the projections guide now say which is
    which, at both call sites, because the asymmetry is exactly what a reader
    finding one of them will next be confused by.

v0.19.0

Choose a tag to compare

@github-actions github-actions released this 05 Sep 13:29
4a3af2c

Added

  • Paired(..., parents=): an edge narrowed to a subset of the partner table.

    Two through tables over one partner model chose independently -- every
    mechanism here computes a column from the row index and nothing else -- so
    they overlapped by construction, and any rule of the form these two
    relationships must not overlap
    was unsatisfiable however the shape was
    written. Declared as an invariant, such a rule correctly fired and rolled
    every build back with nothing a declaration could change.

    The family is larger than it looks: reviewer-is-not-author,
    approver-is-not-requester, auditor-is-not-audited. It is every
    separation-of-duties constraint there is.

    It behaves as FanOut(parents=) does, which gained the same narrowing in
    0.13.0 for the same class of problem one primitive over: keys are read through
    the database as a predicate rather than filtered afterwards, named partners
    are put back in declaration order so the weights do not follow a sort order
    nobody wrote down, and a named key matching no row is refused rather than
    silently dropped.

    Narrowing can make a declaration impossible that was fine over the whole
    table
    , because what binds is the busiest group rather than the product. The
    existing build-time refusal already says so and now says it about the narrowed
    set.

  • Projection(..., values={...}) and SqlValue: a projected column of the
    table's own.

    A projection copies a column from the source it names or takes the model's
    own default, and a projected table's measure column is neither -- the
    score on a review, the amount on a generated line, the reading on a sample.

    Leaving it to a model default was legal and was the wrong answer for this
    package specifically
    : one value across every projected row is
    n_distinct = 1, the exact shape a planner cannot use. A library whose whole
    purpose is planner realism was building tables it had made unplannable, and
    the declaration looked correct.

    sql= already answered this and answered it expensively -- it replaces the
    whole SELECT, so the join stops being derived from the model graph and can
    drift from it afterwards, the copied columns are written by hand, and the key
    strategy has to be spelled in SQL. values= gives up none of that:

    Projection(
        ReviewScore,
        per=Review,
        copying=Criterion,
        values={"score": SqlValue("({per}.id * 31 + {source}.id * 17) % 5 + 1")},
    )

    {per} and {source} are substituted with the aliases the derived statement
    uses. They are placeholders rather than the aliases themselves because the
    aliases are this package's private business.

    It takes SQL rather than a distribution, and that is a decision. A
    Distribution computes from draw(stream, row), which is SplitMix64 --
    expressible in PostgreSQL only through numeric modular arithmetic and casts
    across the sign boundary, where one mistake gives a declaration two meanings
    depending on which statement filled the table. That is the divergence
    SqlKeys exists to refuse, and it is not worth buying convenience with.

    Declaring an expression for a column the source already carries, or for a
    column the model does not have, or alongside sql=, is refused at declaration
    time.

    Two things the expression carries that the declaration should not have to. A
    literal % is escaped, because the statement is executed with bound
    parameters and an unescaped one is an incomplete placeholder to psycopg and to
    Django's SQLite wrapper alike -- the paramstyle is no more the declaration's
    business than the join's aliases are. And the expression is the one part of a
    shape that is not portable, which is documented rather than papered over:
    mod(x, 5) is an integer on PostgreSQL and a REAL on SQLite, so % is the
    spelling the examples use.

v0.18.1

Choose a tag to compare

@github-actions github-actions released this 05 Sep 10:05
dc5a1a4

Fixed

  • scaled_shape now multiplies a projection's max_rows by the factor.
    It passed a Projection through untouched, which is right for its size --
    that grows because the tables it reads did -- and wrong for a declared
    ceiling, which is a number in the same units and did not move. A ceiling sized
    for the world as written therefore refused every growth assertion, because
    those build the same declaration at a larger factor.

    The consumer that asked for max_rows in 0.18.0 hit this on the first run:
    10,000 declared against a world of 3,188 rows, and 29,078 at factor 10.

    Multiplying is the arithmetic rather than an approximation of one, and the
    reason is the same one that makes the size need no factor: every table scales,
    parents included, so a parent has the same number of children at every factor
    and the projection is a sum over factor times as many parents of an
    unchanged per-parent product.

    A declaration with no ceiling is still passed through by identity.

v0.18.0

Choose a tag to compare

@github-actions github-actions released this 05 Sep 09:39
92041d6

Added

  • Projection(..., max_rows=N): a declared ceiling, checked before the insert.

    A projection is the one declaration with no rows=, deliberately -- its
    cardinality comes from the join, which is what reproduces a correlation a
    FanOut on the child would destroy. The consequence is that the largest table
    in a database can be the one nobody declared a size for, and it grows as a
    product: when both sides of the join fan out over the same parents, the
    busy parents multiply, so raising either declared count by four grows the
    result by sixteen. A consumer measured 2,413,223 rows against a declaration
    whose largest number was 300,000.

    The count is taken first and compared, so a declaration that has run away
    costs a scan of the join rather than the time to write every row of it. It is
    exact rather than estimated: the derived form counts the same join the insert
    selects from, and a sql= projection is counted by wrapping the caller's own
    select, because this package cannot know what that statement is one row per.

    The refusal names the number it would have written and the tables the join is
    over, because the surprise is never the ceiling -- a reader told only that a
    limit was exceeded still has to work out which of the two counts moved.

    A declaration that does not ask is not charged for the answer: with no
    ceiling, no count is taken. There is no default ceiling and there will not be
    one -- how many rows is too many is a judgement about size, which this package
    does not make on a caller's behalf anywhere else either.

v0.17.1

Choose a tag to compare

@github-actions github-actions released this 05 Sep 09:18
d204df2

Fixed

  • A scaled world left the identity sequence pointing at rows that came back,
    a regression introduced by 0.17.0 and found by the consumer that prompted it.

    0.17.0 made scaled_world empty the tables its shape declares, so a session
    world could sit under one, and the rows come back because the emptying happens
    inside the transaction it rolls back. The sequence does not: setval is
    not transactional, so the counter kept whatever the scaled build moved it to.
    A scaled world is usually smaller than the session world it was built over,
    which left the counter below the ids that had just returned.

    The symptom was an IntegrityError on a primary key, in a later test, for a
    row the failing test never wrote.

    The sequences are now recomputed once the transaction has ended -- from
    max(pk) of whatever actually survived, rather than from a number captured on
    the way in, so it is correct in both directions and correct too when the
    caller's block raised partway through the build.

    The Postgres statement-count constant moved from 17 to 19: one reset per
    declared table, a property of the declaration rather than the factor, so what
    those tests pin is unchanged. The portable constant did not move, because
    Django emits no sequence reset for a SQLite table without AUTOINCREMENT.

v0.17.0

Choose a tag to compare

@github-actions github-actions released this 05 Sep 09:03
3319458

Added

  • Product, Offset and Copied: the commonest arithmetic, said as data.
    They compute nothing Derived could not. They exist because of what Derived
    costs: it takes a callable, a callable cannot be honestly digested, and a
    shape holding one is refused by template_database. So a column as ordinary
    as total = quantity * unit_price took a whole declaration out of the reuse
    that turns a forty-second build into a hundred-millisecond clone -- and
    build() kept working, so nothing said what it had cost.

    The refusal stays; it is right. These three are pure data, implement
    Canonical, and hash. Derived remains the answer for computation that
    really is code.

    Offset is also the same-row half of After, which is parent-scoped only.
    Its gap is fixed rather than spread across a window, because a due date thirty
    days after an issue date is a term and not a distribution.

Changed

  • A scaled world can now be built over a session world. The two pytest
    surfaces this package ships wanted the same tables in any application with one
    model graph, and were mutually exclusive: the session world holds its rows for
    the whole run, so the scaled build met a table that was not empty and refused.
    The documented answer -- give the two different models -- is not available to a
    project whose plan assertions and growth assertions are about the same flow,
    because that is what the application is.

    It was also order-dependent, which is what made it worth fixing rather than
    documenting better: a suite whose growth tests happened to be collected first
    passed, and the same tests named in the other order failed.

    scaled_world now empties the tables its shape declares before building.
    Nothing is snapshotted, and nothing needs to be: this happens inside the
    atomic block it already rolls back, so whatever was there comes back when the
    block ends. A table whose keys are Disjoint is left alone, mirroring the
    exemption build already makes, so the documented hybrid -- parents made by
    your code, children made here -- keeps working.

    build() keeps its refusal. It has no transaction of its own to undo and
    would be destroying rows for good.

    The two statement-count constants moved: one TRUNCATE on PostgreSQL, one
    DELETE per declared table elsewhere. Both are properties of the declaration
    rather than the factor, so what those tests pin -- fixed on PostgreSQL at
    every factor, a curve off it -- is unchanged.

Fixed

  • After across a parent column of a different kind is refused.
    date + timedelta is a date, so a DateTimeField child whose parent column
    was a DateField was filled with dates: COPY accepted them and they landed
    as naive midnights. A world whose orders went on sale in 2024 had shows
    starting in 1900, and the only signal was Django's own per-row
    RuntimeWarning in the middle of a build that prints thousands of lines.

    Both field types are on _meta and both are known before a row is generated,
    so this belonged in the class of things already refused at declaration time.
    Only After is checked, deliberately: After is parent + offset written
    into this column, so the two describe the same quantity by construction, while
    a Derived reading a parent column may legitimately convert -- turning a
    timestamp into a count of days is an ordinary thing to declare.

v0.16.0

Choose a tag to compare

@github-actions github-actions released this 04 Sep 10:14
3233392

Fixed

  • A FanOut(parents=[...]) partition no longer depends on what the keys are.
    Keys are read back ordered by primary key and the sizes are assigned by
    position, so the weights followed the sort order of the values rather than
    the declaration. With integer keys that is merely surprising; with the UUID
    primary keys a factory row has on a modern schema it means the same shape
    builds differently every run
    -- measured, one declaration gave the
    first-named parent 5, 11 or 79 rows across twelve builds. That is the promise
    the package rests on, and the same reason UuidKeys derives rather than
    draws, so a narrowing that broke it was a defect and not a preference. Named
    parents are put back in the order they were named before anything is weighed.
    • Reversing the list is now a different declaration, which is what a
      reader writing one would expect and was the symptom reported.
    • Weights are still scattered across positions, so the reasoning that
      argued for scattering survives: a caller's order decides which key lands
      where, and nothing correlates a parent's key with its child count.
    • Worth knowing for anyone testing near this: a test written over
      integer parents passes against the broken behaviour, because the
      database's sort order matches insertion order there. Only UUID parents show
      it.
  • shape_from_factory takes defaults=, and names a factory that cannot be
    called.
    Most factories in a real codebase take arguments and need them; a
    TeamFactory wanting a permission_role, called with nothing, left a
    required column empty and failed as a raw IntegrityError naming a
    constraint. That said nothing about what the caller ran, which is the one
    thing every other refusal here avoids. Anything the factory raises is now
    named with the call number, because "it failed" and "it failed after three"
    are different bugs.

v0.15.0

Choose a tag to compare

@github-actions github-actions released this 04 Sep 08:34
9fdafaa

Added

  • shape_from_factory runs a factory you already have and returns source, not
    a Shape.
    That is the design rather than a limitation of it: a shape this
    package builds is declared, which is what makes it reviewable and
    assertable, and a shape learned from a sample is neither -- one used directly
    would change whenever the factory did, silently. So it returns text, and a
    person decides what of it to keep.
    • What it finds matters more than what it writes, because a faithful
      reading of a typical factory produces the flat world this package exists to
      argue against. Factories are written for single-object tests, so they fix
      values and reach foreign keys in the two most unrealistic ways there are.
      Every such column is a finding, and the report leads with them.
    • A sub-factory is the sharpest case and the reason to run it.
      company = SubFactory(CompanyFactory) creates one parent per child, which is
      a fan-out of degree one: every parent has exactly one row, the average is the
      truth, and a join over it cannot be misestimated. It is invisible in the
      factory's own source, so it is detected by watching which other tables grew
      and by how much. A round-robin over four parents is the same defect with a
      different number, and is reported too.
    • A relation is rendered as a FanOut and never as a value distribution,
      because Table refuses the second -- emitting it would hand back source that
      cannot be built, which is worse than a wrong number. A factory that already
      varies everything gets a declaration and no findings: if the quiet case
      were not quiet, the loud one would stop meaning anything.
    • Nothing is left behind. The calls run inside a transaction that is rolled
      back, so it can be pointed at a development database without writing to one.
      The sample size is stated in the output, because the answer moves with it.

Fixed

  • Md5Keys could not fill a projection on about half of all tables.
    key_sql wrote to_hex(<stream>::bigint), asking PostgreSQL to re-derive a
    number Python produced as unsigned -- and bigint is signed, so any
    stream above 2^63 raised NumericValueOutOfRange before a row was written.
    The stream is a hash of the table and field names, so it is a coin flip per
    table rather than anything about a schema or a seed: measured, 49% of table
    names land above the limit. The stream is a constant by the time the statement
    is built, so the sixteen hex digits are embedded and nothing is re-derived.
    • Reported by a consumer, and it had never executed for them. In 0.13.0
      the join-ambiguity refusal answered first, so key_sql was unreachable;
      fixing that in 0.14.0 with through= is what exposed a feature 0.13.0 had
      shipped and nobody could run.
    • The test that was supposed to prove the two halves agree used
      stream=12345
      -- a number chosen for a test rather than one the producer
      makes. It now uses a pair field_stream actually produced, one either side
      of the limit, and asserts that the pair straddles it so a later change to
      how streams are derived cannot quietly make both tests vacuous.

v0.14.0

Choose a tag to compare

@github-actions github-actions released this 03 Sep 19:53
8968cce

Added

  • Paired fills the second key of a many-to-many, and the edge count stays
    exact.
    Two fan-outs on one through table are refused because a fan-out is a
    partition computed from the row index alone, so two of them partition the same
    rows without either seeing the other -- nothing enumerates the pairs and a
    collision is a matter of the seed. person=Paired("company", Zipf()) is the
    half that looks: company partitions the rows as any fan-out does, and within
    each of its groups person takes that many distinct partners. A duplicate
    pair is impossible by construction rather than removed afterwards, so nothing
    is deduplicated and the build never reports an achieved count against a
    requested one. The design notes for this milestone had given that up; they
    were wrong, and the rule that cardinality is declared survives M2M.

    • What binds is the busiest group, not the product. Every row of one
      group needs a different partner, so the constraint is
      largest group <= partners rather than rows <= groups x partners -- and a
      heavy tail puts a large share of every edge on one group, Zipf(1.2) over
      five thousand of them putting 21% on the top one. A declaration that looks
      sparse against the product can still be impossible, and it cannot be known
      until the partition is resolved: this is the one structural refusal that
      waits for the build rather than firing where the shape is written.
    • The second side is derived and approximate, and the numbers are
      published.
      Both marginals plus the edge count is over-determined, and
      fixing all three is a constraint satisfaction problem that cannot stream
      into COPY. Measured against exact weighted sampling over 200,000 edges,
      the busiest partner comes out 1.09 times as busy, the 99th percentile about
      30% high, and about 7% fewer partners are touched.
    • Partners are allocated across weight bands and strided within one, so
      nothing is ever drawn and asked whether it is taken. Rejection sampling
      degenerates exactly where the shape is most interesting -- 245 probes per
      row on that world -- and the usual fix, sampling the ones to leave out when
      a group wants more than half, is ten times faster and a different sampling
      rule
      , so a group one row over the halfway mark would get a materially
      different membership from one row under. The band count is derived from
      the declared weights
      rather than chosen, which systematic allocation is
      what makes possible: it converges as bands get finer where largest-remainder
      runs away, so "enough" is a limit rather than a number somebody tuned.
  • FanOutPlan.group_of answers which parent a row belongs to, which a
    pairing needs to choose partners inside this row's group.

  • Projection(..., through=Model) says which model a derived join runs on.
    An abstract base carrying created_by/updated_by to User -- a pattern most
    Django schemas have -- makes any two models in the schema joinable, so the
    derivation had candidates everywhere and could settle nothing: not for one
    pair but for every pair, which put the derived form out of reach for a whole
    application and left sql= as the only route to this feature's own motivating
    example. The refusal now names through= first, and names a model that can
    actually resolve the join rather than the first one alphabetically -- an audit
    model is reached by two edges from each side, so it is a real candidate and
    never a usable answer. Where every candidate is reached more than once,
    through= cannot narrow anything and the refusal says so instead of
    suggesting it.

Changed

  • The two-fan-out refusal names the declaration that replaces it. Its remedy
    used to be a statement of your own, because nothing else could keep the
    constraint. It now names Paired first and keeps Projection(columns=, sql=)
    as the escape hatch for an edge set that has to be a particular one.

Fixed

  • A rounded Uniform can size a fan-out.
    FanOut(Uniform(1, 10, places=0)) was refused with "needs numeric fan-out
    sizes, but produced Decimal('3')" -- which reads as a contradiction, and was
    one keyword away from FanOut(Uniform(1, 10)), the spelling this package's own
    docstring recommends and two of its refusals suggest. places=0 is the natural
    way to say that a fan-out size is a count. The guard was
    isinstance(weight, (int, float)) while the next line already did
    float(weight), so it was stricter than the arithmetic it protected. It is now
    the numeric tower plus Decimal, which numbers.Real deliberately excludes,
    and Fraction and numpy scalars come along with it. bool stays refused. The
    message also names a remedy now, which it was alone in not doing.

  • columns= accepts the column an INSERT actually lists. The refusal that
    sends a reader there says "the select's columns are what the insert lists", and
    what an insert lists for a foreign key is event_id -- which was then refused
    as "no field named event_id". That is a contradiction rather than a correction,
    and the surrounding documentation reinforced the wrong reading. Both spellings
    are accepted and resolve to the same column, the messages say so, and the same
    dict is read for statistics= so the two cannot disagree about what a column
    is called.