Skip to content

Constant-fold f-strings with constant fields in the CFG optimizer #154910

Description

@overlorde

Feature or enhancement

Proposal:

The flowgraph optimizer folds constant expressions in most positions
(binary/unary ops, tuples of constants, in against literal
collections), but f-strings whose fields are all constants are left
unfolded:

>>> import dis
>>> dis.dis(lambda: f'{1}{2}')
   RESUME                   0
   LOAD_SMALL_INT           1
   FORMAT_SIMPLE
   LOAD_SMALL_INT           2
   FORMAT_SIMPLE
   BUILD_STRING             2
   RETURN_VALUE

This could compile to a single LOAD_CONST '12', the same way
'1' + '2' already does.

The transformation is safe under a narrow guard:

  • FORMAT_SIMPLE on an operand that is an exact str, int,
    float or bool constant is pure and calls no user code
    (subclasses can override __format__/__str__, so exact-type
    checks are required).
  • FORMAT_WITH_SPEC with a constant spec on those exact types is
    pure for the same reason, though it could be left to a follow-up.
  • BUILD_STRING over constant strings is pure.

Conversion functions (!r, !s, !a) lower to intrinsics before
FORMAT_SIMPLE, so f'{1!r}' is out of scope unless the intrinsic
is also handled; the initial change can simply not match those
sequences.

Constant fields do appear in real code — mixed f-strings like
f'{HEADER}{x}' gain partially (the constant fields fold, shrinking
BUILD_STRING's operand count), and fully-constant f-strings show up
in generated code and in string-building via comprehensions.

The natural place is a fold_format_simple alongside
fold_tuple_of_constants in Python/flowgraph.c, using the existing
get_const_loading_instrs / instr_make_load_const helpers.

gh-77273 covered the opcode-level inefficiency of f-string formatting
and was resolved by the FORMAT_SIMPLE/FORMAT_WITH_SPEC redesign;
this proposal is about folding the constant cases of those opcodes,
which is not currently done anywhere in the compiler.

With a prototype of the fold, a fully-constant f-string costs the same
as the equivalent string literal (~34 ns/call instead of ~1770 ns/call
on a debug free-threaded build — I'll post release-build pyperf numbers
on the PR). I have a patch ready and will follow up with it.

Has this already been discussed elsewhere?

This is a minor feature, which does not need previous discussion elsewhere

Links to previous discussion of this feature:

No response

Linked PRs

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)performancePerformance or resource usage

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions