Releases: IACBI/promptbase-profile-exporter
Release list
v0.15.0
Changed
pb-historyreports no longer slow down as the history grows. Every report used to
add up every listing of every snapshot for the trend chart; each snapshot now
stores its totals when it is taken. With 30,000 listings and 40 snapshots the
report took 2.2 s and now takes 0.46 s, the same at any history length (the totals
alone: 1.6 s, now 1 ms).- History format 2. A file written by 0.14.3 or earlier is still read as it is,
and its reports are unchanged. The firstpb-history snapshotinto it upgrades it
in one transaction (0.4 s for 40 snapshots of 30,000 listings), so two scheduled
runs at once are safe. Afterwards 0.14.3 and earlier refuse the file as a newer
format; keep one version of the tool per history file.
v0.14.3
Added
- A project site on GitHub Pages, https://iacbi.github.io/promptbase-profile-exporter/,
with a searchable catalog and a trends page rebuilt every day from a public profile
by the dashboard recipe in docs/github-action.md.
Changed
--layout filesleaves a file alone when it already holds exactly the new text,
and writes the others several at a time. On Windows, re-exporting 5,000 prompts
after one change took 8.5 s and now takes 0.9 s, and a first export went from
6.4 s to 4.4 s. Unchanged files keep their modification time.pb-convertrenders the catalog once and checks it in memory instead of writing
and re-reading a scratch copy. Converting 30,000 prompts to HTML took 5.4 s and
now takes 1.7 s (CSV 3.2 s to 1.9 s, NDJSON 2.2 s to 1.4 s), with identical output.- Faster writing and reading of large catalogs, with byte-identical files:
- The writers encode once and skip work for extra fields nobody asked for, so the
HTML writer is about 30% faster. - Reading back an HTML catalog (
--compare, validation) is twice as fast. - Reading a TXT catalog is about three times as fast.
- The writers encode once and skip work for extra fields nobody asked for, so the
Security
- Release files are built with a hash-pinned setuptools (
python -m build --no-isolation); the build used to fetch whatever setuptools PyPI served at
release time, and the result was then attested. - The Action no longer upgrades pip from PyPI before installing the exporter.
--layout filesnames a prompt whose slug starts with a Windows device name and a
dot (con.x)con_.x, notcon.x_: Windows before 11 opens the device for both.- Titles printed to the terminal and written into diff and
pb-historyreports drop
bidirectional control characters, which can make text display in another order
than it reads. - A diff report (
--diff-output) is written atomically and replaces a symbolic link
at its path instead of writing through it.
v0.14.2
Changed
- Comparing catalogs (
--compare,--update-file,pb-diff) is about 16x faster on
large catalogs: text that is already equal is no longer normalised. A 30,000-prompt
comparison took 7.7 s and now takes 0.5 s, with an identical report.
Fixed
pb-historywrites numbers in full: a gain of 2,000,100 views read+2.0001e+06
in the Markdown and HTML reports, and no digit of a small price is rounded away.- Two first
pb-history snapshotruns at the same time on a new file (two
scheduled jobs) no longer fail: the check and the creation of the file's tables
run under one write lock. Five of six concurrent runs used to stop with an error. - A
pb-historyHTML report keeps titles on one line and drops control characters,
as the Markdown report already did; a history file with an unreadable snapshot
time is an error, not a traceback. --sinceor--untilwith an offset that moves the time past year 1 or 9999
(0001-01-01T00:00:00+05:00) is now an invalid date: the command line stopped
with anOverflowErrortraceback and the web UI answered with an error 500.--comparetogether with--update-fileis refused.--update-filecompares with
the file it rewrites, so the other catalog was ignored without a word.- A response that is not UTF-8 (from a proxy or a captive portal, say) is retried
and then reported like any unreadable response, not a traceback; an infinite
or NaN number in PromptBase's data (a count, price, or rating) is reported as bad
data instead of a crash or an invalidInfinityin a JSON export. @(or@@) alone is an empty profile, not a query for a user with no name.- A
--configfile saved with a byte order mark is read (Windows Notepad adds one),
and an option on the command line now replaces one from the file that it excludes
(--paid-onlyoverfree_only,--verboseoverquiet), as "the command line
wins" promised; argparse used to stop with a usage error. - A description that starts with a bulleted list is read back from a Markdown
catalog. Its "- item" lines were taken for more metadata, so the description came
back empty and--compareor--update-filereported a change that never
happened. - A TXT export of a listing with an empty title passes its post-write check; the
count required at least one character after "Title:" and the export failed. - A text holding a lone surrogate (valid in JSON, not in UTF-8) is written with
U+FFFD in its place; every writer used to fail after building the whole catalog. - A CSV tag that itself looks like a JSON list (
["a"]) is written in the JSON form,
so it reads back as written;pb-convertturned it intoa.
Security
- The web UI no longer shows the server's absolute paths: the output directory and
the written files are shown relative to the working directory, and a failed write
reports the reason and the relative file instead of the full OS error. The same
holds for a written file that fails its check and an unreadable comparison catalog. - Errors the HTTP server produces itself (501 for an unsupported method such as PUT
or OPTIONS, 400 for a malformed request) now carry the same security headers as
every other response.
v0.14.1
Fixed
- Found by the new fuzzing:
- A CSV catalog with a row longer than its header is now an error that names the
line. With--compare-csv-safeor--from-csv-safeit used to end in an
AttributeErrortraceback, and without them the surplus values were silently
kept under no column. - A profile written as a malformed URL, such as
https://[x/profile/acb(an
unclosed[reads as an IPv6 address), is now the usual "not a valid profile
URL" error;pbandpb-historyused to stop with a traceback. - Records that have neither a slug nor a title (a hand-edited CSV, say), or that
repeat a slug, are paired with an identical record first, so a catalog compared
with itself no longer reports them as both added and removed.
- A CSV catalog with a row longer than its header is now an error that names the
- A time stored as epoch milliseconds is now converted by counting from the epoch,
not withfromtimestamp, which depends on the platform: a catalog holding a time
outside years 1 to 9999 (a corrupt or hand-edited file) made every writer, and so
pb-convert, stop with anOverflowError, and on Windows any time before 1970
failed. Such a time is now unknown, like a missing one. For every real time the
text is unchanged (checked on all 4,826 timestamps of two live profiles and
200,000 random ones). Found by the weekly-length Atheris run.
Changed
- Development only: the code that reads catalogs, config files, and remote text is
fuzzed with Atheris (coverage-guided) in a newfuzzworkflow: catalog loaders
in every format,pb-convert's record reader and every writer, the text
sanitizers, profile and alert parsing, comparisons, and config files, each with
the guarantees it must keep. A seeded smoke mode runs the same checks without
Atheris, and the test suite runs it briefly.
Security
- The web UI no longer expands
~in the output directory field. Expanding it ran
a path operation on the submitted text before the working-directory check (on
POSIX,~namelooks up another account), and it pointed at the server user's
home, which the check then refused anyway, so nothing could escape. Now~is
an ordinary folder name inside the working directory, as in the comparison
catalog field. Reported by CodeQL (py/path-injection).
v0.14.0
Added
pb-history comparelines up every profile in a history file (or the ones named
with--profile) from its latest snapshot: listings, total views, sales,
favorites, and reviews, average rating, median price, share of free listings,
and sales per listing.--days Nadds per-day growth, andn/amarks a profile
whose history is shorter. Markdown, JSON, or HTML.pb-history report --alert RULEflags the listings that moved enough to act on:
a counter's gain (sales+1,views+50%) or an event (new,removed,price).
The hits come first in the Markdown, HTML, and JSON report, and the command exits
with2when any rule fires, so a scheduled workflow can open an issue or send
a message.
Fixed
- A comparison no longer depends on the order of the records. A new listing that
shared an existing listing's title, if it came first, used to be reported as a
change to that listing, and the unchanged original as added. Now every slug match
is made before any title match, a title only pairs records when one side has no
slug (a TXT catalog), and each old record is paired once; two old TXT records
with the same title used to leave one of them unmatched. Records that share a
title are paired with an identical old record first, so reordering them is not
reported as a change. Neither live profile
checked (@acb, @emanema) has repeated titles today; property tests found it.
Changed
-
The GitHub Action guide shows how to publish a dashboard on GitHub Pages every
day: the trends report, the searchable catalog, and an index linking them. Its
build step was run as written, with and without an earlier snapshot, and the
pages were checked in a browser (the catalog's search works, nothing is loaded
from other sites). -
A profile's listings and their descriptions are fetched at the same time instead
of one after the other, at most two connections at once. Fetching @emanema's
2,683 prompts went from 29.4 to 18.3 seconds (median of 3 live runs, same records
either way); a small profile like @acb is not measurably faster. -
Development only: every tool CI installs is pinned by hash and installed with
--require-hashes(requirements-dev.txtfor the checks,
requirements-release.txtfor building), and the exporter is installed with
--no-deps, since it has no dependencies.scripts/lock_requirements.pyrebuilds
the files, resolving the dependencies for each Python version CI uses. The
OpenSSF Scorecard counted these installs as unpinned. -
Each release also carries its signed Sigstore bundle
(promptbase_profile_exporter-X.Y.Z.sigstore.json), so a download can be verified
against that file withgh attestation verify --bundle. -
Development only: property tests compare random catalogs that were changed in
known ways (new listings, some reusing a title; removals; edits; counter moves;
re-wrapped text; TXT catalogs with repeated titles) and require the diff to
report exactly those changes, count each record once, and mirror when the sides
are swapped. -
Development only: CodeQL scans the Python code and the GitHub Actions workflows
(security-extended queries) on every pull request, onmain, and weekly; the
OpenSSF Scorecard grades the repository's supply-chain practices weekly and
feeds the new README badge.
v0.13.0
Added
- Each release now carries a wheel and an sdist, built by a GitHub Actions workflow
from the release tag with a signed build provenance attestation;
gh attestation verifychecks a downloaded file (see SECURITY.md). There is
still no PyPI package.
Deprecated
- Python 3.10 reaches end of life on 31 October 2026 (PEP 619). This is the last
release line that supports it: the first release after that date will need
Python 3.11 or newer, which also makes--configTOML files work everywhere.
Security
- Text from PromptBase can no longer change how a report renders. Titles and values
in a Markdown diff report (--diff-output report.md, and the Action's run summary)
are escaped, so a title cannot add links or headings or hide the rest of the
report behind<!--. Titles printed to the terminal or a CI log are kept on one
line with control characters removed, so a title cannot fake extra log lines, a
GitHub Actions workflow command (::error::), or a terminal escape sequence. - Responses from PromptBase are capped at 64 MiB, compressed or not (a real page is
well under 1 MiB), and a response or document nested too deeply is a clean error
instead of a traceback. - On a web server bound beyond loopback, the comparison catalog must be named like
an export, as downloads already were, so another machine cannot read values from
other JSON or CSV files in the working directory or probe which files exist. - The Action passes profiles after
--, so a profile value that starts with-
is never read as an option, and writes its artifact path output in the delimited
form, so a newline in a path input cannot add another output.
Fixed
- A server that keeps returning the same page (a cursor that does not advance) now
stops the fetch after the second request with a clear error; it used to repeat
the query up to the 100-page safety limit first. - Comparisons keep listing kinds apart: a prompt and an app are no longer paired
because they share a slug or, more often, a title. Comparing @acb's prompts with
its apps used to report 132 "changed" listings; it now reports 243 removed and 211
added. A TXT catalog, which records no kind, still matches by title. - An app is no longer required to carry a
typefield: apps have no model type, so
a future response without one is not reported as schema drift. - In a
pb-historyMarkdown report, a slug containing backticks can no longer end
its code span early. pb-web --openon a server bound to::openshttp://[::1]:port/.- In the Action, a cut-off run summary now points to the full report in the step's
log and thediff-outputinput; it used to point to a temporary file that is
deleted when the job ends.
Changed
- The pinned
actions/checkout(v7.0.1) andactions/setup-python(v7.0.0) are
updated in the Action, the workflows, and the documentation examples. - Development only: resilience tests run the client against a real local HTTP
server that answers with 5xx and 4xx errors, drops the connection before or in
the middle of a body, truncates a compressed body, responds after the timeout,
and pages with and without a working cursor. - Development only: CI installs exact versions of its tools from
requirements-dev.txt(Dependabot updates them), checks out without keeping the
token in.git/config, and every test job, Windows and macOS included, is now
required before a pull request can merge. - The GitHub Action guide shows a scheduled
pb-historyworkflow that keeps the
history in the Actions cache and posts each report to the run's summary page;
its commands were run as written from a clean install of v0.12.1. The nightly
canary now also records two livepb-historysnapshots and reports on them. - Documentation only: the README (both languages) now lists
pb-history, the web UI's
Preview button and--open,--config, andndjsonwhere it had not; the CLI guide's
contents list is complete; the coverage floor is stated as 85% everywhere.
v0.12.1
Fixed
- A CSV catalog is now read with its line breaks intact: a description or title
with\r\nor\rinside a quoted cell used to come back as\n, so--compare,
--update-file,pb-diff, andpb-convertcould see a change that never
happened. None of the profiles checked live (@acb,@emanema) has such a
value today; it was found by the new round-trip fuzz test. - A CSV export now quotes a cell that contains a lone carriage return on every
supported Python. Python 3.10, 3.11, and 3.13 left it unquoted, so a reader took
it as the end of the row; the same fuzz test failed there in CI. Output is
otherwise byte for byte what it was.
Changed
- Development only: a seeded fuzz test writes random records full of hostile
characters (quotes, separators, Unicode line breaks, emoji, formula prefixes) in
every lossless format and requires them to read back exactly; ruff now also
checks bugbear, comprehension, simplify, ruff-specific, and security rules; the
coverage gate is 85% (it was 70%, and the suite is at 93%).
v0.12.0
Added
pb-history(promptbase-history, orpython -m promptbase_exporter.history)
keeps snapshots of a profile's views, sales, downloads, favorites, reviews,
rating, and price in a SQLite file, and reports what moved:pb-history snapshot @acb --db history.sqliteon a schedule, thenpb-history report --db history.sqlitefor the totals before and after, the top movers by a chosen
counter, new and removed listings, and price changes, as Markdown, JSON, or a
self-contained HTML page with trend charts (--since DATEor--days Npick the
baseline). It uses the standard library'ssqlite3only, stores values through
parameterised queries, versions its file, and refuses a file that is not its own.- The web UI has a Preview matches button: it applies your filters, sort, and
limit and lists what an export would select (the counts, the text and image
split, how many have no description, and a table of the first 100 with links),
without writing anything. It is the first button, so Enter in a text field
previews rather than exports, and it passes the same cross-origin and Host
checks as an export.pb-web --openopens the page in your browser once the
server is listening.
Fixed
- In the web UI, the export result table now names the kind it exported ("Bundles",
"Apps") instead of always saying "Prompts", and a table that is wider than a
phone screen scrolls inside its own box instead of making the whole page scroll
sideways.
v0.11.0
Added
--config FILEreads options from a.jsonor.tomlfile (TOML needs
Python 3.11 or newer), with the long option names as keys plusprofiles, so
a recurring export is one short command. The file's entries become command-line
arguments in front of the real ones, so they go through the same validation and
anything typed on the command line overrides them. A problem in the file stops
the run before anything is fetched and names the file. Theprofileargument
is now optional when the file listsprofiles; with neither, the usual usage
error (exit code 2) is shown. Profiles and options may now alternate on the
command line (pb @a --mode all @b), which used to be rejected.- A new
ndjsonformat (--format ndjson,.ndjson;.jsonlis accepted on
input and inferred from--output-file): one compact JSON object per line
with the JSON format's fields and no enclosing array, forjq -c,
pandas.read_json(lines=True), and warehouse loaders. U+2028, U+2029, and
U+0085, whichstr.splitlines()treats as line breaks, are escaped so a
record is always one physical line. It works with--compare,
--update-file,pb-diff,pb-convert, the web UI, and the Action, and
each line of a written file is checked. --layout files(Action inputlayout) writes one Markdown file per prompt
instead of one file per catalog, into a folder per catalog
(exports/acb_all_prompts/<slug>.md), for note tools that index front matter.
The YAML front matter holds the record's fields as JSON scalars, which are valid
YAML, so every string is quoted and a title such asyesor2026-01-01is never
read back as a boolean or a date. File names are the slug made safe (no path
separators or leading dots, Windows device names suffixed) and, when it had to
change, carries a digest of the original slug, so a prompt's file never depends on
which other prompts a run selects; nothing is ever deleted. U+0085, the
U+007F to U+009F controls, U+2028, U+2029, and U+FEFF are escaped in the front
matter so a YAML reader reads every value back exactly. It needs Markdown and
cannot be combined with--output-file,--compare, or--update-file. The web
UI has a "Layout" list for it.
v0.10.0
Added
--extra-fields(Action inputextra-fields, a checkbox group in the web
UI) addstags,engine,nsfw,featured,updated,last_sale, and
unique_salesto the markdown, JSON, CSV, and HTML exports. Only the fields
you ask for are requested from PromptBase, so a default export is
byte-identical to before and asking for all of them adds about 8 KB to a
243-prompt download. A time PromptBase does not record isnull(empty in
CSV) rather than a zero, andtxt, which cannot hold them, is an error. In
CSV the tags are joined with,, or written as a JSON array if a tag
contains a comma, so the cell can always be read back.--item-type(Action inputitem-type, a "Kind" list in the web UI) exports
a profile'sbundles orapps as well as itsprompts (the default).
Bundles and apps take their descriptions from PromptBase's publicBundles
andAppDetailsrecords, are written to<user>_<mode>_bundles.<ext>and
..._apps.<ext>so they never overwrite a prompt catalog, link to
promptbase.com/bundle/<slug>and/app/<slug>, and carry anitem_type
column; a prompt catalog keeps its original columns. Apps mirror the
profile's prompts, are free, and have notype. Skills are not supported.
The HTML catalog's markup gained adata-nounattribute so its heading and
search count name the right kind.pb-convert(promptbase-convert, orpython -m promptbase_exporter.convert)
rewrites a saved catalog in another format without fetching anything. Its
result is the file a direct export would have written, byte for byte, extra
fields and bundle or app catalogs included. It reads JSON, CSV, and HTML
catalogs and refuses TXT and Markdown, which lack columns and would need
invented values; it stops on a missing column, an empty number, or an
unreadable value instead of guessing (a boolean where a number belongs, an
entry that is not a record, or an extra field missing from some records
included), checks the result in a scratch file before it replaces anything,
and never overwrites its source.--csv-safeand--from-csv-safecover
protected CSVs.- With
compareorupdate-file, the GitHub Action now shows the diff report
on the workflow run's summary page (step-summary, on by default,false
turns it off). A report over 900 KB is cut, with a note to usediff-json;
the summary is still written whenfail-on-difffails the step. The Action
docs gain a recipe that opens a pull request when the catalog changes.
Fixed
- The commands no longer crash with a
UnicodeEncodeErrorwhen their output is
redirected to a pipe or a log file on Windows and a message names a prompt
title (or a path) that the console's legacy code page cannot encode, such as
one with an emoji. Such characters are printed as?; catalogs are still
written as UTF-8.