Releases: ctdio/textopt
Release list
textopt@0.3.0
Minor Changes
-
a3354ac: A score an adapter synthesized after catching an error is never written to the
evaluation cache. Adapters say so withfailedonScoreResult, which takes no
judgement about the provider;transientstill decides what is retried and kept
out of a candidate's mean, and still classifies nothing by default. A run that
finished with failures nothing classified reports how many under the new
unclassifiedFailureswarning code.A repeat run over a warm
file-cachethat hit failures now re-runs those
instances rather than reading their zeros back, so it can spend more rollouts
and reach a different winner than the run before it. -
12789b5: A
reflectcall that throws is retried under the run's existingretry
policy, rather than ending the run. Nothing is classified on the way past: a
proposal model is a pure request, so the attempt after a transport failure is
free to succeed and a genuine bug failsattemptsmore times and surfaces
unchanged. Every attempt counts against the optimizer's reflection ceiling —
GEPA'sreflection.maxCalls, OPRO's and SIMBA'smaxReflectionCalls— which a
round can now overrun by up toattempts. Setretry: { attempts: 0 }to keep
the old behaviour.withRetries(model, policy)is exported for applying the
same policy to any otherTextModel. -
0aa5911: Every optimizer emits a
rolloutevent carryingcompleted/totalalongside
the phase, split and candidate it belongs to, so a run reports progress between
batches instead of going quiet for the length of a validation sweep. An adapter
opts in by callingargs.onRolloutfrom itsevaluate— the AI SDK and
LangChain adapters already do, passing it asonSettledto
mapWithConcurrency. A consumer switching exhaustively over an optimizer's
event union must handle the new member.
Patch Changes
- 0aa5911:
candidateAcceptedcarriesobjectiveScores, the per-objective mean over the
instances the candidate was measured on. A single objective collapsing while the
aggregate holds is how a degenerate metric channel announces itself, and it was
previously visible only in the result. - b62adf6:
createPromptAdapter({ run, score })fromtextopt/gepais the single-prompt
case named:runreceives the candidate's text asinstructionandscore
grades the output. It reads which component to run off the candidate and throws
when there is more than one, because a component no module runs is text the
search rewrites every iteration for no effect — that system wants
createPipelineAdapter. Neither helper is GEPA-only, and the docs now say so:
aGepaAdapteris the baseAdapterwith reflection's evidence added, so the
same adapter passes unchanged to SIMBA, OPRO, MIPRO and both searches. - 0aa5911:
createReporter({ on })takes a handler map whose keys are checked against the
optimizer's event union, so a misspelled event name is a compile error rather
than a reporter that runs to completion having seen nothing. A reporter built
this way also declares what it handles, and a run warns at start about a
handler named for an event no optimizer emits, or about a reporter whose
handlers all miss the run it is attached to. A handler for an event some other
optimizer emits is not warned about: one reporter written for several searches
is what the shared events are for. - 0aa5911: Two reporters ship with the library:
consoleReporter({ level })prints one
line per event —quietfor acceptances and the finish,verbosefor
everything — andjsonlReporter({ path })fromtextopt/file-reporterappends
each event as a JSON line, leaving structured data structured instead of
flattening it into log prose.
@textopt/langsmith@0.1.2
@textopt/langchain@0.3.0
Patch Changes
- a3354ac: The adapter marks a run or scoring failure it caught with
failed, so the zero
it stands in for is never written to the evaluation cache.isTransientstill
decides only what is worth retrying. - 0aa5911: Every optimizer emits a
rolloutevent carryingcompleted/totalalongside
the phase, split and candidate it belongs to, so a run reports progress between
batches instead of going quiet for the length of a validation sweep. An adapter
opts in by callingargs.onRolloutfrom itsevaluate— the AI SDK and
LangChain adapters already do, passing it asonSettledto
mapWithConcurrency. A consumer switching exhaustively over an optimizer's
event union must handle the new member. - Updated dependencies [0aa5911]
- Updated dependencies [a3354ac]
- Updated dependencies [b62adf6]
- Updated dependencies [0aa5911]
- Updated dependencies [12789b5]
- Updated dependencies [0aa5911]
- Updated dependencies [0aa5911]
- textopt@0.3.0
textopt@0.2.0
Minor Changes
- c2c4d3e:
GepaOptimizerrefuses an acceptance policy that cannot accept anything at the
configuredminibatchSize. A sign-flip test over three instances bottoms out at
p = 0.125, sopairedPermutationAcceptance({ alpha: 0.05 })at the default
minibatch rejected every proposal and returned the seed after spending the whole
budget. Acceptance policies report the batch they need asminimumPairs; raise
minibatchSizeto at least that, or loosenalpha. - c2c4d3e:
createFileCacherequires anamespacenaming the system its scores measure,
and never serves an entry written under a different one. A durable log outlives
the model behind an alias, the decoding settings, and the scorer version — every
part of a measurement a cache key does not name. Pass the same string you would
pass ascacheNamespace. - c2c4d3e:
GepaAdapter.evaluatereturns aReflectiveBatch— anEvaluationBatchwith
feedbackrequired. An adapter that returned scores and no prose left the
reflection prompt rewriting instructions from empty feedback blocks, which is
blind search reported as a normal run. Adapters that already returnfeedback
need no change; the rest now fail to compile. - c2c4d3e:
JudgeCriteriontakesweightandgate. A gated criterion that grades below
its bar scores the instance 0 whatever the other criteria said, so a hard
requirement is no longer something a search can trade away against three
cosmetic ones. Whenexpectedis passed, the default judge prompt also forbids
the feedback from restating it — feedback is rewritten into a reusable
instruction, and a fact copied out of the gold answer becomes an answer key
memorised in the prompt. - c2c4d3e: Every result and
finishevent carrieswarnings: what a run could see about
its own measurement that its numbers cannot say. A run given novalidationSet
reports that selection ran on the instances reflection read, and a seed the
metric scores identically on every validation instance reports that there was
nothing to rank. PassvalidationSet: "reuseTraining"to accept the reuse by
name. Custom optimizers implementingOptimizerResultor emittingRunFinished
must now populatewarnings.
Patch Changes
-
3f5824b: The long-form guides ship in the tarball, under
docs/, so an installed copy
documents the version installed rather than whatevermainhas become. Two of
them are new:data-prep.mdon splitting a dataset a search can be trusted
with, andmetric-preflight.mdon checking a metric separates candidates —
and moves over a wide enough interval — before a budget is spent on it.Doc comments on the API carry the traps that belong beside the code: that
weight: 0removes a criterion from the aggregate but not from
objectiveScores, that agateand a heavyweighton the same criterion
enforce a requirement twice and narrow the range a search has left to move in,
and that a SIMBA run has to be funded past its finalist reserve before any step
happens.
@textopt/langsmith@0.1.1
@textopt/langsmith@0.1.0
Minor Changes
- 4b7df44:
@textopt/langsmithships to npm.createLangSmithReporterwrites any optimizer's run to LangSmith as one experiment per accepted candidate, and matches LangSmith'sClientstructurally, so it adds no runtime dependency on thelangsmithSDK.
@textopt/langchain@0.2.0
textopt@0.1.0
Minor Changes
-
ca8a541: A harvested rollout cannot end the demo block it is stored in.
Demo blocks are delimited by
<demo>, and the values inside them were written
raw. A system that quotes its own prompt back produces the one output that
breaks: a rollout worth keeping whose text carries</demo>, which closes the
block early and leaves the rest of it as loose text.parseDemosthen returned
a demo that was not the one stored, and SIMBA reparses and rewrites its demo
components at every step, so the loss compounded over a run instead of showing
up once.<demo>,<input>and<output>are now escaped in serialized demo values and
unescaped on the way back, so a demo containing a demo round trips as itself.
The escape is escaped first, so a value that already reads<demo>survives
too. Demo blocks written by earlier versions still parse; blocks whose values
contain those tags will render them escaped from now on, which changes the text
a component holds and so the candidates a run compares.A custom
renderDemois responsible for its own escaping: the library cannot
know which part of what a renderer emits is the delimiter it meant to write. -
b25abd2: SIMBA's advice proposer sees what each component already says.
buildAdvicePromptnamed the components it wanted advice for but never showed
their text, while the advice it produces is appended to that text rather than
replacing it. A proposer that cannot read what it is appending to writes blind:
it restates guidance the component already carries, and it cannot correct
guidance that is wrong, since contradicting a line it never saw is not something
it can choose to do. SIMBA's reference implementation passes the current
instructions for exactly this reason.AdvicePromptArgsnow carriescurrent, a map from component name to what that
component holds, and the prompt renders each as a<component name="…">… </component>block followed by the instruction not to restate advice already
present. A customAdvicePromptBuilderreceives the extra field and may ignore
it; anything constructingAdvicePromptArgsby hand, or parsing the built
prompt's component list, has to be updated. -
fec8f51: Ceilings hold where a run actually spends, checkpoints describe whole rounds,
and an instance id names one row.Harvesting takes a
maxCostUsdof its own.harvestRolloutsand
harvestFewShotExamplescheck it between batches, and MIPRO and bootstrapped
few-shot search pass what is left of the run's ceiling into each pass. MIPRO
also stops building demo sets once the ceiling is reached. A demo menu is many
evaluations on a separate evaluator, so a ceiling it never consulted bounded
only the trial loop that followed it, and a run could spend its whole allowance
choosing demos and never score a candidate.OPRO and MIPRO checkpoint after the sweep their cadence schedules, not before
it. A snapshot names a round, and a resumed run schedules its next sweep an
interval past the round the snapshot names, so a checkpoint taken first
described half a round and the resumed run skipped that sweep entirely. A MIPRO
trial whose rollouts all failed transiently now checkpoints and runs its cadence
like any other: the rollouts were bought and the counter moved either way.Bootstrapped few-shot search reads every sweep it dispatched before it leaves a
wave. Stopping on the first failure abandoned the sweeps behind it, which went
on calling the adapter, and spending, after the caller had been handed the
error.An adapter reading of
NaN,Infinityor a negative token count is refused
where it enters. Folded into the totals it silently disabledmaxCostUsd:
every later comparison against aNaNcost is false, so the ceiling stopped
holding without saying so.createJudgerefuses ascalethat is zero or
negative for the same reason — every grade is divided by it.Instance ids fall back to the row's position for a datum a content hash cannot
read, not only for one that will not serialize. AMap, aSetand a class
instance holding its state privately all serialize to{}, so distinct rows
shared an id and were served each other's cached scores. The six optimizers now
share onedefaultInstanceIdrather than six copies of it.A file cache terminates a record its previous process left half-written. The
truncated record was already lost; appending onto the line it left open lost the
next one too.The LangChain adapter counts a provider that reports tokens in both shapes once.
Integrations that fillllmOutput.tokenUsageandusage_metadatafor the same
call were billed twice, and the message-level shape now wins with the legacy
total as a fallback. Usage a scorer reports is no longer dropped when the run
itself counted none — the guard consulted only the callback total, so a judge's
spend went unreported. -
b25abd2: A demonstration lands in a component without erasing what else it says.
SIMBA's
appendDemorebuilt a demo component from its demos alone, so any text
the component held that was not a demo block was gone the first time a rollout
was harvested into it — including the adviceappendRulehad just written
there. The two mutations could not share a component, which is why
instructionComponentsdefaulted to the componentsdemoComponentsdid not
name, and why a candidate with a single component got one mutation instead of
two. SIMBA's reference implementation appends demos and instructions to the same
predictor; a component is the closest thing this library has to one.replaceDemosrewrites the demo blocks in a text and leaves the rest of it
alone, and bothappendDemoand the loop's demo-dropping now go through it. A
component named indemoComponentscan hold instructions too, and the default
instructionComponentsfalls back to every component when every component holds
demos, rather than leavingappendRulewith nowhere to write and throwing.Runs where demo and instruction components were already disjoint are unaffected
except that a demo block now keeps its position in the text rather than
replacing it. Runs where they overlapped were losing text and are not
comparable to their old results. -
155cb19: First public release: GEPA, SIMBA, OPRO, MIPRO, bootstrapped few-shot search,
and random search behind a shared optimizer interface, with a LangChain
adapter. The shared substrate handles budgets in rollouts, dollars and wall
clock, transient-failure retry, durable caching, checkpoints and resume for
every optimizer, a model-graded judge, a multi-module pipeline adapter, and
compare()for deciding between two runs with a p-value rather than a hunch.
Every run reports throughreporters: any number of observers, each with an
onEventcalled synchronously and aflushawaited as the run ends, including
when it ends by throwing. Each search emits its own event union, but every one
of them announces an accepted candidate with the text that scored and its row
over the validation set, and reports the held-out sweep per instance rather
than as a lone mean. A reporter that reads only those narrows with
isCandidateAcceptedandisRunFinishedand drops into any optimizer.The Vercel AI SDK, Braintrust and LangSmith packages are still in beta and ship
from the repository only. -
38dfb12: A ceiling bounds what it was told it bounds, and a reading it cannot be checked
against is refused wherever it enters.The held-out sweep is reported apart from the search.
maxCostUsdand
maxWallClockMsbound the search loop, and atestSetis measured once after
that loop has already stopped — charging it would let the size of a held-out set
decide which candidate wins.metricCallsalready excluded those rollouts and
reported them astestMetricCalls;usagedid not, so a run that honoured a $4
ceiling and then swept a hundred held-out rows reported spending $104 and looked
like it had overrun. Their dollars are nowtestUsage, alongside
testMetricCalls, in all six optimizers. SIMBA and bootstrapped few-shot search
counttestMetricCallsthe way the other four already did — the rollouts the
sweep bought, not the rows it was handed, which differ once one is retried or
served from the cache.EvaluatorgainsunchargedUsage()for the same split, andusage()now
returns only what the search spent. Nothing bounds the held-out sweep, so what a
run costs end to end isusageplustestUsage: budget for atestSetthe way
you would budget for one full validation sweep.A usage reading that is not a non-negative finite number is refused whether it is
a number or not. The guard only fired onNaN,Infinityand negatives, so a
JavaScript adapter reportingcostUsdas a string concatenated onto the totals
and left the ceiling checked against text. Resumed usage and usage absorbed from
a harvest are checked too: a hand-edited checkpoint could poison the totals with
nothing to say it had.The LangChain adapter reads the legacy total again when the message-level shape
counts nothing. Preferringusage_metadataon its presence rather than on what
it carries meant an integration attaching a zeroed one — as some do to
intermediate generations — reported zero tokens for a call whose real total was
sitting inllmOutput.tokenUsageall along. -
2606e36: Rollout accounting survives a resume, and the numbers a run reports name the
candidate it returns.maxCostUsdis a ceiling on the run rather than on the segment: every snapshot
now carries the usage already spent, and a resumed run folds it back in instead
of restarting its token and dollar totals at zero. Harvesting is part of that
total —harvestRolloutsandharvestFewShotExamplesreport theusagetheir
own evaluat...
@textopt/langchain@0.1.0
Minor Changes
-
155cb19: First public release: GEPA, SIMBA, OPRO, MIPRO, bootstrapped few-shot search,
and random search behind a shared optimizer interface, with a LangChain
adapter. The shared substrate handles budgets in rollouts, dollars and wall
clock, transient-failure retry, durable caching, checkpoints and resume for
every optimizer, a model-graded judge, a multi-module pipeline adapter, and
compare()for deciding between two runs with a p-value rather than a hunch.
Every run reports throughreporters: any number of observers, each with an
onEventcalled synchronously and aflushawaited as the run ends, including
when it ends by throwing. Each search emits its own event union, but every one
of them announces an accepted candidate with the text that scored and its row
over the validation set, and reports the held-out sweep per instance rather
than as a lone mean. A reporter that reads only those narrows with
isCandidateAcceptedandisRunFinishedand drops into any optimizer.The Vercel AI SDK, Braintrust and LangSmith packages are still in beta and ship
from the repository only.
Patch Changes
-
fec8f51: Ceilings hold where a run actually spends, checkpoints describe whole rounds,
and an instance id names one row.Harvesting takes a
maxCostUsdof its own.harvestRolloutsand
harvestFewShotExamplescheck it between batches, and MIPRO and bootstrapped
few-shot search pass what is left of the run's ceiling into each pass. MIPRO
also stops building demo sets once the ceiling is reached. A demo menu is many
evaluations on a separate evaluator, so a ceiling it never consulted bounded
only the trial loop that followed it, and a run could spend its whole allowance
choosing demos and never score a candidate.OPRO and MIPRO checkpoint after the sweep their cadence schedules, not before
it. A snapshot names a round, and a resumed run schedules its next sweep an
interval past the round the snapshot names, so a checkpoint taken first
described half a round and the resumed run skipped that sweep entirely. A MIPRO
trial whose rollouts all failed transiently now checkpoints and runs its cadence
like any other: the rollouts were bought and the counter moved either way.Bootstrapped few-shot search reads every sweep it dispatched before it leaves a
wave. Stopping on the first failure abandoned the sweeps behind it, which went
on calling the adapter, and spending, after the caller had been handed the
error.An adapter reading of
NaN,Infinityor a negative token count is refused
where it enters. Folded into the totals it silently disabledmaxCostUsd:
every later comparison against aNaNcost is false, so the ceiling stopped
holding without saying so.createJudgerefuses ascalethat is zero or
negative for the same reason — every grade is divided by it.Instance ids fall back to the row's position for a datum a content hash cannot
read, not only for one that will not serialize. AMap, aSetand a class
instance holding its state privately all serialize to{}, so distinct rows
shared an id and were served each other's cached scores. The six optimizers now
share onedefaultInstanceIdrather than six copies of it.A file cache terminates a record its previous process left half-written. The
truncated record was already lost; appending onto the line it left open lost the
next one too.The LangChain adapter counts a provider that reports tokens in both shapes once.
Integrations that fillllmOutput.tokenUsageandusage_metadatafor the same
call were billed twice, and the message-level shape now wins with the legacy
total as a fallback. Usage a scorer reports is no longer dropped when the run
itself counted none — the guard consulted only the callback total, so a judge's
spend went unreported. -
38dfb12: A ceiling bounds what it was told it bounds, and a reading it cannot be checked
against is refused wherever it enters.The held-out sweep is reported apart from the search.
maxCostUsdand
maxWallClockMsbound the search loop, and atestSetis measured once after
that loop has already stopped — charging it would let the size of a held-out set
decide which candidate wins.metricCallsalready excluded those rollouts and
reported them astestMetricCalls;usagedid not, so a run that honoured a $4
ceiling and then swept a hundred held-out rows reported spending $104 and looked
like it had overrun. Their dollars are nowtestUsage, alongside
testMetricCalls, in all six optimizers. SIMBA and bootstrapped few-shot search
counttestMetricCallsthe way the other four already did — the rollouts the
sweep bought, not the rows it was handed, which differ once one is retried or
served from the cache.EvaluatorgainsunchargedUsage()for the same split, andusage()now
returns only what the search spent. Nothing bounds the held-out sweep, so what a
run costs end to end isusageplustestUsage: budget for atestSetthe way
you would budget for one full validation sweep.A usage reading that is not a non-negative finite number is refused whether it is
a number or not. The guard only fired onNaN,Infinityand negatives, so a
JavaScript adapter reportingcostUsdas a string concatenated onto the totals
and left the ceiling checked against text. Resumed usage and usage absorbed from
a harvest are checked too: a hand-edited checkpoint could poison the totals with
nothing to say it had.The LangChain adapter reads the legacy total again when the message-level shape
counts nothing. Preferringusage_metadataon its presence rather than on what
it carries meant an integration attaching a zeroed one — as some do to
intermediate generations — reported zero tokens for a call whose real total was
sitting inllmOutput.tokenUsageall along. -
Updated dependencies [ca8a541]
-
Updated dependencies [b25abd2]
-
Updated dependencies [fec8f51]
-
Updated dependencies [b25abd2]
-
Updated dependencies [155cb19]
-
Updated dependencies [3336175]
-
Updated dependencies [38dfb12]
-
Updated dependencies [7809e1d]
-
Updated dependencies [2606e36]
-
Updated dependencies [b2be4d6]
- textopt@0.1.0