You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Short answer: nobody has published a benchmark that settles it, so what follows is every figure this repo has collected on the question, each one kept with the sentence it was published in and the date it was read. Corrections and counter-numbers are the point of this thread.
What the loop is. Rather than tuning one prompt until the output is right, the agent runs a fixed cycle — act, observe, correct — and the developer edits the loop instead of the prompt.
The figures that exist, and where each came from
5x — "Last week, a developer in the Chinese AI community shared groundbreaking results using 'loop engineering' to boost AI agent efficiency by 5x." (2026-08-09)
5x — "A developer in China's AI community achieved 5x productivity gains using loop engineering, reducing MVP development time from four prompt tuning sessions to a single command." (2026-08-09)
74% — "By consolidating market insight, media strategy, creative generation, smart deployment, and data analysis into a unified system, it achieved a 74% improvement in optimization efficiency." (2026-08-09)
40 seconds — "One implementation reduced average response times from hours to 40 seconds and increased transaction volumes by 50%." (2026-08-09)
8 hours — "This system reduced manual processing time from 8 hours to minutes while improving content organization quality." (2026-08-09)
89% — "I don't buy the 89% claim." (2026-08-09)
That last row is in the table on purpose. A figure that the write-up itself refuses is still a figure someone published, and hiding it would make this list look more settled than it is.
What none of these are. None is a controlled measurement. Every one is a self-report by the party that benefited from it, restated in a write-up. 5x appears twice from the same origin, which is one claim, not two. Treat the whole list as "what is being claimed", not "what is true".
The one thing worth replying with: if you have run an agent in a loop rather than re-prompting it, what did the number actually turn out to be for you — and on what task? A one-line reply with a task and a ratio is worth more than any of the rows above, and it goes into the table with your wording kept.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Short answer: nobody has published a benchmark that settles it, so what follows is every figure this repo has collected on the question, each one kept with the sentence it was published in and the date it was read. Corrections and counter-numbers are the point of this thread.
What the loop is. Rather than tuning one prompt until the output is right, the agent runs a fixed cycle — act, observe, correct — and the developer edits the loop instead of the prompt.
The figures that exist, and where each came from
5x— "Last week, a developer in the Chinese AI community shared groundbreaking results using 'loop engineering' to boost AI agent efficiency by 5x." (2026-08-09)5x— "A developer in China's AI community achieved 5x productivity gains using loop engineering, reducing MVP development time from four prompt tuning sessions to a single command." (2026-08-09)74%— "By consolidating market insight, media strategy, creative generation, smart deployment, and data analysis into a unified system, it achieved a 74% improvement in optimization efficiency." (2026-08-09)40 seconds— "One implementation reduced average response times from hours to 40 seconds and increased transaction volumes by 50%." (2026-08-09)8 hours— "This system reduced manual processing time from 8 hours to minutes while improving content organization quality." (2026-08-09)89%— "I don't buy the 89% claim." (2026-08-09)That last row is in the table on purpose. A figure that the write-up itself refuses is still a figure someone published, and hiding it would make this list look more settled than it is.
What none of these are. None is a controlled measurement. Every one is a self-report by the party that benefited from it, restated in a write-up.
5xappears twice from the same origin, which is one claim, not two. Treat the whole list as "what is being claimed", not "what is true".Write-up with the full context: https://xyzs996.github.io/llm-api-pricing/articles/ai-agent-loop-engineering-karpathy-s-method-for-5x.html
All 296 figures as JSON/CSV: https://xyzs996.github.io/llm-api-pricing/figures.html
The one thing worth replying with: if you have run an agent in a loop rather than re-prompting it, what did the number actually turn out to be for you — and on what task? A one-line reply with a task and a ratio is worth more than any of the rows above, and it goes into the table with your wording kept.
All reactions