Replaying LLM extraction offline — and a measured note: crawl4ai honours both OPENAI_BASE_URL and OPENAI_API_BASE #2274
xizhuomengcontin
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
LLM extraction is the part of a crawl that costs money, and it is also the part you end up re-running the most — you tweak the schema, the instruction or the chunking, and pay for the whole extraction again to find out whether anything changed.
I wanted to stop paying for that loop, so I checked something first and then built the workflow around it.
The measured bit (useful on its own)
perform_completion_with_backoff— the function behindLLMExtractionStrategy— picks up a base URL from the environment even when you do not setLLMConfig(base_url=...). I ran a local HTTP server that returns a canned chat-completion and logs the path it is asked for:Both spellings work. So you can point crawl4ai's extraction at a local endpoint, a gateway, or a proxy without touching your code or your
LLMConfig. That is worth knowing regardless of the rest of this post — it is not true of every framework I have probed this way (one honours only one of the two names; one ignores both because it pins the origin in its own config).What I do with it
I run the crawl once under a recorder that moves that variable, then replay the extraction offline as many times as I want:
orca replay last --from N --model <other>replays the run up to step N from the recording and then continues on a different model — everything before the fork is byte-identical, so the model is the only variable when you are comparing extraction quality.Disclosure: OrcaReplay is mine (Apache-2.0, Node 20+, no account). The measurement above is about crawl4ai and holds whatever you use it with.
Three things it does not do
Stating these because the failure mode of tools like this is people trusting them past their limits:
Happy to be told the base-URL behaviour above is documented somewhere I missed, or that there is a neater way to do the same thing with
LLMConfigalone.All reactions