Releases: lncrawl/tor-pool
Release list
v0.4.0
Added
SESSION_PORT_BASE— one credential-free SOCKS port per instance, at base+N, for callers that cannot send a username. A browser is the case that forced it: Chrome refuses--proxy-serveroutright when the URL carries credentials, and Firefox cannot supply them either. Dropping the username instead would fall back toDEFAULT_SESSIONand key by client IP, so a browser and the crawler reusing its work would land on two instances and two exit relays — and a clearance replayed from an address that did not earn it reads as the site refusing you.GET /api/sessions/{key}now answers withsession_portso a caller is told its port rather than computing one from a setting it cannot see. RequiresAUTH_DISABLEDand refuses to start without it, since a credential-free port beside listeners that demand a password would undo them silently.
v0.3.0
Added
AUTH_DISABLED=trueturns off every credential check — SOCKS5, the HTTP proxy's407and the API's bearer token — so a plainsocks5h://127.0.0.1:9250works and the dashboard opens straight to the pool. Only set it where nothing else can reach the ports: whoever can open a socket gets your Tor bandwidth, the session table and the ability to restart instances.GET /api/auth/statusreports whether authentication is required.- Typed failure reports.
POST /api/sessions/{key}/failureacceptskind:captcha,blocked,rate_limited,transportorother. A captcha means the exit IP is burnt, so two reports quarantine the instance; a 429 means the exit still works, so it counts for less than one and never tripsQUARANTINE_CONSECUTIVE.reasonstill works, and a bodylessPOSTis still worth one ordinary failure. failure_score,quarantine_scoreandfailures_by_kindin each instance's health, with matchingtorpool_*metrics — the score, not the report count, is what says how close an instance is to quarantine.- Documentation at lncrawl.github.io/tor-pool, built from
docs/on every push tomain.
Changed
compose.ymldefaults toAUTH_DISABLED=true, since it publishes every port to127.0.0.1. If you deploy from a checkout of this repo, setAUTH_DISABLED=falsein your.envbefore upgrading — a restart otherwise turns authentication off, and if you had widened a*_PUBLISHline the pool becomes reachable and unauthenticated in the same step. The image's own default is stillfalse, sodocker runand your own compose file are unaffected.- Tokens and token ids are base62 —
tp_plus 22 alphanumerics instead of base64url. Same bits, no-or_for a terminal to word-break or a double-click to half-select. Tokens issued before this keep working. DELETE /api/sessions/{key}needsproxy, notadmin, so a client can release the session it created. A scraper holds aproxytoken by design, so until now every session it opened sat in its slot untilSESSION_TTL, and enough of them ran the pool out of capacity.GET /api/sessionsstaysadmin.
Fixed
- The dashboard no longer polls the session list from a hidden tab. Opening Sessions started a 3-second
GET /api/sessionsloop that then ran for the life of the page. - The Python examples in
README.mdanddocs/scraper.mdrun againstlncrawl-scraper1.x, and the stickiness promise now namesPIN_EXIT_RELAY— without it one instance can hand a caller more than one exit IP.
v0.2.0
Authentication. Every existing caller breaks until it presents a credential — see
the migration note below.
Breaking
-
The proxy password is now a credential. The SOCKS5 username is still the session
key, but the password — previously read and discarded — must be a token. A caller
offering no credentials is refused during the SOCKS5 handshake, and the HTTP proxy
answers407with aProxy-Authenticatechallenge.before socks5h://my-session:x@host:9250 after socks5h://my-session:tp_7Kq2mXvR8nB4jL6wYtZaPc@host:9250 -
The REST API and dashboard require a credential.
GET /health,GET /metrics,
POST /api/auth/loginand the dashboard's static assets stay public; everything else
under/api/answers401. An unmatched/api/path is authenticated before it
404s, so the surface no longer reports which endpoints exist.
No released lncrawl-scraper is affected: its tor-pool support is still unreleased and
carries a token field from the start.
Migrating. On the first boot after upgrading, a dashboard password and a
proxy-scoped token are generated and printed once to the container log. Take the token
and use it as the password in every proxy URL. To provision them from config instead,
set ADMIN_PASSWORD and PROXY_TOKEN. Both live in DATA_DIR, so mount the volume or
they are regenerated on every recreate.
Added
- Issued tokens. Mint, name, scope and revoke credentials from the dashboard's new
Tokens tab. Only a digest is stored, so a secret is shown exactly once. Revoking takes
effect immediately — before the change reaches disk, so it cannot be undone by a
restart. - Two scopes.
proxycovers traffic plus the session routes a caller uses to manage
its own sessions;admincovers everything. Give a scraperproxy: underadminthe
credential in its config could also resize the pool and read every session key. - A sign-in screen, and
ADMIN_USER/ADMIN_PASSWORDto configure the operator
login. Changing either invalidates every outstanding session immediately, which is the
answer to a credential exposure — there is no separate "sign out everywhere". PROXY_TOKEN, a proxy credential fixed by configuration for deployments
provisioned from files rather than by hand. Verified like any token but never
persisted, so the environment stays authoritative.LOGIN_TTLandLOGIN_RATE_LIMIT. Repeated wrong passwords from one address are
refused with429; a wrong username and a wrong password are answered identically.authevents in the audit log for sign-ins and for tokens issued or revoked.
Refused proxy credentials are logged to stderr instead: the event ring is bounded, so
one entry per rejected connection would let anyone flush the audit history in seconds.- Weekly rebuild of the moving image tags against the current Alpine
tor, so a Tor
security release reaches users without waiting for a tor-pool release. It refreshes
edge,latest,X.YandX, never the exactX.Y.Z— that one promises the same
bytes every time. Each target has to boot a pool and bootstrap a circuit before it is
published.
Fixed
- A destination Tor was never going to reach no longer quarantines the instance. Tor
refuses a private or loopback address, and that refusal was scored against the
instance — three requests for127.0.0.1were enough to quarantine a healthy one, and
enough of them emptied the pool. The request is still counted as failed; only the
remediation ladder is spared. Hostnames are unaffected, since a name that will not
resolve genuinely can mean a broken circuit. - The SOCKS5 handshake has a read deadline. A client that connected and sent one byte
held a goroutine and a file descriptor indefinitely. - The HTTP proxy authenticates every request on a keep-alive connection, not just the
first, so a later request cannot ride on an earlier one's credential.
Changed
DEFAULT_SESSIONnow decides what happens when an authenticated caller names no
session, rather than when a caller sends no credentials.- A session key is still not a tenancy boundary: any valid token may claim any key, so
sessions separate exit identities rather than callers. Now documented rather than
implied.
v0.1.0
The first release with notes. The earlier v0.0.x tags were created automatically by
every push to main, so their numbers counted pushes rather than changes and none of
them was ever accompanied by an entry here — this release is also where that stopped.
Added
- A pool of Tor instances in one container, behind a single sticky SOCKS5 and HTTP proxy
endpoint. - Sticky sessions. The SOCKS5 username (or
Proxy-Authorizationuser) is a session
key; a caller keeps the same instance, and so the same exit IP, until it rotates.
Callers with no credentials are pinned by client IP. - Instant rotation.
POST /api/sessions/{key}/rotatereassigns a session to an
already-built instance, skipping Tor's ~10s NEWNYM cooldown. - Failure-driven remediation. Failures are counted per instance from transport
errors and from client reports, and a bad instance escalates through new circuit →
wipe-restart → restart with exponential backoff. - Management dashboard with live updates over SSE: instance grid with per-instance
actions, sessions view, filterable audit log, and timeline charts. - REST API for instances, sessions, events and history, plus live pool resize.
- Prometheus metrics at
/metrics, and a/healthcheck that reports routability
rather than process health. - Multi-arch images (
linux/amd64,linux/arm64) onghcr.io/lncrawl/tor-pool. PIN_EXIT_RELAYlocks each instance to a single exit relay, so one instance really
is one exit IP until it rotates. Off by default: a pinned instance depends on one relay.BOOTSTRAP_STALL_TIMEOUTrestarts an instance that stops making bootstrap progress,
keeping its state on the first attempt and wiping it on the next.exit_confirmedandpinned_exiton the instance API, surfaced in the dashboard: an
exit no traffic has used yet is shown as the guess it is.
Changed
latestnow means the newest release, not the last push tomain. Every push used
to bump a patch tag and movelatest, so a README fix became a version number and
unreleased work reached everyone trackinglatest. Pushes tomainpublishedge;
releases are cut deliberately fromCHANGELOG.md. The weekly rebuild is gone with it —
tornow updates when you pull a newer image rather than on a timer.- Conflux is off by default (
TOR_CONFLUX). Each set Tor pre-builds has its own exit
relay and successive requests land on different sets, so one instance handed a caller
several exit IPs with no rotation at all. POST /api/instances/{id}/rotatereturns as soon as the instance is out of service,
finishing Tor's cooldown in the background, instead of holding the request open for up to
~13 seconds.
Fixed
- Rotation no longer drops requests in flight. Retiring an instance's circuits spared
only the ones carrying a connected stream, so a request still waiting for its exit to
reach the destination had its circuit closed underneath it. Measured at 4–5% of requests
failing while rotating under load, against 0% at rest. Any circuit with a stream on it is
now left standing, whatever state that stream is in. - A rotation no longer quarantines the instance it rotated. The failures a rotation
causes were scored against the instance, so a few rotations were enough to quarantine a
healthy one — whose remediation rotated it again. Failures inside an instance's own
rotation window are no longer counted against it. POST /api/pool/rotatekeeps the pool serving. It rotated every instance at once,
leaving nothing to route to for a second or two. It now sweeps one instance at a time and
returns immediately, reporting whether a sweep was already running.- The reported exit IP no longer jumps after a rotation. Tor holds several
exit-bearing circuits and builds more preemptively, and the API named whichever looked
newest — an exit no traffic had used. Only a circuit carrying a stream now confirms an
exit, an inferred one can never displace a confirmed one, andexit_confirmedsays which
it is. - A session is no longer routed to an instance that is mid-rotation. Diverting covered
the sessions pinned when the rotation began, but not the ones arriving during it. - A stalled bootstrap is now remediated. Tor can wedge part-way through with a live
process, which neither the supervisor nor the failure ladder catches, leaving the pool
quietly under strength — instances were observed sitting at 45% indefinitely. See
BOOTSTRAP_STALL_TIMEOUT. - The maintenance loop cannot be stalled by a control port. The exit poll shared a loop
with session sweeping and process supervision, and one instance's NEWNYM cooldown blocked
all three for up to ten seconds — a pool-wide rotation, for tens of seconds. Control
commands also had no I/O deadline, so a Tor that stopped answering wedged it forever. - HTTP proxy: keep-alive requests are routed individually. A client sending requests
for several hosts down one proxy connection had the second delivered to the first host.
Each request is now routed and dialled on its own, which also means a rotation takes
effect on the next plain request rather than when the client happens to reconnect. - HTTP proxy: IPv6 destinations work. A bracketed literal was passed to Tor as a
hostname to resolve. - A control connection lost while Tor keeps running is redialled, instead of leaving an
instance that serves traffic but can never be rotated or report its exit again. - Rotating an instance that has not bootstrapped is refused with 409 rather than spending
the NEWNYM cooldown on a Tor with no circuits — which silently swallowed the rotation
asked for once it was ready. - Rotating a session that lands back on its own instance (a one-instance pool, or one
instance routable) now rotates that instance's circuit, instead of reporting success
while changing nothing. - Instance indexes are reused instead of counted upwards, so enough resizes can no longer
hand an instance a SOCKS port that is another instance's control port. - Remediation backoff grows with the attempts at the current rung, not with the instance's
lifetime count — an instance that misbehaved last week no longer starts at maximum
backoff. - Retired instances no longer leave their per-instance counters behind, a resize honours
SPAWN_STAGGER,POST /api/instances/{id}/drainanswers 404 for an instance that does
not exist, and?newnym=1is accepted alongside?newnym=true. - Fixed data races on an instance's process handle during a restart, and on the NEWNYM
cooldown timestamp.