Skip to content

History

NihilDigit edited this page Sep 30, 2026 · 4 revisions

History

What was removed or rewritten, and why. The reasons are the point: each of these is the kind of thing that comes back if nobody remembers why it went.

2.0.0: the cache leaves the handle

Breaking, on purpose, with its only two consumers (Piko and Animeko) moving in step.

  • PikPakFileCache. The block cache used to belong to PikPakFileHandle, which made the handle carry the file's identity, its links and its bytes at once (Architecture, review item 4), and every new thing done with the bytes had to go through the object minting links. Now handle.openCache(blockStore, coroutineContext) or client.fileCache(source, size, storeKey, ...) builds the cache, and streams, prefetches and downloads are opened on it.
  • Downloads through the cache. PikPakFileCache.download(ranges) into a DurableBlockStore: the store's held blocks are skipped, blocks are fetched in the order asked and written before the job completes, and blocks a stream has in memory are written from there. Built for Animeko, whose cache engine fetched each block a second time when an episode was downloaded while it played — through its own piece fetcher, beside the player's reader — and on a free account's 20 GiB of downstream a day that doubles what an episode costs. The store stays the caller's: the SDK still writes nothing to disk itself, and a sparse file with a bitmap is a few hundred lines any caller can own. See Playback.
  • A download on a watched account keeps to two single-block requests, the file being played included. The file's own stream would otherwise queue behind eight download requests it cannot take back.
  • A foreground stream counts as foreground for 5 s after its last read, not for as long as it is open. A paused player held every background file on the account at two requests; Animeko worked around it by closing its cloud stream after 5 s idle and reopening it on the next read.
  • close() makes the report a handle still owes, in the client's background. closeAndReport() existed because close() could not suspend, and a caller that used the wrong one stranded a leased object; Animeko's engine carried a retire path just to call the right one.
  • contentKey is public: the key a store files a handle's blocks under, which a caller now passes to fileCache itself.
Removed Why Instead
PikPakFileHandle.openStream, .prefetch The cache left the handle handle.openCache(), then openStream / prefetch on it
blockStore, coroutineContext on PikPakFileHandle and fileHandle Configured the cache openCache(blockStore, coroutineContext)
PikPakFileHandle.closeAndReport close() makes the report close()
PikPakStreamReader(source, size, ...) public constructor A private cache per reader is the shape 1.0.0 moved away from; the one remaining use was a reader over a bare source client.fileCache(source, size).openStream()

1.3.0: leases, account tier

Added, not breaking: leaseDetail, fileHandle(leased = true) and LeaseBudget (Playback); the daily cloud-download count and the account tier on getQuota and getTransferQuota; gcidByCid sampling helpers and a gcid's thumbnail URL. The lease mode an earlier version had dropped came back without the rebuild risk being measured; see Magnets and Instant Create.

1.2.0: urgent reads

Added, not breaking:

  • PikPakStreamReader.urgent and STREAMING_PRIORITY (50). A stream's blocked read used to go out at BLOCKING_PRIORITY whatever it was for, and a player filling its buffer blocks on every read; those fills filled the top band a seek needed. Now a blocked read is at the top only while the caller marks the stream urgent, and at 50 otherwise. See Playback.

Changed:

  • A read the stream is stopped on never explores an unmeasured host, at either band. The line used to be BLOCKING_PRIORITY; a buffer fill at 50 is still a read the player stops on if it is late.

Considered and not done:

  • Raising the priority of requests already queued when a stream turns urgent. It needs the gates to find and reorder waiters by demand. With the top band otherwise nearly empty, a request issued after the flag is set wins the next slot anyway; revisit if a stall is measured waiting on a request queued before it.
  • Deciding urgency inside the SDK, from a seek or a newly opened stream. Behind a loopback proxy a seek and a preloading player opening a file both arrive as a new stream at an offset; only the caller knows which one someone is looking at.

1.1.0: edge hosts, root domains, a download ceiling

Added, none of it breaking; every default keeps 1.0.0's behaviour except host steering, which only acts once it has measured a reason to:

  • Edge-host steering (steerEdgeHosts, on by default). A signed link is not bound to its host: the same path and signature work on any host of the link's family. The reader records each host's speed account-wide, sends reads off a host measured at least four times slower than a sibling, and lets background reads try an unmeasured sibling now and then so there is something to compare with. See Playback and Measurements.
  • Root domains (PikPakDomain, PikPakClient.domain, a var). The API is reachable under four roots with the same tokens; the root can be switched while the client runs.
  • probeDomain: one unauthenticated round of requests to a root, returning whether PikPak's gateway is behind it and how long it took. See Getting Started.
  • BandwidthLimiter, and limiter on downloadTo: a ceiling several downloads share, changeable while they run, paid per block before the block is requested. Playback never takes it.

Considered and not done:

  • Pinning API addresses. A community list of addresses for api-drive and user was measured: the best was about 25 ms faster than what DNS returned, inside the noise; one address with a valid certificate served another service, and two did not answer. On a machine behind a fake-ip proxy a pinned address is either replaced by the proxy, which re-dials by SNI, or bypasses the user's routing rules. No gain and a way to break proxied setups.
  • Several hosts serving one range at once. Eight connections on one healthy host already fill the line (8.8 MB/s through the proxy measured); what was left to gain was avoiding a slow host, which steering does. The maintainer ruled it out for that reason.
  • Choosing a root automatically. No speed difference between roots was measured, so the SDK has nothing to choose on; the caller probes and decides.
  • The region check. access.<root> is a gate in PikPak's own clients, not in the API. The SDK never called it and does not; see Measurements.

1.0.0: one cache, many readers

The block cache, the workers and the read position used to be one object, PikPakStreamReader, so a cache had exactly one position. Both consumers built the same four things around that, independently and with the same constants: a tail prefetch at 512 KiB and BLOCKING + 1 that bypassed the reader because seekTo cancelled the head fetches; a scheduler admitting two files at a time; a disk layer capturing bytes by reading them through a cursor; and a reader serving disk, tail and network in turn. Piko's proxy also had to cancel one HTTP request before serving the next, so an interleaved MP4 cancelled its own audio and video reads in turn. Now the cache is BlockCache, one per handle, and a stream is a cursor on it. The public constructor still builds a private cache, so a reader over a bare RangeSource behaves as before.

Also in 1.0.0:

  • A dead block no longer retires the reader. It fails the reads waiting on it; the next read tries again. Retiring the reader took the cache down with it, and the only recovery was a new reader that fetched everything again.
  • Ties go to the older demand. Files prefetched together used to take turns block by block and finish late together. See Playback.
  • Hosts. A 2 s first-response deadline moves a read off a silent host; silent hosts are remembered account-wide, but only when other hosts are answering. The engine's per-host cap was raised to the account budget: it had been 8, and two files on one host queued out of sight past the deadline and were recorded as dead.
  • Links in hand are reused, and a handle can be built from a detail (fileHandle). A transcode used to cost three detail lookups before its first byte.
  • Authentication became one ladder. login() ignored an expired in-memory session's refresh token; the 401 path never read the store, so a request made before login() went from a missing header to a password sign-in while a valid session sat on disk. A failed sign-in after a dead refresh token was attempted twice, asking the password supplier twice. The captcha refresh was keyed on a snapshot but the headers were read live, so the refresh could skip itself.
  • POSTs are no longer replayed blind. A lost answer to a completed multipart upload was retried, found the upload gone, and the cleanup then deleted the finished file.
  • Breaking, and why now. OfflineTask became DriveTask: it has always been the record for restore, extraction, pack and copy tasks as well, and the offline name sent readers looking for a second type. A deprecated typealias keeps old code compiling. Every gcid the SDK hands out is upper case: the server lists upper case, but PikPakHash and gcidByCid returned lower case, so a hash computed locally never == the listed one. Stored lower-case gcids still work in requests — the server accepts either case — but compare them after uppercase().
  • Smaller fixes. batchMove did not split long id lists like the other batch calls; the folder memo survived a batch that failed halfway and could hand out a trashed folder; a single-file offline prune matched by name and deleted a file PikPak had renamed "name(1)"; a magnet tree deeper than eight folders lost files silently; readers built from a handle threw at EOF where the interface promised a short read; a cancelled rate-limiter waiter kept its reservation; Android's default session path was unwritable and failed every login after a successful sign-in.

Removed in 1.0.0

Removed Why Instead
CdnConnectionStats, the OkHttp event listeners No reader anywhere RangeAttempt.timeToHeaders separates handshake from reuse
PikPakClient.httpRetries, HttpRetryStats No consumer; superseded Retries per request on RangeAttempt
downloadSingleConnection, downloadSingleConnectionFromUrl No consumer; a 5xx cost the square of the retry budget; a finished file of unknown size was re-downloaded downloadTo(concurrency = 1)
downloadFromUrl No consumer; a fixed link that cannot refresh or change host fileHandle(detail).downloadTo(...)
rangeReader(fileId) without mediaId Duplicated the other overload rangeReader(fileId, mediaId = null)
PikPakFileHandle.readToEnd, public readBytes overrides No caller; the overrides broke the EOF contract RangeSource.readBytes
PikPakStreamReader.MAX_ATTEMPTS, public lastSeekLatency Test hooks —
openStream(size, concurrency, parentCoroutineContext) parameters They configured a cache shared by every stream, or were silently ignored Set them on the handle
instantCreateResolvable No caller instantCreate, then read the detail you were going to read
retryOfflineTask Did not work: every retried task failed again createUrlFile(parent, task.params["url"])
restoredFileIds The task has no id mapping List the destination
getVipInfo, VipStatus No caller; duplicated getTransferQuota getTransferQuota
searchFilesRecursiveList A one-line .toList() searchFilesRecursive(...).toList()
emptyTrash No caller; request shape never verified; irreversible —
invalidateFolderId, RateLimiter.Default / .Unlimited No caller; deprecated clearFolderIdCache, RateLimiter.default() / .unlimited()
compileOnly ContentNegotiation and serialization Never used —

The stream reader before 1.0

A read-ahead that never reached its depth. Read-ahead and the memory cap both defaulted to 64 MiB, and eviction exempts every block inside the read-ahead window. Equal values meant the window was the whole cache: nothing could be evicted, no block could be claimed, and read-ahead collapsed to "the reader advances one block, the scheduler refills one". The configured depth never existed in steady state. The fix was the sizing, not the exemption — every block inside the window is one the same scheduler asked for, so evicting it only makes the next claim fetch it again. Read-ahead became 32 MiB against a 64 MiB cap, with a check that the cap leaves room for the window plus every block the workers can hold; the wide-block threshold went from 32 MiB to 8 MiB in the same change, since at 32 it equalled the window and doubling could only start once read-ahead was complete. It hid because the test asserted the cache stayed under its cap, which a stalled window satisfies as well as a working one; the replacement asserts the depth reached.

runBlocking out of common code. The surface used to bridge into blocking calls in two places, which came with a rule stated only in prose: never call read or seekTo from the dispatcher the reader ran on, or deadlock. read, seekTo and the internals are suspending now. close stayed non-suspending on purpose, for lifecycle callbacks; it works because each fetch's completion is parented to the reader's job, so cancelling the scope wakes every parked read — without that, close would have to walk the in-flight table under the lock the workers drain under, and would have had to suspend.

0.6.0: the download surface

Three whole-file download modes collapsed into one. parallelDownloadFromUrl was deleted: once downloadTo slid its window, all it still had was streaming parts to temporary files instead of holding a window in memory — 4 MiB at the defaults — and it paid for that with no resume and no link refresh. The names changed hands at the same time so that the name a reader reaches for without the docs is the right one; the single-connection pair was renamed to say what it was (and removed in 1.0.0).

downloadTo used to fetch concurrency blocks, wait for all of them, write the batch, then issue the next. Every round ended with the connections idle behind its slowest block, and that cost most on the weak links fan-out is for, where block latency varies most. It is a sliding window now: writes stay in order, but the block at the head is written the moment it lands and its replacement goes out at once. The window still counts finished-but-unwritten blocks against its size, so a slow head leaves the connections behind it idle until it lands; removing that means buffering beyond concurrency blocks, memory for throughput, not taken unmeasured. The regression test puts the slow block last in the first window, where the two models differ; reinstating the barrier fails it.

Fact. kotlinx-io 0.9.0 has no seekable file handle — source, sink(append), metadataOrNull, delete, atomicMove, no truncate. It constrained nothing: writing at N offsets would leave holes, and a file with holes needs a bitmap, where an in-order append keeps length equal to progress and resume free.

Moved back to Animeko

PikPakFileDownloader — start/stop, progress and error as StateFlow, retry forever — was Animeko's SequentialDownloader almost verbatim, and the one piece that introduced a concept PikPak does not have: a background task with a lifecycle. It depended only on public API, so it went back as a file move. Two lessons it carried:

  • Retry must live in exactly one layer. The first version retried in its outer loop while downloadTo retried inside with its own backoff, so error stayed null while the inner loop kept trying and a caller saw a download that merely stopped moving.
  • start/stop must not hand a Job between them. The first version assigned it inside a launched coroutine, so a stop() a microsecond later cancelled nothing. A MutableStateFlow<Boolean> collected with collectLatest cannot race.

It also carried out a known bug: error was cleared only when the whole file finished, so a download that failed, recovered and failed again never showed null in between.

indexMagnet, removed

It submitted a magnet, polled the offline task to a terminal phase and walked the result, as a Flow of progress — because a magnet nobody has uploaded takes minutes or fails. The gcid path (Magnets and Instant Create) removed the wait it narrated, and it went with pollUntilTerminal, MagnetIndexProgress, OfflineTaskFailedException, awaitIndex and readMagnetIndex. Two problems it had: submission was unconditional, so a second call produced a second pack (a probe's donor file was named archlinux-…(6).iso, the seventh copy); and CreateUrlResult.InstantComplete, meant for "PikPak already had this", could not be reproduced — an already-held magnet submitted twice produced a fresh task both times.

FileRelocator, removed

An early handle made the caller supply "find this content again", because only Animeko knew how — it held the magnet and the path — and onRelocated wrote the new id back. With the gcid in hand the SDK rebuilds by itself. CachingVariantLinkSource went too: it cached links by file id, and the thing worth caching turned out to be the id itself.

Never ported

  • Animeko's DownloadScheduler ("at most two downloads, yield to playback"). The two-level budget with priority replaces it at the right granularity: connections are the scarce resource, not downloaders. Animeko's PikPak engine brought one back beside its own piece fetcher; with downloads in the cache it can go again.
  • Episode matching (TorrentFileLabel) — matching Bangumi episodes to file names is not PikPak's business.
  • Adaptive connection counts. See Measurements.

Clone this wiki locally