Skip to content

Playback

NihilDigit edited this page Sep 30, 2026 · 4 revisions

Playback

Reading a PikPak file at any offset, fast enough for a player, over links that expire, file objects that vanish and edge hosts that stop answering.

The stack

PikPakStreamReader   a read position with a read-ahead window              (one per consumer)
        │
PikPakFileCache      256 KiB blocks, workers, prefetches, downloads, store  (one per file)
        │
PikPakFileHandle     gcid identity, link minting, rebuilds, leases          (one per file)
        │
RangeReader          one signed link: gates, retries, host steering         (replaced on expiry)
        │
streamRangeFromUrl   one HTTP range request on the CDN client

Everything above RangeReader talks to the layer below through RangeSource: one remote file, readable at any offset, with a priority on each read.

interface RangeSource {
    suspend fun <T> read(start: Long, length: Long, priority: Int = 0, block: suspend (ByteReadChannel) -> T): T
    suspend fun readBytes(start: Long, length: Long, priority: Int = 0): ByteArray   // short at EOF, like a file read
}

The file cache and downloadTo take a RangeSource rather than a RangeReader because a handle retires its reader — on expiry and on rebuild — and anything holding the instance it was handed once would keep reading through one the handle has already replaced. Taking the interface means every read asks again.

Priority decides who takes the next free slot, never who keeps one. A slot is not taken back from a read already running, so the length of low-priority reads is what bounds how long a high-priority read waits. Measured on a saturated 50 Mbit/s line across eight connections: 1 MiB background requests made playback wait up to three seconds; halving them brought the worst wait under one.

Quick use

val handle = client.fileHandle(detail, mediaId = null)       // or PikPakFileHandle(client, gcid, size, name, initialFileId = id)
val cache = handle.openCache()                                // one per file; see below
cache.openStream().use { stream ->
    stream.seekTo(position)
    val n = stream.read(buffer, 0, buffer.size)               // at most one block per call; -1 at EOF
}
cache.close()
handle.close()

read and seekTo suspend. close does not, so a player can release its input from a lifecycle callback with no coroutine, and closing a stream from another thread is a supported way to abort a read parked on it.

RangeReader: one signed link

  • Two gates. Every read takes a slot of the file's gate (8, the per-link cap) and then one of the account's (16). Per file first: that bounds one file to its own budget of account slots, so a busy file cannot starve another. The file's gate belongs to the handle, not to the reader, because reads on a retired reader keep running; a replacement with its own gate would let one link carry up to twice its cap until they drain.
  • Priority, then age, then arrival. A contended slot goes to the higher priority; among equals, to the older demand — a RequestOrder carried on the coroutine context, set by the stream reader — and only then to the earlier arrival. See Prefetching for why.
  • Expiry (401/403) asks the URL provider for a fresh link; concurrent reads share one refresh.
  • 503 is the per-link cap: waited out with growing pauses, up to 30 times, without counting as a failure.
  • A truncated body resumes from the offset already delivered. The channel the caller reads is stitched across however many requests it took.
  • A silent host. A request that gets no response headers within 2 s is abandoned for a fresh link, which PikPak signs for a host picked anew. The clock restarts whenever the HTTP layer receives and retries a 5xx or 429 — that host answered — so only complete silence counts. A body that goes silent for 3 s once bytes were flowing is treated as a dropped connection and resumed.
  • Host health, account-wide. A silent host is recorded on the client for ten minutes, but only while some other host has answered in the last ten seconds: a dropped Wi-Fi silences every request at once, and recording those would blacklist the whole pool. A link any handle mints onto a recorded host is minted again, up to twice.
  • Host steering (steerEdgeHosts, on by default). A slow host answers promptly, so the silence check never fires, and re-minting keeps landing on the same few hosts (Measurements). But a link is not bound to its host: the same path and signature work on any host of its family (dl-z01a-* with dl-z01a-*, same root domain). So each attempt may go to a sibling host instead of the link's own:
    • Speed is recorded per host, account-wide, from every attempt of every file that delivered at least 256 KiB, as the rate from its first byte to its last. A host's figure is the lower median of its recent samples (at least two, within ten minutes).
    • A read moves when the link's host is at least 4× slower than a sibling, or silent, and only to a sibling that has itself delivered on this account — the fastest one.
    • Background and read-ahead reads explore: at most one attempt every 5 s account-wide goes to a sibling with no figures yet, until three hosts of the family have them. Without this a single file never compares its host with anything. Reads a stream is stopped on — STREAMING_PRIORITY and above — never explore, since the sibling may be a slow one.
    • Candidates are hosts this account's links have named, then a built-in list of hosts seen in 2026. A listed host is only a candidate: it is used as a destination only after it has delivered.
    • A sibling that fails a rerouted attempt in any way — refused handshake, 403, 404, silence, an empty body — is left alone for ten minutes, and the read goes straight back to the link's own host. That failure is not counted against the read and does not refresh the link: it says nothing about the link.
    • Only this account's hosts are compared with each other, so a slow line or a slow consumer slows every host alike and moves nothing.
    • The link itself is never rewritten: expiry, refresh and rejection still concern the link, and RangeAttempt.host names where each attempt actually went. steerEdgeHosts = false keeps every attempt on the link's own host; silent-host avoidance stays.

Every HTTP attempt can be observed through onRangeAttempt (on the handle, or onAttempt on a bare reader) as a RangeAttempt: offset, bytes, time to headers, time to first byte, bytes per 200 ms, the retries the HTTP layer made underneath, the host, and how it ended — Complete, Expired, Throttled, Rejected, Failed, or Cancelled when the reader stopped it because nobody wanted the rest. Most attempts a player makes end Cancelled; that is not a fault.

PikPakFileHandle: the file, not the file object

PikPak stores content by hash. The handle holds the gcid as the file's identity and treats the file id as a cache of it:

  • Expiry → read the detail again, take the fresh link. One request, about 180 ms.
  • A rejection, or a detail that answers 404 → the file object is gone. Create a new one from the gcid (instantCreate) and continue. Two requests, about 380 ms. No caller help is needed, and offsets already read stay valid: it is the same bytes.

Details the handle uses:

  • Links are reused. A reader's first link is the one already in hand while it is usable — one passed in from a detail the caller just fetched (client.fileHandle(detail)), or one the size probe minted a moment ago. Without it a transcode cost three detail lookups before its first byte.
  • The variant is fixed for the handle's life (mediaId, null for the original). Switching mid-file would change the bytes under offsets already read.
  • A transcode's length is not in the metadata and costs a one-byte range probe; it is kept per gcid and variant on the client, and can be passed in (streamSize).
  • Objects it is done with. onObjectMinted reports every file object the handle is done with — the initial one once a link has been minted from it, and every one a rebuild replaced — so a caller that creates objects only to mint links can delete them. A report still owed when the handle closes — a rebuild whose detail lookup then failed — is made in the client's background by close(); a client closed first leaves that object for the caller's sweep.
  • Leases (leased = true) do that deletion inside the SDK; see Leases.
  • contentKey names what the handle reads — gcid/mediaId, gcid/origin for the original — and is the key a BlockStore files its blocks under.
  • Two rebuilds may race and orphan one file object. Serialising them would hold a lock across two round trips on the path that is already slow; the orphan is reported, not dropped.

Leases

A signed link outlives the file object that minted it, permanent deletion included (Magnets and Instant Create), so content can be read with no object left in the drive:

val detail = client.leaseDetail(resolvedFile, parentId = leaseFolder)     // create, read the detail, delete in the background
val handle = client.fileHandle(detail, leased = true)                      // rebuilds are leased too
client.leaseBudget = LeaseBudget(capacityBytes = freeBytes)                // once the free space is known
  • An object takes its full size of storage from its create until its delete lands, and 15 % of it from the monthly upload allowance. LeaseBudget bounds the bytes held at once: a lease takes its size before the create and gives it back once the object is gone, so on a small account later leases wait for earlier deletes instead of failing. Waiters are served in order, and a file larger than the whole budget still goes through, alone.
  • A leased handle deletes every object it creates on a rebuild, in parentId, and reports it through onObjectMinted too.
  • Deletes run in the client's background. One still pending when the client closes is lost; the object waits in parentId for the caller's sweep.
  • Open question. A rebuild needs PikPak to still accept the gcid after every object holding it was deleted. rclone's tracker reports the server's record of a deleted file's hash changing within the hour; it has not been measured here. LeaseRebuildProbeTest measures it (Development).

The file cache

PikPakFileCache holds the bytes of one file for everything reading it: fixed 256 KiB blocks keyed by offset, least recently used evicted first, capped at 64 MiB, filled by up to 8 workers. handle.openCache(blockStore, coroutineContext) builds one reading through the handle; client.fileCache(source, size, storeKey, ...) builds one over any RangeSource. The cache owns no link and the handle owns no bytes: until 2.0.0 the handle held the cache, and with it every stream.

  • A stream is a cursor: a position, a read-ahead window (32 MiB by default, readAheadLimit to lower it), and a role. Several can read at once. A player that opens a second connection to seek opens a second stream; an interleaved MP4 read in two places at once keeps both windows filling. Bytes one stream fetched are there for the other.
  • Nothing is fetched until a stream reads. A stream opened only to prefetch would otherwise pull a whole read-ahead from offset zero.
  • Blocks, not a sliding window, because mpv opens a Matroska file by reading the head, jumping to the tail for the Cues and jumping back; a window would discard one end each time.
  • Wide fetches. From a start or a seek a fetch is one block, so a cold open spreads across all connections. Once 8 MiB ahead of the position is cached or on its way, fetches take two blocks, halving the per-request latency of a settled stream.
  • Cancellation. A fetch no stream's window and no prefetch covers any more — its stream moved or closed — is cancelled. What the stream needs back is the connection.
  • Failure is local. A block that fails three times (each attempt already carrying RangeReader's own retries) fails the reads and prefetches waiting on it, and nothing else. The next read of it tries once more before reporting. The stream stays usable.
  • Memory. Blocks inside any window are not evicted, which is why the cap must clear one window plus every block the workers can hold. Several streams can together cover more than the cap; then the block a stream is stopped on may evict the one needed last — the cached block farthest ahead of its stream.

wastedBytes counts blocks evicted before anyone read them. A read that waits says read-ahead was too shallow; this is the only signal that it was too deep. deliveredBytes counts what the CDN handed over for every stream, prefetch and download on the cache: one number for the file.

Priorities

Priority Constant What asks for it
101 INDEX_PRIORITY A container index the demuxer needs before any frame: Matroska Cues, a trailing moov
100 BLOCKING_PRIORITY The block an urgent foreground stream is stopped on
50 STREAMING_PRIORITY The block any other foreground stream is stopped on
10 READ_AHEAD_PRIORITY Foreground read-ahead
8 WARM_PRIORITY A foreground prefetch: the file the user will see next
5 BACKGROUND_BLOCKING_PRIORITY The block a background stream is stopped on
1 BACKGROUND_READ_AHEAD_PRIORITY Background read-ahead and background prefetch

INDEX_PRIORITY sits above blocking on purpose: until the index arrives there is no first frame, so nothing is really blocked yet.

Urgent or streaming. A stream cannot tell a seek from a buffer fill: a player fills its buffer with blocking reads while seconds of video are still in hand, and a seek arrives as just another stream opened at an offset. Until 1.2.0 both went out at BLOCKING_PRIORITY, and the fills crowded out the seek — in a feed with three players they were nearly half of all requests, 88 to 94 in every 200, and a seek, as the newest demand, lost every tie to them. So the caller says: PikPakStreamReader.urgent is set while someone is actually waiting — a seek that has not shown its frame, a stall, a first frame the user has scrolled to — and cleared when the wait is over. Everything else a stream is stopped on goes at STREAMING_PRIORITY, still above every read-ahead and prefetch. With it set only then, the top band carried one to three requests per ten seconds, and seeks in Piko's feed resumed in 250 to 650 ms (desktop and a phone, transcoded clips). Requests already queued keep the priority they were issued at.

Roles

A stream is FOREGROUND (what the user is watching) or BACKGROUND (warmed for later). Changing the role keeps the cache. While the account has any foreground demand, a file with none of its own keeps at most two requests in flight, each a single block: priority cannot take back a slot a background read holds, so bounding the background reads is what bounds playback's wait.

A foreground stream counts as foreground demand while a read is parked on it and for 5 s after its last read or seek. A paused player, or one sitting on a full buffer, stops reading; before 2.0.0 it held every background file on the account at two requests for as long as its stream stayed open. When it reads again it counts again at once, and what the background holds by then is at most two single blocks. A foreground prefetch counts while it runs.

Prefetching

val job = cache.prefetch(listOf(0L until headBytes), StreamRole.FOREGROUND)    // no stream needed
stream.prefetch(listOf(size - 512 * 1024 until size), PikPakStreamReader.INDEX_PRIORITY)
job.await()      // completes when cached, fails if a block cannot be fetched
job.cancel()     // withdraws the request; fetches only it wanted are cancelled

A prefetch fetches ranges without a read position, so a later seek does not cancel it. Ties at the gates go to the older demand: files prefetched in the order they will be played finish in that order. Before this, a worker queued one block at a time and the gates served equal priorities by arrival, so eight files prefetched together took turns block by block and all finished late; the only fix a caller had was to prefetch two at a time itself.

Keeping blocks on disk

val cache = handle.openCache(blockStore = myStore)

interface BlockStore {
    suspend fun read(file: String, offset: Long, length: Int): ByteArray?
    suspend fun write(file: String, offset: Long, bytes: ByteArray)
}

The cache asks the store before the network and offers it every block the network delivered. What to keep and for how long is the store's decision — the SDK still writes nothing to disk on its own. file names content, not a file object (the handle's contentKey), so a store hit survives a new file object and a new link. Reads run on the workers, so a slow store slows the stream; writes go through a queue of 16 blocks and are dropped when it is full, so a slow disk misses offers instead of filling memory. A store that throws is treated as holding nothing.

Downloading through the cache

interface DurableBlockStore : BlockStore {
    suspend fun missing(file: String, ranges: List<LongRange>): List<LongRange>
}

val cache = handle.openCache(blockStore = myDurableStore)
cache.download(listOf(0L until head, size - tail until size, head until size - tail)).await()

A download is a prefetch whose blocks must end up in the store. It is how to keep a file that is also played: a block a stream already has in memory is written from there, a block the download wrote is what a stream reads next, and nothing is fetched twice. downloadTo below cannot do that.

  • What the store holds is skipped. missing is asked once per call, so calling download again after a failure, a cancellation or a restart continues where the store left off.
  • Order. Blocks are fetched in the order the ranges give them — a container's head and index first, then the middle.
  • Written, then done. Each block's write is awaited on the worker that fetched it, so a disk that cannot keep up slows the download instead of filling memory, and the job completes only once every block is held. Blocks only a download wants are not kept in memory, so a whole film going to disk does not churn what the streams read from.
  • Failure. A block that cannot be fetched (after the retries every read carries) or cannot be written fails the job, and only the job: readers keep the bytes. Retrying is the caller's — call again.
  • Priority. A background download by default. While anything on the account is watched, the file being downloaded included, a download keeps at most two requests in flight, each a single block; otherwise it fetches two blocks per request, like a settled stream. A download never counts as foreground demand itself.

The store's side of the contract is stricter than a BlockStore's: write returning means the block is held, and a block that cannot be kept makes it throw; read returns only held blocks, so a store that preallocates its file must not hand back the zeros of a block never written; a block reported held must survive a crash, which for a file with a bitmap means fsync before the bitmap. Writes arrive at 256 KiB boundaries (the last block shorter), from several coroutines at once.

Transcoded variants

A video exists as the uploaded original and, once PikPak has transcoded it, as MPEG-TS variants at 1080P, 720P and 480P — separate resources with their own links and lengths.

val v = detail.resolveVariant(VariantPreference.Resolution("720P"))   // falls back to the original
val handle = client.fileHandle(detail, mediaId = v.mediaId)
handle.variants()                                                     // what PikPak holds for this content

resolveVariant is the only place a missing or still-transcoding variant falls back to the original. After that the choice is locked: a refresh keeps the same mediaId and fails when that variant is gone, rather than serving another variant's bytes at an offset the caller has committed to. To reopen later, persist the gcid and the mediaId, not the file id.

Fact. Transcodes belong to the content, not the file object: an instant-created copy of an already-transcoded gcid arrives with every variant present. Nothing schedules a transcode — one with none still had only the original after two minutes and after its bytes were read (2026-09-11) — so a caller finding only the original should read it rather than wait. Transcodes carry no embedded subtitles and a lower audio bitrate, and a transcode is MPEG-TS: its length is a multiple of 188, and any 188-aligned slice of it plays from its next keyframe.

Downloading to disk

client.fileHandle(detail).use { it.downloadTo(dest, totalSize = detail.sizeBytes) }

For the plain file whose length is its progress. A file that is also played goes through the cache instead.

Blocks are fetched over concurrency connections (4 by default) and appended strictly in order, so the file is at every moment a valid prefix and its length is the progress: an interrupted download resumes from what is on disk and needs no bitmap. Only the writes are ordered — the fetches slide, so a slow block delays the write, not the next request. Cancelling is pausing. concurrency = 1 is a single-connection download.

downloadTo reads through RangeSource and nothing else, so it shares no blocks with a file cache or its store: a file downloaded this way while it plays is fetched twice.

A download ceiling

val limiter = BandwidthLimiter(bytesPerSecond = 512L * 1024)   // hold one for the process
handle.downloadTo(dest, totalSize, limiter = limiter)
limiter.bytesPerSecond = 2L * 1024 * 1024                      // applies at once
limiter.bytesPerSecond = null                                  // no ceiling

Every download given the same limiter, and every connection of each, shares one ceiling. Only downloadTo takes it: streams, prefetches, cache downloads and size probes never do, so a limited download cannot slow what is playing.

  • Paid per block, before the request. A block waits for its budget with no connection open, then arrives at the line's speed. Throttling the body instead would hold connections open and nearly idle, which the CDN's socket timeout and the reader's own 3 s body-silence check both read as a dead connection. Over a few blocks the rate holds; within one block it does not. At a low ceiling a smaller blockSize makes the flow smoother.
  • A token bucket that may go into debt. A block is granted once the balance is not negative and takes its whole size, so a block larger than a second's budget still goes, and delays whoever comes next. Idle time earns at most one second of credit.
  • In order. Waiters are served in the order they asked, which is the order downloadTo issues blocks, so the block the file is waiting to append is never overtaken.
  • Changes apply at once, to a download already waiting too: it recomputes its wait against the new rate. null releases every waiter.
  • A block that fails and is fetched again is paid for again, so a flaky line stays under the ceiling rather than on it.

acquire(bytes) is public, so the same class can pace a caller's own transfers, such as an upload source.

A bare RangeReader

val reader = client.rangeReader(fileId, mediaId = null)     // refreshes through getFile(fileId)
reader.readBytes(0, 64 * 1024)

For short-lived reads by file id. It does not survive the file id dying; a handle does.

Clone this wiki locally