Skip to content

Driver Amazon S3

Gustavo Viana edited this page Aug 7, 2026 · 15 revisions

Driver: Amazon S3

Driver for AWS S3 General Purpose Buckets. Ships in the FileHub.AmazonS3 package. A hub instance is scoped to a single bucket; an optional root path narrows visibility to a prefix.

Bucket type — General Purpose only. This driver targets the standard, multi-AZ S3 service. S3 Express One Zone (Directory Buckets — names ending in --x-s3) is not supported and a hub bound to one will fail at construction with Access Denied on the first HEAD. Directory buckets use a different auth flow (CreateSession / s3express:CreateSession), are single-AZ, and lack several features the driver depends on: cross-region CopyObject, object tagging / user metadata, presigned URLs with the same TTL semantics, storage classes, versioning, and GetBucketPolicyStatus. If you need that storage tier, use the AWS SDK directly — it falls outside the FileHub abstraction.

using Amazon.Runtime;
using FileHub.AmazonS3;

using var hub = AmazonS3FileHub.Create(
    AmazonS3HubOptions.FromCredentials(
        bucketName:  "reports",
        credentials: new BasicAWSCredentials(key, secret),
        region:      "us-east-1",
        rootPath:    "archive/2026"));

Always using or register as a singleton — the hub owns the AWSSDK HTTP client by default.

Construction

Two layers:

  • AmazonS3HubOptions.From* — typed factories, one per valid auth strategy. Captures only the params relevant to that strategy; impossible to mix mutually-exclusive fields. Use these by default.
  • new AmazonS3HubOptions { … } — raw object initializer. Escape hatch for unusual combinations.

Both produce an AmazonS3HubOptions passed to AmazonS3FileHub.Create(...) (or CreateAsync(..., ct) under a SynchronizationContext).

Auth strategies

Factory Required Optional Use
AmazonS3HubOptions.FromProfile(bucketName, profile?, region?, rootPath?) bucketName profile (default "default"), region (read from profile when omitted), rootPath Reads ~/.aws/credentials. Hub creates and owns the SDK client.
AmazonS3HubOptions.FromCredentials(bucketName, credentials, region, rootPath?) bucketName, credentials, region rootPath Explicit AWSCredentials (IAM role, instance profile, env vars, STS). Hub creates and owns the SDK client.
AmazonS3HubOptions.FromClient(bucketName, IAmazonS3, region, rootPath?) bucketName, client, region rootPath Reuse an existing IAmazonS3. Caller keeps ownership; hub disposal is a no-op.
AmazonS3HubOptions.FromClient(bucketName, AmazonS3Client, rootPath?) bucketName, client rootPath Same as above but takes an AmazonS3Client directly — region is read from client.Config.RegionEndpoint.

When rootPath is not empty the factory ensures the prefix marker exists on the bucket before returning.

// Profile (~/.aws/credentials)
var hub = await AmazonS3FileHub.CreateAsync(
    AmazonS3HubOptions.FromProfile("reports", profile: "prod", region: "sa-east-1", rootPath: "archive/2026"));

// Explicit credentials
var hub = AmazonS3FileHub.Create(
    AmazonS3HubOptions.FromCredentials("reports", new BasicAWSCredentials(key, secret), "us-east-1"));

// Reuse a pre-configured client (region from AmazonS3Client.Config)
var hub = AmazonS3FileHub.Create(AmazonS3HubOptions.FromClient("reports", existingAmazonS3Client));

SDK configuration (SdkConfig)

When the hub creates its own client (FromProfile / FromCredentials strategies), AmazonS3HubOptions.SdkConfig passes an AmazonS3Config through to it — retry mode, timeouts, proxy, anything the SDK exposes. The hub pins nothing: when SdkConfig is null, plain AWS SDK defaults apply. The resolved Region always takes precedence over a RegionEndpoint set on the config; leave Region null to use the config's. SdkConfig is mutually exclusive with Client — an external client already carries its own configuration.

var hub = await AmazonS3FileHub.CreateAsync(new AmazonS3HubOptions
{
    BucketName = "reports",
    Credentials = new BasicAWSCredentials(key, secret),
    Region = "us-east-1",
    SdkConfig = new AmazonS3Config
    {
        RetryMode = RequestRetryMode.Standard,
        MaxErrorRetry = 4,
        Timeout = TimeSpan.FromSeconds(30),
    },
});

The legacy positional factories on the hub class — AmazonS3FileHub.FromProfile, FromCredentials, FromClient (and their *Async siblings) — were removed in 2.0. Use Create(AmazonS3HubOptions.FromX(...)); the options builder makes the argument order explicit and avoids the silent-swap footgun of the old positional pairs.

Interface

public interface IAmazonS3FileHub : IFileHub { }
public sealed class AmazonS3FileHub : IAmazonS3FileHub, IDisposable { ... }

Bucket → filesystem mapping

S3 is a flat key/value store; the driver overlays a tree on it:

Concept Mapping
Object key Path under the root (/ separators)
Directory Zero-byte marker with content-type application/x-directory and key ending in /
Listing ListObjectsV2 with prefix + delimiter = "/" — objects become files, common prefixes become subdirectories
Sandbox Everything under rootPath/ is visible; names outside it are rejected

Existence semantics — three variants, one name

Because directories are a convention (keys ending in /), an item named foo inside a given prefix can live in three different places. Asking "does foo exist?" is ambiguous, so the driver exposes two specific probes:

Variant Key in bucket Created by
File prefix/foo (no trailing slash) CreateFile("foo") or a writer that PUTs the exact key
Explicit directory prefix/foo/ (zero-byte marker) CreateDirectory("foo")
Implicit directory no marker, but prefix/foo/bar.ext exists Any tool that puts nested keys without creating the parent first (e.g., aws s3 cp on the CLI, some third-party uploaders)
hub.Root.FileExists("foo");        // 1 HEAD on prefix/foo
hub.Root.DirectoryExists("foo");   // 1 LIST(prefix/foo/, limit=1) — covers explicit + implicit
  • FileExists(name) — single HEAD prefix/name.
  • DirectoryExists(name) — single LIST prefix/name/ limit=1. Covers both marker-backed and implicit prefixes in one call; no HEAD-first probe.
  • dir.Exists() and TryOpenDirectory(name) use the same single-LIST probe.

Callers that legitimately want either ("anything called foo here?") should call both — see API → DirectoryEntry.

IsPublic / GetBucketAsync is probed once via GetBucketPolicyStatus and memoized per hub, so URL-related calls don't repeat round-trips per file.

Lazy stubs (zero-call open-or-create)

OpenFile(name, createIfNotExists: true) and OpenDirectory(name, createIfNotExists: true) defer all server calls. They return a handle synthesised entirely client-side — no HEAD, no LIST, no PUT.

File stub

var file = hub.Root.OpenFile("reports/2026/q1.pdf", createIfNotExists: true);
// 0 server calls. file.Length == -1, default timestamps, empty Metadata.

if (file is ILazyLoad lazy)
    Console.WriteLine(lazy.IsLoaded);  // false
  • Length = -1, default timestamps, empty Metadata.
  • ((ILazyLoad)file).IsLoaded == false.
  • Writing materialises the object — 1 PutObject (SetBytes / stream / multipart). On a successful commit, IsLoaded flips to true.
  • Exists() fires 1 HEAD. On hit, the response populates Length / timestamps / Metadata and IsLoaded flips to true. On miss, the stub stays unloaded.
  • Reading bytes from a missing stub silently returns 0 bytes — call Exists() first if uncertain.
  • TryOpenFileAsync / OpenFile(name) (strict) return a loaded file: 1 HEAD up-front, IsLoaded == true.

Directory handle

OpenDirectory(name, createIfNotExists: true) is symmetric — zero server calls. The "directory" is purely a virtual prefix on the bucket; no marker is written until you explicitly ask:

var dir = hub.Root.OpenDirectory("reports/2026", createIfNotExists: true);
// 0 server calls. The prefix is virtual.

dir.CreateFile("q1.pdf").SetBytes(bytes);   // 1 PutObject — writes only the leaf.
hub.Root.CreateDirectory("reports/2026");   // explicit ask → 1 PutObject for the marker.

Same shape as nested-path CreateFile("a/b/c.txt"): intermediate "directories" are never materialised unless you call CreateDirectory on them by name.

Nested paths in CreateFile / OpenFile / TryOpenFile

The file APIs accept nested paths (/ and \ are both valid). .. segments are still rejected with FileHubException.

hub.Root.CreateFile("reports/2026/q1.pdf").SetBytes(bytes);
// Exactly 1 PutObject — no markers for "reports/" or "reports/2026/".

var file = hub.Root.OpenFile("reports/2026/q1.pdf", createIfNotExists: true);
// 0 calls — stub returned. SetBytes/SetText/stream then issues the 1 PutObject.

if (hub.Root.TryOpenFile("reports/2026/q1.pdf", out var existing))
    Console.WriteLine(existing.Length);  // 1 HEAD up-front.

Cost minimization

The driver never makes a request that isn't strictly needed. The rules:

  • OpenFile(name, true) / OpenDirectory(name, true) / nested-path CreateFile("a/b/c.txt") / CreateDirectory("a/b/c") never create intermediate markers. Result: a single PutObject for the leaf; everything else is virtual.
  • FileExists(name): 1 HEAD.
  • DirectoryExists(name) / dir.Exists() / TryOpenDirectory(name): single LIST(prefix, limit=1). Covers marker-backed and implicit prefixes in one call.
  • CopyAllObjects (dir.CopyTo / dir.MoveTo / dir.Rename): no post-loop marker enforcement. If the source had a marker it's copied along; if not, the destination prefix stays implicit.
  • CreateDirectory(name) still does the marker PutObject — caller explicitly asked for an empty visible directory.

So the canonical workflow

hub.Root.CreateFile("reports/2026/q1.pdf").SetBytes(bytes);

costs exactly 1 PutObject end-to-end, regardless of nesting depth.

Behavior matrix

Operation Server calls
OpenFile(name) / OpenFile(name, false) exists 1 HEAD
OpenFile(name) / OpenFile(name, false) missing 1 HEAD (404) → FileNotFoundException
OpenFile(name, createIfNotExists: true) 0 (lazy stub)
OpenDirectory(name, createIfNotExists: true) 0 (lazy handle)
OpenDirectory(name) / strict 1 LIST(limit=1)
TryOpenFile(name) / TryOpenFileAsync(name) 1 HEAD
TryOpenDirectory(name) / TryOpenDirectoryAsync(name) 1 LIST(limit=1)
CreateFile(name) 1 HEAD + 1 PUT (empty); refuses existing → FileAlreadyExistsException
CreateFile(name, overwrite: true) 1 PUT (empty); clobbers
CreateFile("a/b/c.txt") 1 PUT (no markers)
CreateDirectory(name) 1 PUT (marker)
FileExists(name) 1 HEAD
DirectoryExists(name) / dir.Exists() 1 LIST(limit=1)
GetFiles() LIST pages only; entries unloaded (IsLoaded = false)
Stub Exists() hit 1 HEAD; flips IsLoaded = true
Stub SetBytes / write 1 PUT

Nested directory paths

CreateDirectory and TryOpenDirectory accept nested paths ("a/b/c", "a\b\c") and resolve the whole path in a single request:

Operation API cost
CreateDirectory("a/b/c") 1 PUT (leaf marker only — no per-segment marker objects)
TryOpenDirectory("a/b/c") 1 LIST(limit=1) proving anything exists under the prefix

Path-traversal guards (.., ., absolute paths) always apply.

Pagination

S3 paginates through an opaque continuation token. Index offsets force the driver to walk object-by-object — billed per ListObjectsV2 call. Prefer named offsets past the first ~1000 entries:

hub.Root.GetFiles(offset: FileListOffset.FromName(lastSeen), limit: 100);

Full explanation in Usage → Pagination.

URLs

AmazonS3File implements IUrlAccessible:

var file = hub.Root.OpenFile("january.pdf");
if (file is IUrlAccessible url)
{
    var link = url.IsPublic
        ? url.GetPublicUrl()
        : await url.GetSignedUrlAsync(TimeSpan.FromMinutes(15), ct);
    return Redirect(link.ToString());
}
  • IsPublic is probed once via GetBucketPolicyStatus and memoized per hub.
  • GetPublicUrl() returns the virtual-hosted-style URL https://<bucket>.s3.<region>.amazonaws.com/<key>. Throws InvalidOperationException on private buckets.
  • GetSignedUrl(TimeSpan) returns a pre-signed GET URL via AWS SigV4.

Signed upload URLs — ISignedUploadable

AmazonS3Directory implements ISignedUploadable. GetSignedUploadUrl(name, expiresIn, options?) / GetSignedUploadUrlAsync(...) mint a pre-signed PUT URL a remote client can upload straight to — the backend never touches the bytes. The target key does not need to exist beforehand; the first PUT creates it (an existing key is overwritten).

var dir = hub.Root.OpenDirectory("uploads");
if (dir is ISignedUploadable up)
{
    var url = await up.GetSignedUploadUrlAsync("user-123/avatar.png", TimeSpan.FromMinutes(15),
        new FileWriteOptions { ContentType = "image/png" }, ct);
    // Client PUTs bytes to `url` — and MUST send Content-Type: image/png.
}
  • When options are supplied, S3 binds them (Content-Type / Cache-Control / user-metadata) into the pre-signed URL signature. The client's PUT must send matching headers, otherwise S3 rejects the request with a signature mismatch. Omit options for an unconstrained upload URL.
  • This is a cross-provider difference: OCI cannot bind headers to a PAR and instead throws NotSupportedException when header-binding options are passed. See the Security page.

Metadata refresh

AmazonS3File and AmazonS3Directory implement IRefreshable. AmazonS3File additionally implements ILazyLoad so callers can detect handles whose state was never populated from the bucket (lazy stubs from OpenFile(..., createIfNotExists: true) and entries returned by GetFiles — LIST doesn't return per-object metadata). Property getters (Length, CreationTimeUtc, LastWriteTimeUtc) return cached values and never do hidden I/O — call Refresh() / RefreshAsync() explicitly to re-sync. Writes update the cached length/timestamp client-side so the common SetBytes → file.Length flow works without a refresh.

LastWriteTimeUtc reads S3's native LastModified header from the last HEAD/LIST response. Refresh() fires a HEAD and reloads Length, CreationTimeUtc, LastWriteTimeUtc, and the metadata snapshot returned by GetMetadataAsync (ContentType, CacheControl, StorageClass, ServerSideEncryption, Tags) — all from the same HEAD response, no extra round-trips.

var file = hub.Root.OpenFile("report.pdf");
var known = file.Length; // whatever was known at open / last write
await ((IRefreshable)file).RefreshAsync(ct);

Operations reference

This section documents exactly what each operation does — the AWS calls issued, how metadata flows, and every edge case. The "metadata on destination" column cross-references the Metadata section.

S3 has no atomic rename primitive. Every name-changing operation in this driver boils down to copy + delete.

File operations

Operation S3 calls Metadata on destination
CreateFile(name) 1 × HEAD (existence guard) + 1 × PutObject (empty body, no metadata headers) Refuses an existing key — throws FileAlreadyExistsException if prefix/name already exists. Use CreateFile(name, overwrite: true) to clobber. Destination created with bucket defaults (no SC, no SSE, no Content-Type, no user-metadata).
SetBytes(bytes) / SetText(...) / GetWriteStream() (no options) 1 × PutObject (on stream dispose / method return) No metadata headers — bucket defaults apply.
SetBytes(bytes, options) / SetText(..., options) / GetWriteStream(options) 1 × PutObject (on stream dispose / method return) Values from S3WriteOptions applied as PutObject headers (Content-Type, Cache-Control, x-amz-storage-class, x-amz-server-side-encryption, x-amz-meta-*). Unset fields fall back to bucket defaults.
Rename("new.txt") 1 × HEAD (existence guard) + 1 × CopyObject (same bucket, default COPY) + 1 × DeleteObject (old key) Never overwrites — throws FileAlreadyExistsException if the new key exists (CopyObject would otherwise clobber it). Destination inherits source metadata (Content-Type, Cache-Control, SC, SSE, user-metadata) via CopyObject default COPY.
CopyTo(dir, name) — same credentials, any bucket/region 1 × HEAD (existence guard, default) + 1 × CopyObject (default COPY) issued by the destination client Destination inherits source metadata. overwrite defaults to false: a HEAD guards the destination first (throws FileAlreadyExistsException if present) — best-effort, not atomic. Pass overwrite: true to skip the guard and clobber. The destination file returned carries the source's known Length but an empty metadata snapshot — call Refresh() / GetMetadata() on it if you need its state.
CopyTo(dir, name) — different credentials Base fallback: 1 × GetObject + 1 × PutObject on destination (stream copy through client) Destination created via CreateFile + PutObject with no metadata headers → bucket defaults. Source metadata is not forwarded.
MoveTo(dir, name) — same credentials Same as CopyTo + 1 × DeleteObject on source Same as CopyTo.
MoveTo(dir, name) — different credentials Base fallback: stream copy + DeleteObject Same as cross-credential CopyTo + delete.
Delete() 1 × DeleteObject N/A — Length resets to -1. Idempotent-silent: deleting a key that is already gone is a no-op, not an error.

There is no MetadataDirective = REPLACE path: every CopyObject this driver issues uses the default COPY, so copies/renames/moves always inherit the source's metadata. Metadata is set only at write time via S3WriteOptions.

Cross-region note. CopyObject is always issued through the destination's client so the request hits the destination's region endpoint. Required for cross-region routing, harmless for same-region. Copying us-east-1 → eu-west-1 under same credentials is still server-side (no byte transfer through your process).

// Rename in place — 1 CopyObject + 1 DeleteObject.
hub.Root.OpenFile("old.txt").Rename("new.txt");

// Move between directories in the same bucket — same two calls.
srcDir.OpenFile("report.pdf").MoveTo(dstDir, "report.pdf");

// Cross-bucket with same credentials — still server-side CopyObject.
hubA.Root.OpenFile("report.pdf").MoveTo(hubB.Root, "report.pdf");

// Change an object's storage class — re-uploads the bytes with S3WriteOptions.
// (There is no metadata-only update path; see the Metadata section.)
var file = hub.Root.OpenFile("doc.pdf");
var bytes = file.ReadAllBytes();
file.SetBytes(bytes, new S3WriteOptions { StorageClass = "GLACIER" });

Overwrite. CopyTo / MoveTo default to overwrite: false — a HEAD guards the destination first; if the key exists the call throws FileAlreadyExistsException and the source is untouched. Pass overwrite: true to skip the guard and let S3 PutObject/CopyObject replace an existing key with no ceremony. The guard is best-effort — a concurrent writer between the HEAD and the copy can still be clobbered. Rename always guards (never overwrites).

Atomicity. None of the copy+delete paths are atomic. Between the copy and the delete the object exists at both keys; a crash in that window leaves the source behind. Readers should treat the new key as the source of truth once the call returns.

Directory operations

Directories are just a prefix (see bucket → filesystem mapping). Name-changing operations expand to per-object copies + deletes.

Operation S3 calls Metadata on destination objects
CreateDirectory(name) 1 × PutObject (empty body, key ending in /, Content-Type: application/x-directory) Marker object only. No user-metadata applied.
dir.Rename("newName") For each object N under the prefix: 1 × CopyObject. Then N+1 × DeleteObject (all objects under old prefix + old marker). Each object's copy uses CopyObject default (COPY) — destination objects inherit their source's metadata.
dir.MoveTo(parent, "name") — same credentials Same mechanic as Rename, destination prefix under the target parent. Same: per-object COPY (inherit).
dir.MoveTo(parent, "name") — different credentials Recursive: for each child → base fallback stream copy. Then source deleted. No metadata forwarded — all destinations created with bucket defaults.
dir.CopyTo(parent, "name") N × CopyObject (same-credentials) or recursive stream copy (cross-credentials). Same as MoveTo (per-object inherit for same-credentials; defaults for cross-credentials).
dir.Delete() List all objects under prefix + N × DeleteObject. N/A.
dir.Delete(name) 1 × DeleteObject (if file) or recursive delete (if directory). N/A. Idempotent-silent on a missing target; a non-empty child directory throws DirectoryNotEmptyException unless recursive: true.

Directory-level operations forward each child via CopyObject default COPY, so children inherit their own source metadata. If you need to set new metadata across many objects, iterate and rewrite each with S3WriteOptions.

A directory move with N files costs N × CopyObject + N × DeleteObject + 1 × DeleteObject (marker). Plan accordingly for large trees.

Overwrite (directories). dir.CopyTo / dir.MoveTo also default to overwrite: false — if the destination prefix already exists the call throws FileAlreadyExistsException before any copy runs and the source is untouched. Pass overwrite: true to merge into the existing destination (per-object server-side copy + delete). dir.Rename always guards (never overwrites).

Recursive delete cost — avoid for large prefixes

dir.Delete(), dir.Delete(name) on a directory, and the cleanup phase of dir.Rename / dir.MoveTo all walk every object under the prefix and issue one DeleteObject per key. S3 charges per request, and the driver does not batch — there is no native "delete prefix" primitive on S3, and DeleteObjects (the multi-key batch endpoint) is not used by this driver.

Practical implications:

  • Cost grows linearly with object count. A prefix with 100 000 files becomes 100 001 DeleteObject calls + at least 100 ListObjectsV2 pages — easily dollars of API spend per call.
  • Latency grows linearly too. Calls are issued sequentially; expect minutes-to-hours for large prefixes.
  • Throttling. S3 may return SlowDown / 503 when delete velocity is too high. Retry behavior is whatever the SDK client is configured with — AWS SDK defaults, or your AmazonS3HubOptions.SdkConfig / external client settings; once retries are exhausted the operation surfaces the failure (see partial-failure section below).

If you need to wipe a large prefix, prefer one of these out-of-band paths instead of dir.Delete():

  • S3 Lifecycle rule with an Expiration action targeting the prefix — S3 deletes objects in the background at no per-object request cost.
  • Bucket-level DeleteBucket + recreate when the entire bucket is disposable.
  • AWS CLI / SDK script using DeleteObjects (batch up to 1000 keys per call) when you need synchronous behaviour without paying per single delete.

Reach for dir.Delete() only when you know the directory is small (think tens to a few hundred objects) or when the cost is acceptable for the use case (manual ops, occasional cleanup).

Partial-failure reporting on directory delete

When per-object deletes fail mid-walk (granular IAM denial, transient throttle), the driver does not abort on the first error. It collects every failure, finishes deleting what it can, and finally throws an AggregateException carrying every per-object error:

try
{
    dir.Delete();
}
catch (AggregateException ex)
{
    // ex.InnerExceptions — one entry per failed DeleteObject.
    // The directory is partially deleted; remediate the failing keys
    // (fix IAM, wait out the throttle) and retry, or fall back to a
    // lifecycle rule.
}

The same applies to the cleanup phase of dir.Rename and dir.MoveTo — when those wrap a delete failure into PartialMoveException, the underlying AggregateException is exposed via PartialMoveException.InnerException.

Partial move failures

When a delete after a successful copy fails (permission revoked mid-op, transient network, partial directory delete), file.MoveTo and dir.MoveTo throw FileHub.PartialMoveException:

try
{
    file.MoveTo(otherHub.Root, "name.txt");
}
catch (PartialMoveException ex)
{
    // ex.DestinationPath — copy succeeded, file is here.
    // ex.SourcePath      — original still exists, delete manually.
    // ex.InnerException  — the underlying S3 error from DeleteObject
    //                     (AggregateException for directory moves).
}
  • PartialMoveException : FileHubException : IOException, so IOException handlers catch it.
  • FileNotFoundException on the delete step is not treated as partial move — source is already gone, move is effectively complete.
  • Rename (file or directory) does not wrap — if the delete step fails, the raw exception propagates (AggregateException for directories, AmazonS3Exception / IOException for files). Remediation is the same: delete source manually.

Streams

Read stream (GetReadStream / GetReadStreamAsync)

Each read pulls a 10 MiB window via GetObject with a Range header; the stream transparently advances to the next window as you read. Seeking within the stream jumps to the right offset without downloading the full object. No local caching — each Read may hit S3 again.

Write stream (GetWriteStream / GetWriteStreamAsync) — single PUT, spills to multipart

Writes are buffered in memory up to the configured threshold (32 MiB by default). Payloads that stay under it commit as a single PutObject; past it, the stream transparently spills into a multipart upload using configurable parts (64 MiB by default). Configure the hub with AmazonS3HubOptions.Multipart = new MultipartStreamOptions(threshold, partSize), or override one write with S3WriteOptions.Multipart. Implications:

  • Payloads within the threshold (32 MiB by default): single PutObject; Flush is the commit — calling Flush mid-write issues the PutObject and the stream stays open for more writes (each subsequent Flush fires another PutObject that overwrites the object).
  • Payloads past the threshold: multipart under the hood with a bounded part buffer (64 MiB by default). After the spill, Flush is a no-op (S3 cannot append; the object materializes at CompleteMultipartUpload on dispose) and an error during writes fires AbortMultipartUpload so no orphan parts are billed. S3 permits at most 10,000 parts; increase PartSize for objects beyond roughly 625 GiB with the default.
  • Metadata: open the stream with GetWriteStream(options) to apply S3WriteOptions on commit — both paths honour them (PutObject headers, or bound at CreateMultipartUpload). The options live with the stream — an abandoned write stream never affects a later write. Without options, no metadata headers are sent and bucket defaults apply.
  • Preference: set S3WriteOptions.StreamPreference = WriteStreamPreference.Multipart to start multipart on the first written byte (skip the buffering phase — payload known large); Single never spills (whole payload buffers for one PutObject — caller owns the memory cost). Default Auto = the threshold behaviour above. Because the preference lives in the options object, SetBytes, SetText and CopyFromStream honor it too. See API → WriteStreamPreference.

When multipart is selected explicitly or reached automatically:

  • Bounded memory — the local buffer caps at the configured part size (64 MiB by default); data rolls over to UploadPart as soon as it fills.
  • At most 10,000 parts per upload; choose PartSize according to the largest expected object.
  • Flush is a no-op. The auto-rollover inside WriteAsync uploads complete configured parts; Dispose uploads the trailing smaller part and calls CompleteMultipartUpload.
  • Errors during writes call AbortMultipartUpload so no orphan parts are billed.

One stream at a time

Only one stream open per file at a time — a second GetReadStream / GetWriteStream call before the first is disposed throws InvalidOperationException. Mixed read and write are the same restriction — dispose before opening another.

Multipart upload

AmazonS3File supports two multipart paths: the regular write stream when the backend owns the bytes, and the optional signed-part capability when clients upload directly to S3.

Backend streams bytes — regular write stream

Use when the backend has the bytes (server-side generation, long-running import). Opens a write stream that chunks data using the configured part size, uploads each part, and commits on dispose:

var file = hub.Root.CreateFile("dataset.parquet");
var options = new S3WriteOptions
{
    StreamPreference = WriteStreamPreference.Multipart,
};
using var stream = await file.GetWriteStreamAsync(options, ct);
await someLargeSource.CopyToAsync(stream, ct);
// Dispose commits via CompleteMultipartUpload. Any exception aborts.

Bounded memory usage (64 MiB by default) regardless of total size. S3 permits 10,000 parts, so choose a larger configured part size for very large objects. If the stream is disposed with an error, AbortMultipartUpload fires to avoid orphan parts being billed.

To write by name without materializing an empty placeholder first, open a lazy file stub and then request its write stream: directory.OpenFile(name, createIfNotExists: true).GetWriteStreamAsync(options, ct).

Backend hands off via presigned URLs — IMultipartUploadSignable

Use when the backend should not touch the bytes — typical for large web/mobile uploads where the client pushes directly to S3:

var file = hub.Root.CreateFile("big.zip");
if (file is not IMultipartUploadSignable signable) throw new NotSupportedException();

// Describe the upload. Pick one of the factories:
var spec = MultipartUploadSpec.FromPartSize(totalBytes: 1_000_000_000, partSize: 10_000_000);
// or:  MultipartUploadSpec.FromPartCount(totalBytes: 1_000_000_000, partCount: 100);

// 1. Backend initiates, returns session.Parts to the client.
var session = await signable.BeginSignedMultipartUploadAsync(spec, TimeSpan.FromHours(1));

// 2. Client PUTs each SignedPart.ContentLength bytes to SignedPart.UploadUrl,
//    collects the ETag from each response, returns the list to the backend.

// 3. Backend finalizes.
await signable.CompleteSignedMultipartUploadAsync(session.UploadId, clientReportedEtags);

// If the client gives up:
// await signable.AbortSignedMultipartUploadAsync(session.UploadId);

MultipartUploadSpec validates your intent against backend limits when BeginSignedMultipartUpload runs — S3 requires parts ≥ 5 MiB (except the last) and at most 10,000 parts per upload.

Metadata behavior on multipart:

  • The regular write stream and the signed IMultipartUploadSignable flow accept FileWriteOptions — pass S3WriteOptions for storage class / SSE. On the multipart stream path, values are bound at CreateMultipartUpload and applied to the cached snapshot when the upload completes. Omit the options (or pass null) for bucket defaults.
  • BeginSignedMultipartUploadAsync(spec, expiresIn, options) — applies the options at CreateMultipartUpload; they take effect on the object when the client completes the upload.
  • CompleteSignedMultipartUploadAsync — calls CompleteMultipartUpload with the client-returned ETags. The bytes never passed through the backend, so the file's cached Length/metadata snapshot is invalidated (IsLoaded returns false); the next metadata access — or an explicit Refresh() — fires one HEAD to re-sync.
  • AbortSignedMultipartUploadAsync — calls AbortMultipartUpload.

Metadata

This driver fully applies FileWriteOptions on writes. Metadata flows in one direction per phase:

  • Write — pass S3WriteOptions (a subclass of FileWriteOptions) to any write method. The values are applied as PutObject headers when the write commits.
  • Read — call GetMetadataAsync / GetMetadata(). It returns an AmazonS3FileMetadata (a subclass of FileMetadata); downcast to read StorageClass / ServerSideEncryption.

The read snapshot is immutable — there are no setters and no IsModified flag. To change an object's metadata you write it again with S3WriteOptions (which re-uploads the bytes); there is no metadata-only update path.

Writing metadata — S3WriteOptions

public class S3WriteOptions : FileWriteOptions
{
    // inherited: ContentType, CacheControl, Metadata
    public string StorageClass         { get; set; }
    public string ServerSideEncryption { get; set; }
}
Field S3 mapping Notes
ContentType Content-Type header Any MIME type string. null = bucket default / omitted.
CacheControl Cache-Control header HTTP cache directive, e.g. "public,max-age=3600". null = omitted. Applied on write and read back via GetMetadataAsync.
Metadata x-amz-meta-* user-metadata Free-form key/value. Case-insensitive keys. Up to ~2 KB total per AWS limits (driver doesn't enforce).
StorageClass x-amz-storage-class header "STANDARD", "STANDARD_IA", "ONEZONE_IA", "INTELLIGENT_TIERING", "GLACIER", "DEEP_ARCHIVE", "GLACIER_IR". null = bucket default (typically STANDARD).
ServerSideEncryption x-amz-server-side-encryption header "AES256" (SSE-S3) or "aws:kms" (SSE-KMS). null = bucket default / no explicit header.
var file = hub.Root.OpenFile("report.pdf");

await file.SetBytesAsync(pdfBytes, new S3WriteOptions
{
    ContentType          = "application/pdf",
    CacheControl         = "public,max-age=86400",
    StorageClass         = "STANDARD_IA",
    ServerSideEncryption = "AES256",
    Metadata             = new Dictionary<string, string> { ["owner"] = "team-x" },
}, ct);

Passing the base FileWriteOptions works too — the driver reads the shared fields and leaves StorageClass / ServerSideEncryption at bucket defaults. The options live with the write stream, so an abandoned write stream never affects a later write.

Reading metadata — AmazonS3FileMetadata

public sealed class AmazonS3FileMetadata : FileMetadata
{
    // inherited (read-only): Tags, ContentType, CacheControl
    public string StorageClass         { get; }
    public string ServerSideEncryption { get; }
}
var meta = await file.GetMetadataAsync(ct);
Console.WriteLine(meta.ContentType);        // "application/pdf"
Console.WriteLine(meta.Tags["owner"]);      // "team-x"

if (meta is AmazonS3FileMetadata s3)
{
    Console.WriteLine(s3.StorageClass);          // "STANDARD_IA"
    Console.WriteLine(s3.ServerSideEncryption);  // "AES256"
}

GetMetadataAsync returns the cached snapshot when the driver already loaded it (a strict OpenFile / TryOpenFile that paid a HEAD), and fires a single HEAD otherwise.

When the snapshot is populated

Moment Snapshot contents
Directory.CreateFile(name) Empty. Destination was just created with bucket defaults.
Directory.TryOpenFile(name) / OpenFile(name) Populated from the HEAD the driver already fires to build the file entry — no extra round-trip.
file.Refresh() / RefreshAsync() Re-fires HEAD, replaces the snapshot from the server state.
After a successful write with options Replaced wholesale to reflect the applied options.
GetFiles() enumeration result Empty — ListObjectsV2 doesn't return per-object metadata. GetMetadataAsync fires a HEAD on first access.

How each operation handles metadata

See the Operations reference for the per-operation table. The rules:

  • Write with S3WriteOptions → values applied as PutObject headers. Unset fields fall back to bucket defaults.
  • Write without options → no metadata headers; bucket defaults apply.
  • CopyTo / MoveTo / Rename (same credentials) → CopyObject default COPY; the destination inherits every metadata field from the source. There is no MetadataDirective = REPLACE path.
  • Cross-credential CopyTo / MoveTo → stream copy through the client; destination created with bucket defaults (source metadata not forwarded).

Change an existing object's metadata

There is no metadata-only update API — to change metadata you re-upload the object with S3WriteOptions:

var file = hub.Root.OpenFile("report.pdf");
var bytes = file.ReadAllBytes();

file.SetBytes(bytes, new S3WriteOptions
{
    StorageClass = "GLACIER",
    Metadata     = new Dictionary<string, string> { ["archived"] = "2026-04" },
});

// After this PutObject: report.pdf is STANDARD→GLACIER and carries the new tag.

⚠️ Atenção: this re-uploads the bytes (a full PutObject). S3 has no header-only mutation that this driver exposes. For large objects, weigh the transfer cost against managing storage class via an S3 Lifecycle rule instead.

Fields NOT exposed

These S3 features exist but are not surfaced by the driver in v1. Configure them at the bucket level (default encryption, lifecycle, bucket tagging) or use the AWS SDK directly if you need per-object control:

  • Object Tagging (PutObjectTagging / GetObjectTagging) — separate API from user-metadata, up to 10 key/value pairs with different size limits. The Metadata / Tags above map to user-metadata, not Object Tagging.
  • Object versioning — version-id-aware reads/writes, version listing.
  • Object Lock (retention, legal hold).
  • ACLs — we don't set x-amz-acl or grant headers; use bucket policies or the SDK directly.
  • Checksum algorithms — we rely on the SDK's default (MD5 for single-PUT, per-part ETags for multipart).
  • Object Lambda / Access Points — access via a different hostname; use a pre-configured IAmazonS3 passed to FromClient.

DI

services.AddFileHub<IAmazonS3FileHub>(sp =>
    AmazonS3FileHub.Create(
        AmazonS3HubOptions.FromProfile("reports", region: "us-east-1", rootPath: "archive/2026")));

For multiple buckets, use named hubs — see Dependency Injection.

Disposal

Strategy Owns the SDK client?
Profile, Credentials Yes — hub creates the AmazonS3Client and disposes it on Dispose().
Client No — caller owns it; hub disposal is a no-op on it.

Integration tests

The test project FileHub.AmazonS3.Tests ships three integration-test classes under Integration/ that only run when AWS environment variables are present — otherwise they skip, so CI stays green without credentials.

Test class Covers Required env vars
RealS3ClientIntegrationTests Upload / download / delete, 404 mapping, presigned GET AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION, FILEHUB_S3_BUCKET
RealS3MultipartIntegrationTests Regular write-stream multipart round-trip (6 MiB → 2 parts), presigned multipart PUT flow with real HttpClient Same as above
RealS3CrossTargetIntegrationTests CopyTo cross-bucket, CopyTo cross-region (routing through destination's endpoint) Above plus FILEHUB_S3_BUCKET_B, AWS_REGION_B
# Minimum run — skips cross-bucket/region.
AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_REGION=us-east-1 \
FILEHUB_S3_BUCKET=my-test-bucket \
dotnet test tests/FileHub.AmazonS3.Tests

# Full run — exercises cross-region routing (AWS_REGION_B must differ).
... AWS_REGION=us-east-1 FILEHUB_S3_BUCKET=alpha \
    AWS_REGION_B=eu-west-1 FILEHUB_S3_BUCKET_B=beta \
    dotnet test tests/FileHub.AmazonS3.Tests

The cross-region test (CrossRegion_CopyTo_ServerSide) is auto-skipped when AWS_REGION_B == AWS_REGION, since there's no routing to exercise in that case. The cross-bucket test runs regardless — same or different region, as long as both buckets exist.

Clone this wiki locally