A Go error library for web services (HTTP / gRPC). See DESIGN.md for the full design spec.
- Error codes are the source of truth —
Codevalues 0–16 match gRPC'scodes.Code. HTTP / gRPC statuses and retryability are all derived from a lookup table. - Stdlib-only core — the gRPC conversion is isolated in a separate
grpcerrmodule. - Lightweight call-site tracking —
New/Wrapeach record one caller frame. No stack traces. - Fully compatible with the standard
errorspackage — works withIs/As/Unwrap/Join. - Internal vs. public separation — exactly four channels reach a client (public message, public fields, field violations, retry delay); internal messages and attrs stay in your logs.
slog.LogValuerimplementation and RFC 9457 (Problem Details) responses, including extension members.- Same taxonomy end to end — carry code names across gRPC and rebuild them on the calling side; retry delays (
RetryInfo) and validation details (BadRequest) ride along automatically.
go get github.com/repenguin22/errtrail
go get github.com/repenguin22/errtrail/grpcerr # only if you use gRPC
go get github.com/repenguin22/errtrail/otelerr # only if you use OpenTelemetry
The core module supports Go 1.22+. The grpcerr and otelerr modules require Go 1.25+, following the minimums of google.golang.org/grpc and go.opentelemetry.io/otel.
errtrail is built around one small discipline that mirrors the three places an error passes through in a request:
source ─────▶ middle layers ─────▶ boundary
classify add context log once, respond
Each layer has exactly one job. Get these three right and everything else — the
HTTP/gRPC status, the client message, structured logs, retry decisions — falls
out of the Code for free. The rest of this section walks each layer in turn;
Best practices distils them into a checklist.
Where the error originates, attach everything you know in one place: a Code
(the classification, and the single source of truth for the HTTP/gRPC status), an
internal message for your own eyes, and — only if a client should see them —
a public message, structured public fields, field violations
for validation failures, and a retry delay for dynamic server pushback.
With adds slog attributes that stay in the logs.
if row == nil {
return errtrail.New(errtrail.NotFound, "user row missing").
WithPublic("User not found"). // safe to show a client
With(slog.String("user_id", id)) // logs only, never sent to a client
}Errors from the standard library or a third-party package enter the chain the
moment you Wrap them — that is also where you assign their Code:
if err := row.Scan(&u); err != nil {
if errors.Is(err, sql.ErrNoRows) {
return errtrail.Wrap(err, "user row missing").
WithCode(errtrail.NotFound).
WithPublic("User not found")
}
return errtrail.Wrap(err, "scan user").WithCode(errtrail.Internal)
}You never write an HTTP status or a gRPC code by hand; both are derived from the
Code via a lookup table.
As the error travels up, Wrap it to record the path it took. Don't assign a new
code (the one from the source is inherited) and don't log yet — the trace already
remembers every layer.
func (s *Service) Profile(ctx context.Context, id string) (*Profile, error) {
u, err := s.repo.Get(ctx, id)
if err != nil {
return nil, errtrail.Wrap(err, "load profile") // context, not a new code
}
// ...
}
⚠️ Typed-nil footgun.Wrapreturns*Error, noterror. Returning it unconditionally from a function whose signature iserrorproduces a non-nil error holding a nil pointer, even on the success path. Keep theif err != nilguard (as above) rather thanreturn errtrail.Wrap(err, "...")at the end of a function.
At the top of the request (an HTTP handler or a gRPC method) do two things, each exactly once: log the full internal error, then turn it into a response that carries only public information.
// HTTP handler
slog.ErrorContext(ctx, "request failed", slog.Any("error", err)) // everything internal
_ = problem.Write(w, err) // only public data leaves
// gRPC handler
return nil, grpcerr.ToError(err)The internal message, attrs, and trace never leave your process. The client gets
the public message and the machine-readable code name; when no public message
was set, HTTP clients see the generic status text as the problem title, and
gRPC clients get the code name (e.g. "UNAVAILABLE") as the status message —
never HTTP wording, never an empty message.
If you use OpenTelemetry, mark the active span in the same breath — the
otelerr module records the error (code name + attrs, never the public
message) and sets the span status from the Code, following the OTel
server-span conventions: only server-fault codes (Internal, Unavailable,
DeadlineExceeded, …) mark the span as an error, so NotFound-heavy traffic
doesn't light your trace UI up:
otelerr.Record(ctx, err) // exception event + span status derived from the Code
// And to join logs with traces, attach the IDs at the source:
err := errtrail.New(errtrail.Internal, "db write failed").
With(otelerr.TraceAttrs(ctx)...)For validation-style errors, attach field violations at the source with
WithFieldViolation — one typed channel that feeds both boundaries: the
problem package emits them as the errors extension member, and
grpcerr.ToStatus attaches them as an errdetails.BadRequest. instance
comes from the request at the boundary, not from the error:
// At the source:
err := errtrail.New(errtrail.InvalidArgument, "email failed regexp").
WithPublic("Validation failed").
WithFieldViolation("email", "must be a valid email address")
// At the HTTP boundary:
_ = problem.Write(w, err, problem.Instance(r.URL.Path))
// {"code":"INVALID_ARGUMENT","detail":"Validation failed",
// "errors":[{"field":"email","description":"must be a valid email address"}],
// "instance":"/users","status":400,"title":"Bad Request"}
// At the gRPC boundary the same error yields InvalidArgument with an
// errdetails.BadRequest{FieldViolations: [{Field: "email", ...}]} detail.Arbitrary structured extras beyond field violations still go through
WithPublicField(key, value); they surface as RFC 9457 extension members
under their own key (an explicit "errors" public field overrides the
derived one). That override is HTTP-only by design: public fields never
reach the gRPC wire, so the BadRequest detail always derives from the
typed violations.
RFC 9457's type member (a URI reference identifying the problem type,
defaulting to about:blank) is also supported, but opt-in: it's omitted by
default because errtrail has no documentation site of its own to point type
at. Set problem.TypeURL once at startup to derive one from the Code:
problem.TypeURL = func(c errtrail.Code) string {
return "https://errors.example.com/" + c.String()
}
// Now From(err).Type == "https://errors.example.com/INVALID_ARGUMENT"Over gRPC, set the service domain once at startup so custom code names (not
just their numeric gRPC code) survive the wire as a machine-readable
errdetails.ErrorInfo:
grpcerr.Domain = "myservice.example.com" // opt-in; empty (default) attaches nothingWhen your service is itself a client, bring the downstream error back into the
same taxonomy with FromError. Wire codes map one-to-one, and a custom code is
recovered from its ErrorInfo.Reason when you have registered the same code
locally. Retry decisions then derive from the same Code:
res, err := client.GetUser(ctx, req)
if err != nil {
terr := grpcerr.FromError(err) // wire status -> errtrail.Code; wraps err
if errtrail.IsRetryable(terr) { // Unavailable/DeadlineExceeded/ResourceExhausted/Aborted
if d, ok := grpcerr.RetryDelay(err); ok { // server-recommended delay, if any
return retryAfter(ctx, op, d)
}
return retryWithBackoff(ctx, op)
}
return nil, errtrail.Wrap(terr, "call user service")
}Custom-code recovery requires the ErrorInfo.Reason to name a locally
registered code and its registered gRPC code to match the wire code. If the
client also talks to services outside your taxonomy, tighten it further with
TrustedDomain — recovery then additionally requires the ErrorInfo.Domain
to match:
terr := grpcerr.FromError(err, grpcerr.TrustedDomain("user.example.com"))There's no HTTP equivalent of FromError — gRPC codes map to Code
one-to-one, but an HTTP status is inherently many-to-one (400 could be
InvalidArgument, FailedPrecondition, or OutOfRange), so a generic
FromHTTPStatus would guess wrong as often as it guessed right. If you call
another HTTP service, classify its response explicitly at the call site
instead (e.g. errtrail.Wrap(err, "call user service").WithCode(...), picking
the Code from context you have and the library doesn't).
errtrail's inspection functions all walk an errors.Join-ed chain depth-first,
visiting the first branch first — the same rule used for a plain Wrap chain,
just applied per branch:
a := errtrail.New(errtrail.InvalidArgument, "email invalid").WithFieldViolation("email", "must be valid")
b := errtrail.New(errtrail.InvalidArgument, "age invalid").WithFieldViolation("age", "must be >= 0")
joined := errors.Join(a, b)CodeOfreturns the first non-OK code andPublicMessagethe first non-empty public message, each in depth-first walk order — a single response can only carry one status and one message, so the first hit wins by convention. Note these are separate walks: when the first branch has a code but no public message, the message may come from a different branch than the code. That combination is intentional (each walk finds the best available value), but if the branches diverge, decide deliberately — set an explicitWithCode+WithPublicon theJoinresult rather than relying on which validator ran first.Trace/Attrscollect every branch, so nothing is lost from logs.FieldViolationsconcatenates every branch — it's a list, not a map, so joined validators never collide: both violations above reach the client (theerrorsmember over HTTP,BadRequestover gRPC). This is why validation data belongs in violations.PublicFieldscollects every branch too, but a duplicate key keeps only the first branch's value — the same outermost-wins ruleWrapchains use; a"field"key set by bothaandbwould silently keep onlya's. Use distinct keys per branch, or — for anything list-shaped — useWithFieldViolation, which has no keys to collide.
%+v prints the full chain — message, code, public message, public fields and
field violations, attrs, and the recorded trace — for debugging and tests:
load profile: query user: sql: no rows in result set
code: NOT_FOUND
public: User not found
public.fields: resource=user
public.violations: user_id=does not exist
attrs: user_id=42
trace:
example.com/app/service.(*UserService).Profile (/src/app/service/user.go:88): load profile
example.com/app/repo.(*UserRepo).Get (/src/app/repo/user.go:42): query user
A checklist, with the reasoning behind each rule:
-
Attach the
Codeat the source, once. It is the single source of truth: the HTTP status, the gRPC code, andIsRetryableall derive from it. Deriving a status by hand anywhere else is how they drift apart. -
Wrapto add context; don't reclassify in the middle. A middle layer that sets a new code hides the real cause. Let the source's code propagate; only override withWithCodewhen you are deliberately translating one failure into another (e.g. a downstreamUnavailableyou choose to surface asInternal). Beware:WithCodedoes not clear public data set below it — an innerWithPublic("User not found"),WithPublicField,WithFieldViolation, orWithRetryDelaywould still reach the client through the new response (a 403 hiding a NotFound must not carry the lookup's field violations either). When the point of reclassifying is to hide the original failure, addWithoutPublic()— it blocks all four public channels below it — then set a freshWithPublicif needed. Call it on a freshWrap(err, ...), not onerritself: the node you call it on keeps its own public data above the barrier, soerr.WithoutPublic()still exposes everything attached directly toerr. -
Keep internal and public strictly separate. Exactly four channels reach a client —
WithPublic,WithPublicField,WithFieldViolation, andWithRetryDelay— and nothing else ever does; the internal message andWithattrs are for logs. Treat them all with the same care: a violation's field name and description go verbatim into the HTTPerrorsmember and the gRPCBadRequest. When unsure whether a string is safe to expose, leaveWithPublicunset — the client gets the generic status text (HTTP) or the code name (gRPC) rather than a leaked detail. -
Classify with
CodeOf; match sentinels witherrors.Is. For "what kind of failure is this?" switch onerrtrail.CodeOf(err)— errtrail deliberately does not overloaderrors.Isfor codes, because implicit code matching is hard to predict.errors.Is/errors.Asstill work for sentinel values (sql.ErrNoRows, your ownvar Err…) becauseWrappreserves the cause. -
Log once, at the boundary, with
slog.Any. The trace already records everyWrappoint, so logging at each layer only duplicates it. Passing the*Errortoslog.Any("error", err)expands its code, trace, and attrs viaLogValue; the public message is intentionally omitted from logs. -
Register custom codes at
init. Registration is safe at any time — the registry is swapped atomically (copy-on-write) — but registering atinitkeeps the taxonomy identical for every request the process ever serves. Give each custom code a uniqueSCREAMING_SNAKEname; it is the wire and config lookup key. -
Use
Newf/Wrapfonly for the internal message. Interpolated request data belongs inWithattrs (queryable in logs) orWithPublicField, not baked into a message string. -
Building an error factory? Use
NewSkip/WrapSkip. A helper likeapperr.NotFound(msg)built on plainNewrecords its own line in every trace, so all errors appear to originate at the factory.NewSkip(1, ...)skips the factory layer and points the frame at the real call site —skipcounts wrapper layers, like zap'sAddCallerSkip. (grpcerr.FromErrordoes this internally, so its frame is your call site too.) -
Retry on
IsRetryable, not a hand-maintained list. Built-insUnavailable,DeadlineExceeded,ResourceExhausted, andAbortedare retryable; custom codes opt in witherrtrail.Register(c, name, httpStatus, grpcCode, errtrail.Retryable())— orerrtrail.RetryAfter(d), which also ships a recommended delay to gRPC clients as aRetryInfodetail. It's a transience hint derived only from theCode— whether replaying the request is safe (idempotency, retry budget, server pushback) is still your call. -
Push the actual wait with
WithRetryDelay. A registeredRetryAfteris a static per-code hint; a rate limiter knows the real time to the next token at request time.New(throttled, "bucket empty").WithRetryDelay(limiter.NextAvailable())ships that value as the gRPCRetryInfo, taking precedence over the registered delay. Over HTTP, derive aRetry-Afterheader in the handler —problemdeliberately doesn't emit one:if d, ok := errtrail.LookupRetryDelay(err); ok { w.Header().Set("Retry-After", strconv.Itoa(int(math.Ceil(d.Seconds())))) } problem.Write(w, err)
*Error implements slog.LogValuer, so passing it to slog.Any("error", err) nests it as a structured group instead of a flat string — as long as err is a *errtrail.Error (wrap plain errors with errtrail.Wrap before logging them). The public message, public fields, field violations, and retry delay are deliberately left out; they're for response generation, not logs.
slog.New(slog.NewJSONHandler(os.Stdout, nil)).
Error("request failed", slog.Any("error", err)){
"time": "2026-07-07T23:25:18.408363+09:00",
"level": "ERROR",
"msg": "request failed",
"error": {
"msg": "load profile: query user: sql: no rows in result set",
"code": "NOT_FOUND",
"trace": [
"main.main (/app/main.go:17): load profile",
"main.main (/app/main.go:13): query user"
],
"user_id": "42"
}
}| Package | Dependencies | Role |
|---|---|---|
errtrail |
standard library only | Core. Code, Error, inspection, formatting, slog |
errtrail/problem |
standard library only | RFC 9457 response generation |
errtrail/grpcerr |
google.golang.org/grpc |
*status.Status conversion (separate go.mod) |
errtrail/otelerr |
go.opentelemetry.io/otel |
span recording / status from Code (separate go.mod) |
Register your own codes (values >= 100; 0–99 are reserved), preferably from
init. Registration is thread-safe — the registry is replaced atomically
(copy-on-write), so lookups never observe a partial write even if you register
late — but init is still the recommended pattern so that every request sees
the same taxonomy.
HTTP-first services often need custom codes. The 17 built-in codes are gRPC's, and
HTTPStatusmaps each one to a single fixed status — so they're a many-to-one mapping from the HTTP side.InvalidArgumentis always400, for instance, even if your API would rather distinguish400from422 Unprocessable Entity, orAlreadyExists/Abortedboth landing on409when you'd want to tell them apart by status code alone. This doesn't break anything — register a custom code with whatever HTTP status you need — but if you're building an HTTP-only API, expect to reach for custom codes more often than the 17 built-ins alone would suggest.
const RateLimited errtrail.Code = 100
func init() {
errtrail.Register(
RateLimited,
"RATE_LIMITED", // unique SCREAMING_SNAKE name; the wire/config lookup key
http.StatusTooManyRequests, // HTTP status (must be in [400, 599] — a Code classifies an error)
8, // gRPC code, ResourceExhausted (must be 1–16; 0 is OK)
errtrail.Retryable(), // optional: makes IsRetryable report true
)
}Instead of Retryable(), RetryAfter(d) marks the code retryable and
records a recommended delay: grpcerr.ToStatus then attaches an
errdetails.RetryInfo carrying it, and clients read it back with
grpcerr.RetryDelay(err). (Built-ins can't carry a delay — they aren't
registered through Register.)
errtrail.Register(Throttled, "THROTTLED", 429, 8, errtrail.RetryAfter(2*time.Second))Once registered, RateLimited behaves like a built-in everywhere: HTTPStatus,
GRPCCode, String, Retryable, problem.From, and grpcerr.ToStatus all
resolve it through the same table, and CodeByName("RATE_LIMITED") recovers it
(used by grpcerr.FromError to rebuild the code from the wire).
Register panics on misuse — a code below 100, a duplicate code or name, a
name that doesn't match [A-Z][A-Z0-9_]+[A-Z0-9] (≤ 63 chars, the
ErrorInfo.Reason wire constraints), an HTTP status outside [400, 599], or
a gRPC code outside [1, 16] — so a mistake surfaces at startup rather than
mid-request.
Apple M1 Pro, Go 1.26, as of v1.3. New / Wrap are still 1 alloc each, at
144 B — the same allocator size class as v1.1, despite v1.3 adding a retry
delay field: noPublicBelow sits right after Code in the struct, filling
its padding byte instead of costing one at the end (96 B pre-v1.1).
BenchmarkNew-10 7068086 160.8 ns/op 144 B/op 1 allocs/op
BenchmarkWrap-10 7298881 158.1 ns/op 144 B/op 1 allocs/op
BenchmarkWrapChain3-10 2369467 497.3 ns/op 432 B/op 3 allocs/op
BenchmarkFormatPlusV-10 800520 1428 ns/op 3489 B/op 25 allocs/op
The three modules are versioned independently under Semantic Versioning, each with its own tag line:
git tag v0.1.0 # core
git tag grpcerr/v0.1.0 # grpcerr submodule
git tag otelerr/v0.1.0 # otelerr submodule
A submodule's require points at a tagged core version (no replace); when developing several modules together, put a go.work outside the repo. See CHANGELOG.md for the full history.
All three modules are v1.0+ and follow the SemVer compatibility promise: no breaking change to the public API or to documented wire/response behavior without a major version bump. A minor version may add API surface; a patch is fixes only. Behavior changes of any kind are called out in the changelog.