Skip to content

feat(sdk/go): three-layer proto-to-API coverage checks (RPC methods, converter wiring, message auto-discovery) #3162

Description

@rhuss

User Story

As a Go SDK maintainer, I want automated detection when new proto RPCs or message types are added but not yet exposed through the curated SDK API, so that coverage gaps are visible without manual tracking.

As a proto contributor, I want these checks to be non-blocking on my PR, so that I can evolve the API without being gated by SDK implementation work.

Problem Statement

The Go SDK has a three-layer chain from proto definitions to public API: proto fields → converter → curated client. Today, only the first link is checked.

coverage_test.go uses protobuf reflection to verify that all fields on specific message types are handled by converters. This catches new fields added to existing messages. However:

  1. New message types are not auto-discovered. Each message needs a hand-written test function with a handled field set. A new RPC with new request/response message types has no coverage test until someone writes one.
  2. New RPC methods are not checked. A new service RPC in the proto can land with regenerated stubs, pass CI, and never be wrapped in the curated ClientInterface.
  3. Orphaned converters are not detected. A converter function can exist without any curated client method referencing its output type.

Because the Go SDK commits generated stubs, every proto-changing PR touches sdk/go/, which triggers coverage_test.go. This is the right behavior for field coverage (narrow, mechanical fix). But extending it to RPC-level coverage would block proto contributors on full SDK implementation, which #2825 explicitly rejected.

Impact / Why This Matters

Current behavior: A PR adding a new RPC (e.g., PauseSandbox) regenerates Go stubs, passes all CI checks, and merges. The curated Go SDK silently omits the new capability. SDK consumers discover missing methods through trial and error.

Current workaround: SDK maintainers manually monitor proto changes and audit the curated client surface. The coverage_test.go catches field gaps on known messages but not structural gaps (missing RPCs, missing converters).

Why the workaround is insufficient: As the API surface grows and more contributors change proto definitions, manual tracking doesn't scale. The existing field-level coverage test created a false sense of completeness: it catches one class of gap but not the other two.

Proposed Design

Add two new reflection-based coverage checks and extend auto-discovery for the existing one.

Layer 1 (extend existing): Auto-discover message types

Instead of requiring a hand-written TestConverterCoversAllProtoFields_* function per message type, scan proto service descriptors to collect all request/response message types (and their nested messages) automatically. Compare against the set of messages that have converter coverage. Flag any message type used by a service RPC that has no converter coverage test.

Layer 2 (new): RPC method coverage

Use protoreflect to enumerate all RPC methods across the OpenShell and Inference service descriptors. Compare against methods on ClientInterface and its sub-client interfaces (SandboxClient, ProviderClient, etc.). Flag any RPC with no corresponding curated client method.

Layer 3 (new): Converter-to-client wiring

Use reflect to scan all ClientInterface method signatures (arguments and return types). Collect every domain type referenced. Compare against the set of domain types produced by converter functions. Flag any converter output type not referenced by a client method signature.

All three layers follow the same self-maintaining pattern: reflection-based scanning with a skipped set for intentional exceptions (e.g., internal-only RPCs, types exposed only through raw). No manual lists to keep in sync.

Enforcement model:

Check Enforcement Rationale
Field coverage on existing messages Hard (blocks PR) Narrow, mechanical fix. Already works this way.
RPC method coverage Soft (non-blocking) Requires design decisions about the SDK API surface.
Converter-to-client wiring Soft (non-blocking) Catches orphaned converters, lower urgency.

Soft enforcement options (one or both):

Acceptance Criteria

  • Auto-discovery: new proto message types used by service RPCs are detected without hand-written test functions
  • RPC coverage: new service RPC methods not wrapped in ClientInterface are detected
  • Converter wiring: converter output types not referenced by client method signatures are detected
  • All checks use a skipped set for intentional exceptions with justification comments
  • Field-level coverage remains a hard gate (existing behavior unchanged)
  • RPC and converter checks are non-blocking on proto-changing PRs
  • Gaps detected by soft checks create actionable issues (via ci(sdk): add proto drift detection and sync notifications #3123 infrastructure or separate workflow)

Alternatives Considered

Hard gate for all layers: Rejected. Because Go stubs are committed, every proto-changing PR touches sdk/go/, so a hard RPC coverage gate would block proto contributors on SDK implementation. This contradicts the project's decision in #2825.

External linter instead of test-time reflection: Possible but adds toolchain complexity. The reflection-based approach matches the existing coverage_test.go pattern and requires no new dependencies.

Skip Go-specific checks, rely on daily cron only: The daily cron from #3123 detects stub-level drift but not converter or RPC gaps. The reflection-based checks provide deeper coverage that the cron infrastructure alone cannot.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions