Repository navigation
py-cid ↔ go-cid Feature Parity Analysis #1365
sumanjeet0012
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
py-cid Feature Parity Analysis
Table of Contents
Executive Summary
The Python implementation (
py-cid) and the Go reference implementation (go-cid) are much closer to parity than the other multiformat libraries analysed. Both support CIDv0/v1, builders, prefixes, sets, JSON serialization, and stream parsing. However, Go leads in several important areas: typed error handling, zero-allocation operations (ByteLen,WriteBytes), standard Go serialization interfaces, a universalParse()function, fuzz testing, and benchmarks. Python leads with a cleaner OOP design (separateCIDv0/CIDv1classes), version conversion methods, IPFS path parsing, and more Pythonic set operations.StringOfBase(),Encode(encoder)ValueErrorvs Go's typed errorsCidFromReaderrobustnessUndefsentinelByteLen()/WriteBytes()Parse(interface{})universal parserArchitecture Comparison
Go Architecture (Single Type, Performance-Oriented)
Key design principle:
Cidis a single struct wrapping a binary string (Cid{str string}). CIDv0 and CIDv1 are distinguished only by theVersion()method. This enables zero-copy operations and avoids allocations. The internal string stores the raw CID bytes directly.Dependencies:
go-multibase,go-multihash,go-varint.Python Architecture (OOP, Two Classes)
Key design difference: Python uses separate classes for
CIDv0andCIDv1with a sharedBaseCIDbase. This provides type-level distinction but means more isinstance checks and no single-type uniformity.Dependencies:
py-multibase,py-multicodec,py-multihash,morphys,varint.Feature-by-Feature Gap Analysis
3.1 CID Type Design
Cidstruct (Cid{str string})CIDv0(BaseCID),CIDv1(BaseCID)_version,_codec,_multihashfieldsUndefsentinelvar Undef = Cid{}Defined()c.str != ""defined()Equals(other)__eq____hash__/ hashable__hash__based on (version, codec, multihash)Compare/ ordering3.2 Construction & Parsing
NewCidV0(mhash)CIDv0(mhash)NewCidV1(codec, mhash)CIDv1(codec, mhash)Decode(string)from_string(s)Cast([]byte)from_bytes(b)Castwith trailing bytes checkfrom_bytes_strict(b)Parse(interface{})MustParse(interface{})must_parse(v)CidFromBytes(data)→ (n, Cid, err)from_bytes()doesn't return bytes consumedCidFromReader(r)→ (n, Cid, err)from_reader()exists but is fragiletryNewCidV0()enforcesCIDv0()constructor doesn't validateParse()handles/ipfs/parse_ipfs_path()+from_string()3.2.1 CIDv0 Validation Gap
Go validates CIDv0 on construction:
Python's
CIDv0.__init__()accepts any multihash without validation:This means
CIDv0(b"\x00\x00")succeeds in Python but would panic in Go.3.2.2
CidFromReaderRobustness GapGo's
CidFromReaderis a 90-line function that:bufByteReaderto avoid allocationsio.EOFvsio.ErrUnexpectedEOFcorrectlyPython's
from_reader()is ~60 lines but:reader.read(n)3.3 Encoding & String Representation
String()(default)__str__()→ base58btc for bothStringOfBase(base)Encode(encoder)encode(encoding)accepts string name3.4 Serialization (JSON / Binary / Text)
MarshalJSON/UnmarshalJSON{"/": "<cid>"}to_json_dict()/from_json_dict()+CIDJSONEncoderMarshalBinary/UnmarshalBinaryMarshalText/UnmarshalTextUndef→null3.5 Properties & Accessors
Version()uint64versionproperty →intType()(codec)uint64(multicodec code)codecproperty →str(name)Hash()(multihash)mh.Multihashmultihashproperty →bytesBytes()[]bytebuffer/to_bytes()→bytesByteLen()WriteBytes(w io.Writer)KeyString()key_string()→ latin-1 decoded stringLoggable()map[string]interface{}loggable()→dictPrefix()Prefixstructprefix()→Prefixobject3.6 Version Conversion
to_v1()to_v0()(validates dag-pb)3.7 Prefix
Prefixstruct{Version, Codec, MhType, MhLength}(all uint64/int)Prefix(version, codec, mh_type, mh_length)(str for codec/mh_type)Prefix.Sum(data)mh.Sum()with all hash typesPrefix.Sum()identity hash handlingmh.IDENTITYPrefix.Bytes()to_bytes()PrefixFromBytes(buf)Prefix.from_bytes(data)Prefix.GetCodec().codecproperty)Prefix.WithCodec(c)Prefix.v0(),Prefix.v1()3.7.1
Prefix.Sum()Hash Type SupportGo's
Prefix.Sum()delegates tomh.Sum()which supports all registered multihash functions (sha1, sha2-, sha3-, blake2*, keccak, etc.). Python'sPrefix.sum()only implements sha2-256 and sha2-512, raisingNotImplementedErrorfor everything else:3.8 Builder Pattern
BuilderinterfaceSum/GetCodec/WithCodecBuilderABC withsum/get_codec/with_codecV0BuilderV1BuilderV1Builder.Sum()hash supportmh.Sum()V1Builder.WithCodec()V0Builder.WithCodec()auto-upgradePrefiximplementsBuilderPrefixis separate fromBuilder3.9 CID Set
SetCIDSetNewSet()/ constructorNewSet()CIDSet()Add(c)add(cid)Has(c)has(cid)+__contains__Remove(c)remove()usesdiscard()(no error if missing)Len()__len__()Keys()[]Cidkeys()returnslistVisit(c)→ boolvisit(cid)ForEach(f func(Cid) error) errorfor_each(func)returns None__iter__3.10 Error Handling
ErrInvalidCid{Err}(typed, wrappable)ValueError(generic)ErrCidTooShortValueError("argument length can not be zero")ErrInvalidEncodingValueErrorerrors.Is(err, ErrInvalidCid)supportNewCidV0/NewCidV1Go has a proper typed error system:
Python uses generic
ValueErrorfor everything, making it impossible to distinguish CID-specific errors from other value errors.3.11 Utility Functions
ExtractEncoding(v)mbase.Encodingextract_encoding(cid_str)returnsstrParse(v interface{})MustParse(v)must_parse(v)parse_ipfs_path()Parse()is_cid(s)make_cid(*args)3.12 Testing & Benchmarks
benchmark_test.gocid_fuzz.goFeature Matrix Summary
Undefsentinel valueParse(interface{})universal parserCidFromBytes()returns bytes consumedCidFromReader()robustness (multi-byte varints, OOM guard)StringOfBase(base)methodEncode(encoder)with encoder objectMarshalBinary/UnmarshalBinaryMarshalText/UnmarshalTextByteLen()zero-allocationWriteBytes(w)zero-allocationInvalidCIDError, etc.)Prefix.Sum()full hash type supportPrefix.Sum()identity hash handlingPrefix.WithCodec()methodPrefixasBuilder(implements Builder interface)V1Builder.Sum()full hash type supportSet.Remove()raises on missing (vs discard)Set.ForEach()error propagationPhased Roadmap
Phase 1 – Critical Correctness & API Gaps
Goal: Fix validation gaps and add the most important missing APIs.
Estimated effort: 1–2 weeks
1.1 Add CIDv0 validation in constructor
1.2 Add
UNDEFsentinel1.3 Add
parse()universal function1.4 Fix
from_reader()for robustnessRewrite to properly handle multi-byte varints and add OOM protection:
1.5 Expand
Prefix.sum()hash type supportUse
multihash.sum()to support all registered hash functions:Phase 2 – Serialization & Zero-Allocation Operations
Goal: Add standard serialization interfaces and performance-oriented methods.
Estimated effort: 3–5 days
2.1 Change CIDv1 default encoding to base32
Per the CID specification, CIDv1 should default to base32 encoding:
2.2 Add
string_of_base()method2.3 Add
byte_len()method2.4 Add binary/text marshaling helpers
Phase 3 – Error Handling & Robustness
Goal: Add typed error hierarchy and improve error messages.
Estimated effort: 3–5 days
3.1 Add typed error hierarchy
3.2 Update
Set.remove()to match Go semantics3.3 Add error propagation to
CIDSet.for_each()Phase 4 – Testing, Benchmarks & Polish
Goal: Add comprehensive tests, benchmarks, and fuzz testing.
Estimated effort: 1–2 weeks
4.1 Add benchmark suite
4.2 Add round-trip tests
4.3 Add error case tests
4.4 Add property-based testing with hypothesis
Appendix – Python-Exclusive Features
These features exist in Python but not in Go:
CIDv0/CIDv1classesto_v0()/to_v1()methodsfrom_bytes_strict()parse_ipfs_path()CIDJSONEncoderjson.JSONEncodersubclassCIDSet.__iter__CIDSet.__contains__inoperator supportPrefix.v0()/Prefix.v1()class methodsis_cid()functionmake_cid()flexible constructorCIDSet.__len__len()support__hash__on CIDSummary Timeline
UNDEF,parse(), robustfrom_reader(), full hash support inPrefix.sum()string_of_base(),byte_len(), binary/text marshalingSet.remove()fix,Set.for_each()error propagationTotal estimated effort: 3–5 weeks
The highest-priority items are Phase 1 correctness fixes — especially CIDv0 validation (currently accepts invalid multihashes) and
Prefix.sum()hash type support (currently only 2 of 100+ hash types work). These are real bugs that would cause incorrect behavior in production.The Python implementation's OOP design (separate classes,
to_v0()/to_v1(), Pythonic set) are genuine usability advantages that should be preserved and documented as differentiators from Go's single-type approach.All reactions