Skip to content

v2.8.0

Choose a tag to compare

@Laurianti Laurianti released this 05 Oct 08:12
· 123 commits to main since this release

Minor release of ChromaDotNet.Client and ChromaDotNet.Client.DependencyInjection: hybrid search computed by the client, large reads and writes handled for you, Chroma Cloud by connection string, traces and metrics, and an API aligned with the .NET conventions.

dotnet add package ChromaDotNet.Client

New

  • Hybrid search, as the Python client of Chroma does it:
    • AddAsync, UpdateAsync and UpsertAsync compute the BM25 vectors of the chroma_bm25 indexes of the schema, from the document or from a metadata key;
    • ChromaRank.SparseKnn(text, key) takes a text, which SearchAsync turns into a sparse vector;
    • ChromaSparseVectorIndex.Bm25Function is the ChromaBm25 with the settings of the schema, also for a collection created by Python.
  • A collection with a schema keeps its space. The space goes in the schema, as the Python client writes it: Chroma Cloud rejects the space as a configuration together with a schema.
  • Large reads and writes, by default:
    • writes go in batches of the max_batch_size of the server, and GetAsync reads in pages;
    • on Chroma Cloud, 300 records at a time, its quota;
    • when a server rejects a batch beyond its quota of records, the client sends it again in batches of that quota.
  • Chroma Cloud by connection string: ChromaConfigurationOptions.FromConnectionString("Endpoint=...;Token=...;Tenant=...;Database=...").
  • Traces and metrics with OpenTelemetry: AddSource(ChromaTelemetry.ActivitySourceName) and AddMeter(ChromaTelemetry.MeterName). A span for each operation, with the attributes of the semantic conventions for database clients, and its duration in db.client.operation.duration.
  • Mocks in tests: the members of ChromaClient and ChromaCollectionClient are virtual, and a protected constructor makes a client without a server, as in the Azure SDKs.
  • QueryAsync with ids, with one or more query embeddings.
  • .NET Framework 4.6.2: a build of its own, which runs next to OpenTelemetry without binding redirects.
  • Dependency injection: IServiceProvider.CreateChromaClient(options, httpClientName) makes the client that AddChromaClient registers, for an integration with an HttpClient name of its own.
  • Documentation: every parameter and result of the public API, for IntelliSense; ChromaCollectionSchema.ToString() gives the JSON of a schema.

Aligned with the .NET conventions

2.8.0 brings the API in line with the .NET design guidelines. What to change in the code:

Change What to do
Asynchronous methods end in Async GetOrCreateCollection → GetOrCreateCollectionAsync, Add → AddAsync, Query → QueryAsync, and so on
Read-only lists Parameters, results and models are IReadOnlyList<T>: a List<T> or a collection expression still goes in
Read-only dictionaries Metadata are IReadOnlyDictionary<string, object>: write them new Dictionary<string, object> { ["key"] = value }
One embedding, one URI per record ChromaCollectionEntry and ChromaCollectionQueryEntry have Embedding and Uri; Data is gone, as no Chroma server fills it
Parameters named as the properties new ChromaConfigurationOptions(uri, tenant: ..., database: ...)
Metadata values read as written Strings stay strings and lists are List<object>; WithMetadataValues(ChromaMetadataValues.Inferred) reads them as before
Batches and pages by default WithBatchSplitting(false) sends a write in one request
ChromaException for the failures of a request Network errors, answers that are not the expected JSON and timeouts; other exceptions, like an assembly that does not load, go as they are
AddChromaClient and AddKeyedChromaClient return the services Code that ignores the result compiles as before
One type for each filter ChromaWhereOperator and ChromaWhereDocumentOperator, the types the methods take; ChromaWhere and ChromaWhereDocument are gone