Skip to content

Release v10.0.0

Latest

Choose a tag to compare

@github-actions github-actions released this 03 Sep 14:49
· 10 commits to main since this release

Version 10.0.0 is the SDK for the Pinecone 2026-07 API. It runs on Python 3.10 through 3.14 and sends 2026-07 on every plane by default. The release graduates the schema-based document interface, SchemaBuilder, and the rest of the 2026-07 API out of pinecone.preview onto the main client objects. There are breaking changes: a short list of the most important changes are outlined below, and Migrating to V10 is the complete accounting.

If your workload upserts and queries an existing index, upgrading does not change what your code means. upsert, query, fetch, update, delete, list, and describe_index_stats take the same arguments and return the same shapes they did on 9.x; pc.Index("movies") and pc.index("movies") both still hand you an index, though only pc.index() narrows its return type, so reach index.documents through that one; and none of those operations is deprecated or scheduled for removal. Indexes you created before 2026-07 go on being served by them. Three things do change on that path, each covered below: a batched upsert paces itself differently, some invalid calls now raise locally instead of being refused by the server, and gRPC upsert_from_dataframe collects partial failures rather than raising on the first one.

The breaking changes are concentrated in pc.indexes.create's arguments and return model, and in the surfaces that graduated out of pinecone.preview. Flat create_index(...) with dimension=, metric=, or spec= needs no edits: it takes its 9.1.0 parameters in the 9.1.0 order, positionally or by keyword. Which family of operations serves an index follows the index itself, not the SDK version — an index created before 2026-07 is addressed through the vector operations above, and one created with a 2026-07 document schema through index.documents.

What's new

Documents API. Records are addressed as documents: a JSON object with an _id, the fields you declared in the index schema, and any metadata alongside them. Upsert, search, fetch, list, update, and delete all live on index.documents, on the sync and async clients alike, and index.documents.batch_upsert handles large loads with the same admission gate and partial-failure reporting the vector path uses. See Quickstart.

from pinecone import DenseVectorQuery, Pinecone

index = Pinecone().index("quickstart")

index.documents.upsert(
    namespace="movies",
    documents=[
        {
            "_id": "movie-001",
            "embedding": [0.1, 0.2, 0.3],
            "title": "Arrival",
        },
    ],
)

results = index.documents.search(
    namespace="movies",
    top_k=3,
    score_by=[DenseVectorQuery(field="embedding", values=[0.1, 0.2, 0.3])],
)

SchemaBuilder. An index's fields are declared as a schema, and SchemaBuilder builds that schema in Python instead of by hand-assembling nested dicts. Every add_* method returns the builder, and build() returns a plain {"fields": {...}} dict you hand straight to pc.indexes.create(schema=...). See Schema Builder.

from pinecone import Pinecone, SchemaBuilder

schema = (
    SchemaBuilder()
    .add_dense_vector_field("embedding", dimension=1024, metric="cosine")
    .add_string_field("title", full_text_search={"language": "en"})
    .build()
)

pc = Pinecone()
pc.indexes.create(
    name="product-search",
    schema=schema,
    deployment={
        "deployment_type": "managed",
        "cloud": "aws",
        "region": "us-east-1",
    },
)

Full-text search configuration. String fields carry their own full-text search config: language, stemming, stop words, and character n-grams for substring and autocomplete matching, set per field at index creation. NgramConfig and FullTextSearchConfig are exported from pinecone for building the config as typed objects; n-grams cannot be combined with stemming or stop words. See Schema Builder.

schema = (
    SchemaBuilder()
    .add_string_field(
        "title",
        full_text_search={"ngram": {"min_gram": 2, "max_gram": 4}},
    )
    .build()
)

Backup schedules. Attach a recurring cadence to an index and backups keep happening without a caller. Schedules live on pc.backup_schedules, with create, list, describe, update, delete, and a per-schedule run history. frequency takes exactly daily, weekly, or monthly; there is no cron support, and only one enabled schedule per index is allowed. See Backups and restore.

schedule = pc.backup_schedules.create(
    index_name="product-search",
    name="daily-backup",
    frequency="daily",
    retention_days=90,
)

for run in pc.backup_schedules.iter_history(schedule_id=schedule.schedule_id):
    print(run.backup_id, run.status, run.record_count)

Organization management on the Admin client. Admin was OAuth, organizations, projects, and API keys. It now also manages the people and machines in an organization: admin.users, admin.invites, admin.service_accounts, and admin.role_bindings. Role bindings are the whole authorization model: one role, one principal, one scope, and nothing else confers permissions. See Admin.

from pinecone.admin import Admin

admin = Admin(client_id="...", client_secret="...")

admin.invites.create(
    email="newhire@acme.com",
    role_bindings=[
        {"resource_type": "organization", "role": "OrgMember"},
    ],
)

created = admin.service_accounts.create(
    name="ci-deploy",
    role_bindings=[
        {
            "resource_type": "project",
            "resource_id": project_id,
            "role": "ProjectEditor",
        },
    ],
)
print(created.client_secret)  # returned exactly once

Assistant operations API. Assistant file writes are asynchronous server-side, and the operations behind them are now first-class and can be listed and described, so a job started with timeout=-1 can be followed rather than inferred from file status. Successes and failures are both kept for 30 days. See Assistant.

operation = pc.assistants.describe_operation(
    assistant_name="my-assistant", operation_id="op-1234-abcd-5678"
)
print(operation.status, operation.percent_complete, operation.error)

stuck = pc.assistants.list_operations(
    assistant_name="my-assistant",
    operation_type="upload_file",
    status="Processing",
).to_list()

Dedicated read capacity. An index can be provisioned on dedicated read nodes instead of on-demand capacity, at create time, on configure, and on a restore. One top-level read_capacity= argument covers managed and BYOC indexes, and it reads back from IndexModel.read_capacity. pc.create_index_from_backup(..., read_capacity=...) applies the same configuration to a restored index in one call. See Backups and restore.

pc.indexes.create(
    name="product-search",
    schema=schema,
    deployment={
        "deployment_type": "managed",
        "cloud": "aws",
        "region": "us-east-1",
    },
    read_capacity={
        "mode": "Dedicated",
        "dedicated": {
            "node_type": "t1",
            "scaling": "Manual",
            "manual": {"shards": 2, "replicas": 2},
        },
    },
)

Bulk ingest deadlines and concurrency. Batched upserts take a total_timeout, a deadline for the whole call rather than for one attempt of one batch. When it expires the SDK stops submitting; batches in flight finish, and everything unsent comes back as failed items to retry. max_concurrency defaults to 8 and is capped by an adaptive per-host admission gate, so a backend under pressure gets backpressure instead of every batch at once. Identical on REST sync, asyncio, and gRPC. See How Bulk Ingest Behaves.

response = index.upsert(
    vectors=vectors,
    batch_size=200,
    max_concurrency=16,
    total_timeout=1800,
)
# Batch counters and errors carry information only when batch_size is set.
print(response.upserted_count, response.failed_item_count)
print([(e.disposition, e.retryable) for e in response.errors])

upsert_from_dataframe built for large ingests. The DataFrame path has the same signature on REST, asyncio, and gRPC, and it now gets the full bulk-ingest toolkit: a per-batch timeout, a max_concurrency ceiling, a total_timeout deadline for the whole job, and on_error to choose between collecting partial failures into the response or raising. max_concurrency is unset by default here, unlike upsert, which defaults to 8. Rows that did not land come back in response.failed_items, ready to wrap in a new DataFrame and feed back in. See Reliable Large Ingests with upsert_from_dataframe.

response = index.upsert_from_dataframe(
    df,
    namespace="movies",
    batch_size=500,
    max_concurrency=8,
    total_timeout=1800,
)
if response.failed_items:
    leftovers = pd.DataFrame(response.failed_items)
    index.upsert_from_dataframe(leftovers, namespace="movies", on_error="raise")

Python 3.14. The SDK is tested and classified on 3.10 through 3.14; the requires-python floor stays at 3.10.

Breaking changes

This is the headline list, not the complete one. Read Migrating to V10 for every change, including the ones that affect only narrow call shapes.

  • pinecone.preview is removed with no shim. Preview, AsyncPreview, and every preview.* module are gone. Each call site maps onto the main client:

    9.x 10.0.0
    pc.preview.indexes pc.indexes
    pc.preview.index(name=...) / (host=...) pc.index(name=...) / (host=...), or pc.index("my-index")
    pc.preview.index(...).documents.upsert(...) pc.index(...).documents.upsert(...)
    pc.preview.close() pc.close()
    from pinecone.preview import SchemaBuilder from pinecone import SchemaBuilder

    AsyncPinecone.index() also became a coroutine — AsyncPreview.index() resolved its host lazily, so it needs an await now.

  • IndexModel is built from schema and deployment. .created_at is gone. .spec, .embed, .dimension, .metric, and .vector_type still work as deprecated views computed from the new fields, including through index["..."] and to_dict(), but they raise when the schema gives them no single field to resolve to — more than one vector field, or none — and the message names the fields it found.

  • Bulk ingest defaults changed. max_concurrency went from 4 to 8 on every upsert path, an adaptive admission gate can hold batches back or abandon the rest of a call against a dead backend, and total_timeout bounds the whole call. Pass max_concurrency=4 for the old cap, and inspect response.errors rather than assuming a call that returned did all the work.

  • GrpcIndex.upsert_from_dataframe no longer raises on partial failure. Failures are collected into the response, so an existing except block around it goes dead. Pass on_error="raise" to restore the previous behavior. Partial failures on gRPC upsert_from_dataframe covers what the response carries and how to retry from it.

  • pc.indexes.list() returns a Paginator[IndexModel], not an IndexList, so .names() is not available on that path. pc.list_indexes() is unchanged.

  • Client-side validation raises before the request is sent. top_k outside 1-10000, a page limit out of range, an id over 512 characters, an empty filter={}, mutually exclusive arguments, and malformed namespace names now raise PineconeValueError or PineconeTypeError locally, where 9.x let the server refuse them. Neither is an ApiError subclass, so an except ApiError handler will not see them; except PineconeError covers both.

  • Assistant file progress fields moved to the operations API. AssistantFileModel.percent_done and .error_message are removed; call pc.assistants.describe_operation(...) and read status, percent_complete, and error.

  • BackupModel.dimension and .metric are removed. Read .schema instead.

  • The default API version header is now 2026-07 on every plane. To pin an older version, set it yourself: additional_headers={"X-Pinecone-Api-Version": "2026-01"}.

Deprecated

These paths still work and are not scheduled for removal in 10.x.

  • dimension=, metric=, vector_type=, and spec= on create_index and pc.indexes.create. Pass schema= and deployment= instead. replicas=, pod_type=, and serverless_read_capacity= on configure are deprecated the same way, translated into deployment= and read_capacity=.
  • IndexModel.spec, .embed, .dimension, .metric, and .vector_type, along with the IndexSpec, ServerlessSpecInfo, PodSpecInfo, ByocSpecInfo, and ModelIndexEmbed classes. Read .deployment and .schema.fields instead; that is where the server reports them and the only place they appear for an index with several vector fields.
  • pool_threads on the client and async_req=True on data-plane methods. Use AsyncPinecone or a ThreadPoolExecutor; these exist for backcompat only.
  • RerankModel.Pinecone_Rerank_V0 is still a member but deprecated, and most projects now get a permission error when they ask for it.

None of these emit a DeprecationWarning yet, so -W error::DeprecationWarning will not find your call sites. Grep for them.

Other improvements

Two new members of the exception tree apply across every surface: PaymentRequiredError (402) and FailedPreconditionError (412), both ApiError subclasses.

Client and connection

  • ssl_ca_certs, ssl_verify, and proxy_url reach the transport that opens the socket, so they take effect.
  • pc.index(name) is overloaded on grpc=, so index.documents type-checks under mypy and pyright.
  • pc.preview raises an AttributeError that names its replacement and links the migration guide, rather than Python's bare message.

Indexes and namespaces

  • Flat create_index, configure_index, and list_indexes still work on top of the schema and deployment API.
  • NamespaceDescription.size_bytes is populated.
  • Namespace names and limits are validated before the request is sent.

Inference

  • Inference calls send the model id, not the string form of the enum, when model is an EmbedModel member.
  • New EmbedModel members for the models available at 2026-07.
  • top_n >= 1 is enforced on the async rerank path, and argument conflicts are reported before enum validation.

Assistant

  • Streaming assistant chat no longer aborts at 30 seconds; with no per-call timeout= the read floor is 300 seconds.
  • Assistant chat stream chunks carry content_filter_results, and StreamMessageStart carries an id.

Packaging and types

  • Nested namespace models and TokenResponse are exported at the top level.
  • anyio>=4.0 is declared as an explicit runtime dependency instead of relied on through httpx.
  • The wheel no longer installs a bare LICENSE file into the site-packages root; it ships in the sdist and in the wheel's metadata.
  • The [dev] extra is removed; use uv sync --group dev.

Security

Four fixes in the SDK's own code:

  • Every REST path parameter is encoded at the transport boundary. A caller-supplied value containing /, or a . or .. segment, previously survived into the request path and addressed a different endpoint instead of failing: describe_namespace(name=".") requested the namespace list route, and describe_namespace(name="..") requested /. Encoding now happens in HTTPClient and AsyncHTTPClient rather than at each call site, so it covers every path the SDK builds. The exposure was bounded to misdirected reads: none of the collapsing paths resolved to a DELETE, PUT, or PATCH route.
  • TokenResponse.__repr__ masks the access token. It previously rendered the live Bearer token in full, so any crash reporter that captures frame locals — Sentry's Python SDK does by default — wrote the credential verbatim into the error report. repr() and str() now keep only the last four characters. The masking stops there: to_dict() and JSON encoding still return the token in full, so do not serialize the object wholesale into a log line or a cache.
  • Admin no longer keeps the raw client_secret where a repr() reaches it. The token minter held the secret as a bound argument of a functools.partial, whose repr() renders the arguments it closed over.
  • A plaintext gRPC data plane warns once per process. If you have set PINECONE_GRPC_SCHEME=http against a host that is neither loopback nor RFC 1918 private, the SDK now emits a RuntimeWarning, because the API key crosses a public network in the clear.

And the dependency tree. As of 10.0.0 no advisory against the SDK's dependencies is outstanding. Three of those upgrades reach code that ships to you:

  • h2 (Rust) 0.4.13 to 0.4.19, for RUSTSEC-2026-0258: the HTTP/2 codec queued empty DATA frames without bound, so a hostile or compromised endpoint could drive unbounded memory growth or a length-overflow panic. It reached the shipped gRPC extension as a runtime transitive, through hyper and tonic, and 9.1.0 carried the same version. If you use the gRPC transport, upgrade for this alone.
  • h2 (Python) 4.3.0 to 4.4.1, for CVE-2026-71554. It arrives with the default install, as a transitive of httpx[http2]. 4.4.1 also rejects duplicate Host headers and conflicting content-length values rather than accepting the first of each.
  • pyo3 0.24.1 to 0.29.2, for GHSA-36hh-v3qg-5jq4 and GHSA-chgr-c6px-7xpp, compiled into the shipped extension. Neither advisory is reachable from this codebase, since the affected APIs have no call sites here; the upgrade went in regardless.

The remaining advisories were against development and documentation tooling — soupsieve and tornado, reached through beautifulsoup4 and ipykernel — and have never been present in a published wheel.

The Rust dependency tree is audited on every push, so an advisory against a crate inside the gRPC extension surfaces on the commit that introduces it, and the PyPI publish step is pinned to a commit SHA.

Upgrading

pip install --upgrade pinecone
# or
uv add 'pinecone>=10,<11'

Then read Migrating to V10.