Skip to content

v0.16.1

Choose a tag to compare

@github-actions github-actions released this 15 Jul 15:32
accf9c6

v0.16.1

Real token-level streaming for the Claude Max subscription backend.

Fixed

ClaudeCodeChatModel._astream (langchain-claude-code 0.1.0) requests include_partial_messages=True from the Claude Agent SDK -- which makes the subprocess emit granular StreamEvent text deltas -- but the method only ever consumed the terminal, whole-block AssistantMessage, silently dropping every delta. A subscription-backed turn arrived as one or two large lumps instead of a real token stream.

_SubscriptionChatModel._astream (_claude_cli.py, alongside the existing _build_options / _wrap_langchain_tool overrides for other upstream gaps in this same package) now consumes StreamEvent text deltas and yields each one immediately as it arrives, tracked per content-block index so the terminal AssistantMessage never re-yields (and thereby doubles) text a delta already streamed. A block that produces no StreamEvent at all (older CLI build, future SDK regression) still gets its text emitted whole from the AssistantMessage -- strictly additive, never worse than before.

Verified against a real Claude Max subscription session: a response streamed in 13 chunks over ~10.6s (visible incremental delivery), versus 1-2 chunks arriving all at once under the prior behavior.

Full Changelog

v0.16.0...v0.16.1