Skip to content

v0.3.0

Latest

Choose a tag to compare

@senamakel senamakel released this 04 Oct 08:16
Immutable release. Only release title and notes can be modified.
c71e6b5

What's Changed

  • Extract model provider and embedding APIs by @senamakel in #1
  • Carry runtime model selection metadata by @senamakel in #2
  • Add current local-runtime and inference wire contracts by @senamakel in #3
  • Read an explicit null sequence as an empty one by @senamakel in #4
  • Add Anthropic prompt cache controls by @senamakel in #5
  • fix: accept the common reasoning tag names, not just <think> by @M3gA-Mind in #9
  • ci: bump actions/checkout from 4.4.0 to 7.0.1 in the github-actions group by @dependabot[bot] in #8
  • fix: send stream:true on Responses and parse Codex SSE by @felipeinf in #6
  • Make declared prompt-cache prefixes reach the wire (native Anthropic adapter, OpenRouter breakpoints, DeepSeek hits) by @senamakel in #10
  • fix(openai): stream the Responses API for the Codex OAuth backend by @Guykaganovsky1 in #7
  • feat(embeddings): report provider token usage from embedding models by @YellowSnnowmann in #11
  • Split inference into core, LLM, embeddings, and local crates by @senamakel in #12
  • Add provider-neutral model boundary contracts by @senamakel in #13
  • Add provider-specific prompt cache presets by @senamakel in #14
  • chore(deps): drop dependencies no crate references by @senamakel in #16
  • chore: ignore Dependabot patch releases by @senamakel in #17
  • deps: bump base64 from 0.22.1 to 0.23.1 by @dependabot[bot] in #18
  • deps: bump sysinfo from 0.33.1 to 0.38.4 by @dependabot[bot] in #20
  • deps: bump zip from 2.4.2 to 8.6.0 by @dependabot[bot] in #21
  • deps: bump rand from 0.9.5 to 0.10.2 by @dependabot[bot] in #19
  • feat(llm): drive prompt-guided tool calling through tinytools-agent by @senamakel in #15
  • fix(openai): emit prompt-guided stream scrubber's flushed suffix before Completed by @senamakel in #22
  • Remove unused TinyInference provider dependencies by @senamakel in #23
  • Block-indexed streaming, message origin/custom/system patches, profile behaviour, media blocks (tinyagents#171) by @senamakel in #24
  • feat(llm): expose provider request extensions by @senamakel in #25
  • test: cover hosted voice transcription transport by @senamakel in #26
  • feat: tinyinference-image and tinyinference-video (OpenRouter media generation) by @senamakel in #27
  • fix(openai): carry a 429's Retry-After into the provider error by @shaurya703 in #28
  • fix: keep active request after prompt-guided tool results by @senamakel in #29
  • fix: keep active request with prompt-guided tool result by @senamakel in #30
  • feat: add Jev decisions API crate by @senamakel in #31
  • fix(decisions): tolerate rounding when checking a choice is the maximum by @senamakel in #32
  • Harden decisions client transport and probability validation by @senamakel in #33
  • Add Levanto Sage typed decisions and rename decisions crate by @senamakel in #34
  • Add live Sage decision smoke example by @senamakel in #35
  • feat(decisions): add OpenJEV System One provider by @senamakel in #36
  • fix(openai): read a numeric stream-error code as the HTTP status by @M3gA-Mind in #37
  • feat(core): scrub_credentials in sanitize (from OpenHuman) by @senamakel in #40
  • test: port OpenHuman lifecycle, OpenAI request, streaming and cleanup tests by @senamakel in #38
  • feat(llm): expose contains_business_limit for flattened rate-limit text by @senamakel in #42
  • fix(sanitize): make scrub_credentials idempotent by @senamakel in #44
  • feat(llm): provider error-body phrase matchers in failure.rs by @senamakel in #43
  • feat(providers): budget-exhausted message classifier by @senamakel in #45
  • feat: move string-level failure predicates, emoji extraction and voice clients from OpenHuman by @senamakel in #46
  • feat(local)!: drop model downloads and Ollama lifecycle; endpoint-only local inference by @senamakel in #47
  • feat(llm,providers): one budget-phrase matcher with strictness modes; status-word failure predicates by @senamakel in #48
  • fix(llm): depend on tinytools-agent 0.5 so host patches unify the copy by @senamakel in #49
  • test: move inline tests into *_tests.rs files by @senamakel in #50
  • test: rename legacy test.rs and *_test.rs files to *_tests.rs by @senamakel in #51
  • feat(openai): cache_control breakpoints for Anthropic models on any compatible gateway by @senamakel in #52
  • feat(model): discover context windows from providers; learn limits from overflow errors by @senamakel in #53
  • fix(sanitize): stop scrub_credentials redacting ordinary source code by @senamakel in #54
  • feat(openai): send reasoning budget to OpenRouter as reasoning.max_tokens by @senamakel in #55
  • Fix Anthropic thinking precedence with reasoning budgets by @senamakel in #57
  • feat(model): ModelProfile::hoists_system_messages for DeepSeek routes (openhuman#6962) by @senamakel in #56
  • feat(openai): adaptive parameter omission and per-request bearer source by @senamakel in #58
  • feat(embeddings): add resilient voyage reranking support by @Johnson-f in #59
  • Preserve native media and expose transport input support by @senamakel in #61

New Contributors

Full Changelog: https://github.com/tinyhumansai/tinyinference/commits/v0.3.0