Releases: Mo7ammedd/LLMProxy
Release list
LLMProxy v0.3.0 — Provider operations and verified images
LLMProxy v0.3.0 adds provider operations, shared credential coordination and verifiable container releases.
What changed
- Provider-key dashboard and API: add encrypted keys, enable or disable keys, and inspect fingerprints, usage, failures and cooldowns. One provider account supports up to 64 keys.
- Shared key pools: Redis coordinates rotation and cooldowns across replicas. Managed-key changes propagate automatically; active requests retain their captured credentials.
- Optional live checks: validate credentials and model discovery, or run a small generation probe against a configured model.
- Operational alerts: durable budget, error-rate, response-header latency and exhausted-pool alerts, with acknowledgment, recovery, deduplication and optional webhook delivery.
- GitHub protections: required CI checks and resolved conversations on main, administrator enforcement, secret scanning and push protection, private vulnerability reporting and dependency security updates.
- Verified releases: vulnerability scans for both architectures, Sigstore signatures and scan attestations in both registries, SPDX SBOMs, build provenance and checksums. The README, OpenAPI, deployment examples and operations guides cover the new behavior.
Upgrade from v0.2.0
- Drain existing replicas and back up the database and configuration. Apply the additive
ProviderOperationsmigration once for SQLite or PostgreSQL before restarting replicas when automatic migration is disabled. - Existing configured single keys and key arrays remain compatible. To add keys through the dashboard, set
LLMPROXY_PROVIDER_KEY_ENCRYPTION_KEYto one base64-encoded 32-byte key, identical on every replica. Back it up separately; changing it requires re-encrypting stored credentials. - Review alert thresholds. Live provider checks and webhook delivery are opt-in. Confirm readiness, existing traffic, provider-key status and alerts after deployment.
Complete upgrade notes · Provider operations and configuration
Images and verification
docker pull mohammedtv/llmproxy:0.3.0- Docker Hub:
mohammedtv/llmproxy:0.3.0(public). - GHCR:
ghcr.io/mo7ammedd/llmproxy:v0.3.0(authentication with package read access required). - Platforms:
linux/amd64andlinux/arm64. - Verified index digest in both registries:
sha256:aad94fb63c556702115cfa6d624b487f443d3c18d9c39215f9c0b9e09c924473. - Source commit:
1fa11c5. - Successful release CI.
The release passed 347 .NET tests with PostgreSQL and Redis enabled, the recovery drill, native AMD64/ARM64 container and SDK checks, vulnerability gates and signature/attestation verification. Automated provider traffic uses mocks; live account access and model availability depend on your credentials and configuration.
The vulnerability gate rejects HIGH or CRITICAL findings with published fixes. Full reports also include lower-severity and unfixed findings. Download the attached evidence files together and run sha256sum -c SHA256SUMS. Use Cosign verification to verify origin with this workflow identity:
https://github.com/Mo7ammedd/LLMProxy/.github/workflows/ci.yml@refs/tags/v0.3.0
LLMProxy v0.2.0 — Provider key pools and production gateway features
LLMProxy 0.2.0 adds multiple upstream keys, six provider adapters, expanded inference APIs, durable batches and the administration console, with shared quotas and concurrency for production deployments.
Container images
docker pull mohammedtv/llmproxy:0.2.0- Docker Hub:
mohammedtv/llmproxy:0.2.0 - GHCR:
ghcr.io/mo7ammedd/llmproxy:v0.2.0 - Platforms:
linux/amd64andlinux/arm64 - Source:
1911fc881e39bbe033a7fa57a4d7ce9573a05c6c - Docker Hub digest:
sha256:a8bd1b3e93b12549733b86fbb0be9809704924e2a2e27bee075df4c7ab56d3db
The Docker Hub image preserves the verified build published before the Git tag was added. GHCR's tagged build uses the same source and may have a different digest because of build metadata. The release workflow passed solution/database tests, recovery checks and native container/SDK checks on both architectures.
Quick start · Complete changelog · Release and upgrade guide
Changes and upgrade notes
Added
- Six provider adapters: Microsoft Foundry, Mistral, Cohere, DeepSeek, Groq and Ollama, bringing the total to ten. Foundry supports API keys and Entra ID; trusted local Ollama endpoints have an explicit HTTP opt-in.
- Up to 64 API keys per provider account, rotation, authentication/rate-limit failover, cooldowns, separate credential circuits and credential fingerprints in attempt reports. Named accounts support separate endpoints, prices and concurrency scopes for the same adapter.
- Adapter/model capability checks, feature-combination validation, context limits, cost and latency routing, and explicit reload of models, pricing, accounts and key pools.
- Embeddings, stateless Responses on OpenAI/Foundry, supported image/audio inputs and reasoning controls, plus durable files and batches with worker claims, cancellation, expiry and partial results.
- Local administrator/operator/auditor identities, expiring sessions, an embedded
/adminconsole and management/rejection audit records. - Gateway key expiry and rotation with optional grace, UTC monthly token/USD allowances, shared concurrency admission and request deadlines.
- Usage charts, filtered/cursor reports and CSV, ordinary/cache/tier pricing, individual upstream attempts and idempotent invoice reconciliation.
- Usage/audit/batch retention, SQLite and PostgreSQL migrations, load and multi-instance recovery drills, Redis-loss checks, database restore checks and native ARM64 container/SDK CI.
- Public Docker Hub images for
linux/amd64andlinux/arm64, with multi-platform digestsha256:a8bd1b3e93b12549733b86fbb0be9809704924e2a2e27bee075df4c7ab56d3dbformohammedtv/llmproxy:0.2.0.
Compatibility notes
- Single provider keys continue to work. A nonempty
ApiKeyslist replaces the single key; it does not append to it. Rotation and cooldowns are local to each gateway process. - Responses is stateless. Files serve gateway batches, which execute ordinary requests at ordinary configured prices. See the protocol boundaries for unsupported stateful, hosted-tool and media-generation operations.
- Role scopes apply across the gateway. OIDC/SSO, MFA and owner-scoped operator permissions remain follow-up work.
- Capability validation rejects unsupported request combinations before provider execution. Custom model catalogs need suitable capabilities, aliases, upstream mappings and prices to expose the new protocols.
- Batch inputs and outputs persist request/response content. Include them in backup, access and retention planning even though ordinary usage logs omit bodies.
Upgrading from 0.1.0
- Record the deployed image digest and configuration. Drain or stop the old gateway replicas, then back up PostgreSQL or the complete SQLite data directory using the backup guidance.
- Set
LLMPROXY_IMAGE=mohammedtv/llmproxy:0.2.0in the Compose.env, or select that image in your deployment. Keep the existing data volumes, connection strings, Redis namespace and provider credentials. - Apply the new
GatewayExpansionandProviderKeyPoolsmigrations for your storage engine. Automatic migration remains available; for controlled deployments, run the new image'smigrateCLI once before starting gateway replicas with automatic migration disabled. See migrations and upgrades. - Review custom JSON catalogs against the new defaults. Add the aliases/mappings/prices for embeddings or Responses if needed. Existing single-key environment settings can remain; explicitly forward any new key-pool variables through a Compose service override.
- Start the new gateway, check
/health/readyand/v1/models, then verify existing keys, limits and application requests. Use the separate administrator bootstrap secret to create a local administrator through/adminif you want operator accounts.
Treat rollback as restoration of the pre-upgrade database backup together with the old pinned image and configuration. Do not assume the older binary supports the migrated schema. Restore/recovery commands and the validation limits are documented in the operations guide.
Verification
The release source is commit 1911fc881e39bbe033a7fa57a4d7ce9573a05c6c. Its successful source CI run ran 327 tests with PostgreSQL/Redis, a recovery drill and native AMD64/ARM64 container and SDK checks. Provider calls in automated checks use deterministic mocks; account access, live model capabilities and prices remain deployment-specific.
LLMProxy v0.1.0
LLMProxy v0.1.0 is the initial MIT-licensed, self-hosted LLM gateway built with C# and .NET 10.
- OpenAI-compatible chat completions, model listing, function tools and SSE streaming.
- OpenAI, Anthropic, Google Gemini and Azure OpenAI adapters with configurable model aliases, routing, retries and fallback.
- Hashed gateway API keys, distributed rate limits, token and spending allowances, durable usage and estimated cost tracking.
- PostgreSQL and Redis deployment, persistent standalone mode, health checks, OpenTelemetry and graceful shutdown.
- Non-root Docker images for Linux amd64 and arm64, Docker Compose, API documentation and Python, TypeScript and C# client examples.
Validation: 87 automated tests, plus Docker deployment and OpenAI SDK smoke tests. Automated inference requests use local mock providers; Docker runtime checks execute on amd64.
Images: ghcr.io/mo7ammedd/llmproxy:v0.1.0, v0.1 and latest.
The repository and container package are private. Authenticate to GHCR before pulling. Follow the deployment guide to configure credentials and create a gateway key.
This MVP supports text and function tools with one completion choice. Multimodal requests and the Responses API are outside this release. See the API compatibility documentation for the full supported surface and provider differences.