Skip to content

Releases: Mintplex-Labs/anything-llm

AnythingLLM v1.16.1

Choose a tag to compare

@timothycarambat timothycarambat released this 27 Aug 16:42
35c58d8

Activity and Chain of Thought Overhaul

image

We overhauled our UI for long-horizon agentic tasks that show multiple thinking and tool calls to now roll up into a single clean, collapsible UI element.

This gives you more screen real estate for the actual response while still keeping the important call stack and thoughts the model produces in an easy-to-read format.

Navigation Warning

image

Recently, we added true abort signals across every supported LLM so that when you cancel inference or leave the page, we stop inference to save you on compute costs.

However, this had the unintended side effect of canceling inference when navigating away from a chat, preventing background completion. Now we warn you before aborting the response.

Roadmap item: We intend to make this user configurable so you can customize this behavior and allow background inference for those with more powerful or high-concurrency setups where multiple parallel inference sessions are feasible.

Foundry Local Embedded (Desktop Only)

image

FoundryLocal is an inference engine built by Microsoft that is pre-installed and available on Windows 11 (version 24H2, build 26100 or later).

This uses WinML to load models for your CPU, GPU, or NPU across any Windows hardware configuration. It's a great option for hardware configurations we don't have explicit support for. Keep in mind that model availability may be limited as FoundryLocal is still rolling out. Vision models are not available currently.

We partner with Microsoft for this, if you have bugs or issues we can help forward them to the FoundryLocal team

GenieX for Snapdragon (Desktop Only)

image

Note

This replaces our previous Snapdragon NPU engine. Any previous models from the old engine are deleted on update, and you'll need to re-download these optimized models. You should see a significant performance increase for NPU models and regular GGUFs.

We highly recommend using only GGUF models for this engine. Qualcomm NPU-only models have 4K context windows and will result in a poor agent experience.

GenieX is an open-source inference engine built by Qualcomm (previously NexaAI) that efficiently runs GGUFs and Qualcomm AI Hub models on NPU for X and X2 Elite devices.

AnythingLLM offers fully managed and built-in support by selecting the "AnythingLLM GenieX" provider in the dropdown. Both GGUFs and NPU-only models can leverage the NPU, giving you the power efficiency and intelligence of both.

We partner with Qualcomm for this, if you have bugs or issues we can help forward them to the GenieX team or you can make an issue on their GitHub. s/o @zhiyuan8, @alexchen4ai & @alanzhuly

New Features

  • Native Foundry Local SDK on Windows (x64/ARM64) - the Foundry provider is now fully self-contained and no longer requires a separate Foundry Local install, with GPU/NPU-optimized model variants exposed in the catalog
  • GenieX runtime support for Windows ARM64 devices
  • AWS Bedrock cross-region inference profile support - geo-prefixed profile IDs (us., eu., global.) now appear in the model dropdown and route correctly, plus manual region entry for regions outside the preset list (GovCloud, specialized partitions)
  • LocalAI added as an image generation provider, listing only image-capable models from your server
  • LocalAI context window auto-detection - context windows are read per-model from your server's config, with the manual setting now an optional override
  • Silent/headless install and uninstall flags for the Windows installer - see the Windows installation docs

Improvements

  • Chain of thought and agent activity are now a single collapsible component - one line when collapsed, one step per thought with a connected rail when expanded, and the live thought shown in the header while streaming
  • The new chain of thought rollup is now mirrored in the Assistant Panel
  • Leaving a thread or navigating away mid-inference now warns you before the response is lost, instead of silently killing the generation
  • Foundry Local no longer prompts for setup when using the built-in runtime, and auto-loads models by capability detection
  • Onboarding light mode styling and a redesigned LLM selection UI
  • Beacon now renders thinking blocks and markdown formatting correctly while streaming, with collapsible thinking and a fix for model detail text overflow
  • Windows installer now includes a sidebar panel image
  • Prompt input drafts now persist between navigation

Bug Fixes

  • Fixed a Gemini crash caused by a content-header length mismatch on unpinned undici versions - now compatible with both v6 and v7
  • Fixed a crash when aborting an Anthropic (or Bedrock Anthropic) response mid-stream in chat mode and then sending a follow-up
  • Fixed a memory leak from leaked abort-stream listeners, which also caused pausing a response to wipe the entire chat
  • Fixed a crash when creating a new workspace
  • Fixed missing setup CTAs on modals introduced by the new uniform modal component - Scheduled Job skill setup on both the category and tool views, Community Hub connection key, and the experimental features reject button
  • Fixed the intermediate loader being too tall while waiting for an agent capability response
  • Fixed the stray border around the delete button in Workspace Chats
  • Fixed reasoning getting stuck on "thinking" when a model never closes its think tag
  • Fixed agent websocket sends firing while the connection was still opening

All Changes

New Contributors

Full Changelog: v1.16.0...v1.16.1

AnythingLLM v1.16.0

Choose a tag to compare

@timothycarambat timothycarambat released this 13 Aug 21:38
55b6ebc

Notable Changes

Image Generation

Now you can generate images via /img when a supported provider is set up.

This flow also allows image attachments for edits, followups, or combination prompts if your provider supports image editing.

image

Ability for an agent-tool to generate images is coming next update.

Better File Picker UX

Now in the file picker you can drag entire folders and their hierarchy will be preserved in the upload panel. This will only import the top level files in the folder — it will be deep and recursive in a later update.

You can also now drag and drop a file into an existing folder directly without needing to move it.

Lastly, you no longer need to type https in the URL fetcher - it will just assume https when left out.

We also made some large performance improvements, like lazy-loading of folders, for those with thousands of documents in a folder.

Tool Toggle Mid-Session

Now, during agentic chats you can freely toggle tools on/off at will without needing to reload or start a new session. Only tools that have been configured (if configuration is required) are available.

Stop Generation

A large improvement is now that aborting a response actually fully kills the inference as well so that ghost inference does not continue when no longer wanted mid-stream.

All Changes

New Contributors

Read more

AnythingLLM v1.15.0

Choose a tag to compare

@timothycarambat timothycarambat released this 25 Jun 23:25
70e0d2e

AnythingLLM Is Now An AI Agent Across Your Os

With AnythingLLM 1.15.0 for Desktop we have been working hard on what it means to bring your agent to you. Something that is still local, but outside of the walls of an app or a browser. This release features our first 3 features about this effort - we hope you enjoy them.

Introducing Magic Features

Magic Features bring AI to your entire computer — not just inside AnythingLLM. Dictation, text actions, and autocomplete that work in any app, fully on-device.

All Magic Features are free to use — no signup required. Pro removes the daily limits.

AnythingLLM Pro

Nothing about AnythingLLM is changing. Pro is purely additive — no existing features are affected, nothing is being locked away.

Every Pro feature will always have a free daily tier, no signup required.

You can read more about what AnythingLLM Pro is here.

Magic Echo - Docs

A smarter voice-to-text dictation that works anywhere on your OS. Can replace tools like SuperWhisper or WhisprFlow entirely. Fully on-device.

Speak naturally and your words appear right where your cursor is — transcribed, cleaned up, and punctuated. Echo can see what's on your screen to make dictations smarter and more contextual.

Includes custom dictionary support, voice commands, and more.

MagicEcho-promo.mp4

Magic Beacon - Docs

Highlight text in any app and instantly act on it with AI.

Highlight any text on your screen and instantly act on it — summarize, translate, rewrite, research, or run a custom action. Beacon works in any app without switching windows.

It also has full access to your agent skills, MCPs, and tools — AnythingLLM's entire capability set, available anywhere your cursor is.

MagicBeacon-promo.mp4

Magic Tab - Docs

Grammarly across your entire computer - fully on-device.

As you type, Magic Tab suggests what comes next — in any app, aware of what you're working on so suggestions actually fit. Click into a text field and it'll suggest something before you've typed a single letter.

If you use Grammarly, Tab can replace it entirely — privately, on your device.

MagicTab-promo.mp4


What's Changed

New Contributors

Full Changelog: v1.14.2...v1.15.0

AnythingLLM 1.14.2

Choose a tag to compare

@timothycarambat timothycarambat released this 22 Jun 20:59
5862a3e

Just some small patches for the upcoming 1.15.0 / 2.0.0-preview initial release for desktop

What's Changed

New Contributors

Full Changelog: v1.14.1...v1.14.2

AnythingLLM v1.14.1

Choose a tag to compare

@timothycarambat timothycarambat released this 16 Jun 21:17
619868f

Meeting Assistant Overhaul

Meeting Assistant is desktop app only

We have overhauled a large portion of the Meeting Assistant to make it smaller, faster, and more efficient across all devices and platforms.

  • Now supports Intel, AMD, and NVIDIA GPUs for a 92% smaller binary and 15% faster processing times.

    • If you already have the NVIDIA GPU binary installed, you can safely delete it if you want. It will still work and is backwards compatible.
  • Support for Developer API for transcription on audio (POST: /v1/transcription/transcribe)

  • Meeting Assistant context window overflow handling is much better now - so small models can summarize longer meetings.

  • Introduction of Basic Speaker Identification for 60% better summarizes from any audio.

  • Dual channel stero recordings for meetings now - leading to 80% better speaker identification in "Full Diarization" mode.

Improvements

  • Linux AppImage now 91% smaller in size and caches Ollama engine downloads for faster startup times.
  • Meeting Assistant title fix on meetings post-summary now auto-updates in UI
  • AgentFLow variable highlight so its clear what is and is not a valid variable
  • "Copy chat link" in UI to quickly re-open a chat in the desktop/self-hosted app via deeplinking.
  • Re-enabled audio and video uploads via chat UI - uses Tinyscribe engine now.
  • Export Chat as (PDF, JSON, Markdown, etc) from chat UI.
  • Desktop Assistant Setting - HD screenshots now available for screenshot capture area.
  • Request approval internal function is now available for custom skills.

Bug Fixes

  • Removed DPAIS and HuggingFace providers from AnythingLLM (unmaintained)
  • Fixed memory leak in embedder from it constantly reloading in server process
  • Fixed text clearing bug when dragging and dropping files into the chat and text was already present in prompt.
  • Massive performance improvements to the frontend UI for long running chats.
  • Cohere SDK removed and ported to OpenAI SDK for compatibility.
  • Desktop Assistant Capture Area not showing on windows multi-monitor setups.
  • Strip thinking from fork thread name when forking a chat that had thoughts.
  • Fix toast light mode always showing regardless of system theme.
  • Mistral embedder encoding issue fixed.
  • Better error messages for API
  • Omit temp in Claude Bedrock for Claude 4.8
  • Fixed event emitter leak in server process for web-scraping and summarize process

What's Changed

New Contributors

Full Changelog: v1.14.0...v1.14.1

AnythingLLM v1.14.0

Choose a tag to compare

@timothycarambat timothycarambat released this 09 Jun 13:46
5706b42

Improvements

  • Cerebres provider
  • The default chat thread is now killed when you create a new thread. If you have chats on the default thread, it will be available still. New workspaces or workspaces with no chats on default will no longer show it.
  • All model providers are now opt-out of tool calling by default. Everything will call tools by default unless you opt-out offering better performance for agents everywhere
  • STT Support for Deepgram, GenericOAI, Lemonade, & OpenAI
  • TTS Support for KokoroTTS
  • Web-scraping now will convert to markdown for better parsing and chat followup tasks with minimal context bloat
  • Summary tool was overhauled. Now it will so better summaries with transparency as well as ask before continuing for longer summaries
  • Improvements to the GenericOAI provider
  • 24hour system variable formats
  • Better LaTex rendering support

Bug Fixes

  • Context limit detection issue for agents: #5716
  • SEARXNG double encoding: #5723
  • Timeouts for all fetch requests #5721
  • Escape illegal XML in word docs, etc #5760
  • (Windows) On unisntall, checkbox to remove all AnythingLLM data is now present
  • Tray Fixes when app starts in background or Desktop Assistant feature is toggled.

What's Changed

New Contributors

Full Changelog: v1.13.0...v1.14.0

AnythingLLM v1.13.0 - A Hybrid AI Experience

Choose a tag to compare

@timothycarambat timothycarambat released this 26 May 19:08
9fe6bbf

This release is focused on improving the agent experience and adding new features to the agent system as well as moving towards a more passive, personal, and hybrid AI experience.

Model Router: The First Consumer Hybrid AI Experience

The Model Router feature is the first-ever user-defined intelligent routing system that seamlessly blends local and cloud AI into a single, unified experience that is entirely under your control. Until now, you had to choose: run everything locally, or send everything to the cloud. That tradeoff is over.

With Model Router, you define the rules. Every message you send is automatically analyzed and routed to the perfect model for that specific task, whether that's a lightweight local model for quick questions, a reasoning model for complex math, or your most powerful cloud model for nuanced legal analysis. All from the same chat. All invisible to the user. All defined by you.

What makes this so exciting:

  • Hybrid AI. Mix and match local models (Ollama, LM Studio, etc.) with cloud providers (OpenAI, Anthropic, Google) in a single conversation. No manual switching!
  • You're in complete control. Create calculated rules that trigger on keywords, token counts, time of day, or image attachments instantaneously. Or use LLM-classified rules that understand intent in plain English.
  • Save money without sacrificing quality. Route simple queries to cheap or local models. Reserve expensive API calls for the messages that actually need them.
  • Intelligent caching. Our advanced sticky routing system keeps you on the same model during a conversation thread, so you're not bouncing between models on every message.

This is, we believe, a fundamental shift in how AI assistants work. For the first time, you get the privacy of local models, the power of cloud models, and the intelligence to know when to use each. And it's 100% open source.

Learn how to set up your first router →

chat-routed

Scheduled Jobs: Your AI That Works While You Don't

What if your AI assistant could work for you in the background, automatically, on a schedule you define, without you lifting a finger?

Scheduled Jobs turns AnythingLLM into an always-on AI workforce. Create recurring tasks that run themselves: morning briefings, weekly reports, data monitoring, research digests. Anything you'd normally ask an agent to do, but automated and hands-free. Has job specific skills you can set so the model is not overwhelmed with tools and outcomes are repeatable.

Why this changes everything:

  • Set it and forget it. Define a prompt, pick your tools, choose a schedule, and walk away. Your agent runs exactly when you need it: every morning at 8 AM, every Monday at noon, every hour on the hour.
  • No technical knowledge required. Our visual Cron Builder lets you schedule jobs with simple dropdowns. No cryptic cron syntax, no command line, no code. Just point and click.
  • Full agent power, fully automated. Scheduled jobs have access to the same tools as your regular chats: web search, document analysis, custom skills, MCP integrations, and more. If an agent can do it in a conversation, it can do it on a schedule.
  • Complete run history. Every execution is logged with the agent's full reasoning, tool calls, generated files, and final response. Review past runs anytime, or continue where the agent left off in a new thread.
  • Push notifications. Get alerted the moment a job finishes, even when AnythingLLM is in the background. Click to jump straight to results.

Enterprise tools charge thousands for this kind of automation. Cloud-only platforms require you to trust your data to third parties. AnythingLLM gives you scheduled AI agents that run entirely on your machine, with your data, under your control.

Wake up to a summary of overnight emails. Get weekly progress reports written automatically. Monitor websites for changes. The possibilities are endless, and it all happens while you focus on what matters.

Learn how to create your first scheduled job →

scheduled-job

Automatic Memories & Personalization

AnythingLLM now supports automatic memory extraction and personalization so your AI assistant can remember what you've talked about and use that knowledge to personalize its responses.

AnythingLLM runs a background job to extract memories from your chats and store them in a memory bank. This memory bank is then used to personalize the responses of your AI assistant - you have full control over what is remembered and how it is used
you can even add memories manually to the memory bank if you dont want to have the model spend cycles reviewing chat history.

There are two types of memories:

  • Workspace memories: These are memories that are specific to the current workspace (like what you are working on, projects-specific information, etc.)
  • Global memories: These are memories that are specific to the entire AnythingLLM instance (like your name, preferences, etc.)

Memories are injected into the system prompt of your AI assistant so it can use them to personalize its responses and are a welcome addition to your AI assistant's knowledge base.

Learn how to enable and manage memories →

memories-sidebar

Agent Surveys (special tool)

Agent Surveys is a special tool that allows your AI assistant to ask clarifying questions before proceeding. This is useful when you are working with a complex task and the agent needs more information to proceed.

This is off by default and must be enabled in the agent settings. Answers to the questions are saved alongside the chat message so the agent can use them in future turns.

Learn how to enable and manage agent surveys →

multi-choice

What's Changed

Read more

AnythingLLM v1.12.1

Choose a tag to compare

@timothycarambat timothycarambat released this 22 Apr 22:41
f144692

Notable Improvements

Streamed Document Embedding

Now, when you upload a document to the workspace the process per-document is now reported during embedding. This is a huge improvement in performance and user experience. During this process you can add and remove documents to the queue as well as even close and navigate away from the page without losing your progress.

queue-embedding

App integrations

There are now built in integrations for the following apps with minimal to zero setup required for Agent skills:

Other Improvements

  • Image Lightbox in main UI
  • Enabled Korean, Chinese, & Japanese character support for PDF generation via custom mdpdf fork
  • Better citations for app integrations
  • DDG default web-search in agent skills
  • Open documents in native application on machine when generated by Document Generation Agent
  • Auto approve agent skill via ENV setting
  • Ollama bumped to 0.20.7 (Qwen3.5 support, Gemma 4, etc)
  • New Customization > Chat setting for Unload model when closed to unload the model when the user closes the chat window.
  • Generic OpenAI Capability detection/ENV setting
  • Update Lemonade to support 1.10.0 changes
  • Catalan translations
  • Name field added to API keys
  • Chat ID reported in agent sessions so now you can regenerate chats, TTS, and more actions without page reloads.

What's Changed

New Contributors

Full Changelog: v1.12.0...v1.12.1

AnythingLLM v1.12.0

Choose a tag to compare

@timothycarambat timothycarambat released this 02 Apr 20:56
f6f1c80

Major Features

Automatic Mode for native tool calling

For Select providers that support native tool calling, you no longer need to use @agent to use tools. You can now just use the tools without asking.

If your prompt input does not have the "@" symbol, your chats will automatically use tools as needed.

docs-agent-example.1.mov

Intelligent Tool Selection

We have added a new feature called Intelligent Tool Selection. This feature allows you to load unlimited tools for your agent to use into context with better performance and save up to 80% on token usage every single chat.

settings-menu-icon-location

Filesystem Agent

We have added a new feature called Filesystem Agent. This feature allows you to use the filesystem of your host machine to search for files and directories.

fs-main-panel

Document Generation Agent

We have added a new built-in agent for Document Generation. With document generation, you can generate text files, PDFs, Excel files, Docx, and even entire PowerPoint presentations.

docgen.1.mp4

Telegram Bot

AnythingLLM Docker and Desktop now support a Telegram bot so you can connect to your AnythingLLM instance anywhere in the world.

Supports:

  • Text chat (streaming & thinking)
  • Image understanding
  • Voice messages & Attachments
  • Automatic mode and @agent support
  • Workspace and thread selection
  • Model selection
  • Citations
  • Any agent skill available in AnythingLLM
image-understanding

What's Changed

New Contributors

Full Changelog: v1.11.2...v1.12.0

AnythingLLM v1.11.2

Choose a tag to compare

@timothycarambat timothycarambat released this 18 Mar 16:57

More UI Improvements

changelog-1.11.2-uiv2.mp4

Now, in the main chat UI we added some much desired UI improvements and fixes.

  • New prompt input
  • Better Citations UI and reporting
  • Metrics for Agent calls
  • Report document and web-search citations during Agent calls!
  • Ability to each toggle on/off Agent skills from the prompt
  • Ability to select the provider and model for the workspace without leaving the page.

What's Changed

New Contributors

Full Changelog: v1.11.1...v1.11.2