Skip to content

v1.32.0 — hosted-agent workspace overhaul + liteparse + per-user quota knob

Choose a tag to compare

@Exzentttt Exzentttt released this 03 May 00:42
· 402 commits to main since this release

Hosted-agent workspace got a major cleanup, agents can now parse PDFs and Office docs out of the box, and operators can lift the per-user hosted-agent quota without rebuilding the backend image.

What was wrong

Owners watched files vanish from their hosted agent's workspace. An agent would write something like .deep/contacts/note.md, or rely on the seeded .deep/memory/main/MEMORY.md, and on the next sync the file was gone from the workspace tree. Many owners assumed their agents were silently losing memory between sessions.

Hosted agents could not parse PDFs, DOCX, XLSX, PPTX, or images from inside their sandbox — even though the underlying framework already supported a parser tool.

The per-user hosted-agent count was hard-coded to 1. Raising it on a self-hosted install required editing source and rebuilding the backend image.

There was also a fresh-install regression: every newly created hosted agent crashed with HTTP 503 on its first chat message.

Root cause

The runner's /files endpoint (the workspace listing API the backend polls) hid every directory whose name started with a dot. Combined with the backend's prune-on-sync logic — the synchroniser deletes any database row not present in the latest runner listing — this meant the bootstrap-seeded .deep/memory/main/MEMORY.md and any agent-authored .deep/* paths got wiped on every sync. The "hidden by default" model was wrong: only .deep/checkpoints, .deep/plans, .deep/todos.json, .venv, __pycache__, and node_modules are truly volatile noise. Everything else under .deep/ was meaningful agent state.

The parser tool (liteparse, the framework's PDF/Office/OCR toolset) ships as a Python extra plus a Node CLI from @llamaindex/liteparse. The runner image was built on python:3.12-slim with no Node.js installed and no liteparse extra pulled in, so the framework's _HAS_LITEPARSE capability flag stayed False and the toolset returned empty results at runtime.

The 503-on-first-chat regression came from a brief intermediate fix that emitted include_liteparse: true into the persisted agent.yaml. The framework's DeepAgentSpec model rejects unknown keys (extra='forbid'), so loading the new spec raised a Pydantic validation error and the chat endpoint surfaced it as a 503.

Fix

  1. Workspace path layout: the .deep/ prefix is gone. Memory now lives at memory/MEMORY.md, checkpoints at checkpoints/, plans at plans/, todos at todos.json. Changes in agent-runner/main.py and backend/app/services/hosted_agent_service.py. A one-shot migration rewrote both the database (agent_files.file_path) and the runner-mounted on-disk workspaces using cp -n, so no agent lost data.

  2. Visibility model: the dotted-directory blacklist is replaced with an explicit noise-skip set (.venv, __pycache__, node_modules, .git). The _is_visible helper in agent-runner/main.py is removed; the /files walker now returns everything else. The workspace UI shows whatever the agent actually wrote.

  3. Liteparse toolset: the runner Dockerfile now installs Node 20 and npm install -g @llamaindex/liteparse, and pyproject.toml pulls the pydantic-deep[liteparse] extra. The runner enables the toolset by passing include_liteparse=True through DeepAgent.from_file overrides — the kwarg routes to the framework's create_deep_agent passthrough rather than into the strict spec model. agent.yaml no longer carries the flag, so DeepAgentSpec(extra='forbid') validates cleanly and the 503 regression is gone.

  4. Hosted-agent quota knob: MAX_HOSTED_AGENTS_PER_USER is now a runtime environment variable wired through deploy/docker-compose.prod.yml. Default stays at 1. To raise it on a self-hosted install, set the value in .env.prod and restart the backend container — no image rebuild required.

Tests

282 backend pytest tests, 106 agent-runner pytest tests. Two pre-existing backend failures unrelated to this release (test_list_issues_returns_200_not_500, test_commits_endpoint_passes_branch_param) and two pre-existing agent-runner collection errors (test_checkpoints.py, test_history_integration.py — Pydantic validation error at import time) remain and will be triaged in a separate ticket. End-to-end smoke check on production: a fresh hosted agent successfully called write_file and liteparse in one chat turn, both tools fired, the resulting file appeared in the workspace tree, cleanup completed normally.

What owners need to do

Nothing.

If you want to lift the per-user hosted-agent quota on your own install, set MAX_HOSTED_AGENTS_PER_USER=N in .env.prod and restart the backend container.

Followup fix (commits 585d737, 6839e80)

The /chat/stream endpoint returned tool_calls[].result as an empty field even after the tool had visibly executed. When a hosted agent uses the execute tool (or any tool routed through DeferredToolRequests), the streaming runner splits execution into two phases: an agent.iter() pass that yields tool-call events, followed by a separate agent.run(deferred_tool_results=...) that actually runs the tools and generates the final reply. The new_messages() slice from the deferred run contains only the ToolReturnPart result and the final text — no ToolCallPart — so _extract_response returned extra_tools = [] and the streaming-captured all_tool_calls entries never received a result field. A second compounding issue: the streaming loop added a duplicate entry for the same tool before the deferred loop added its own, causing the dedup in the final merge to pick the unpatched (no-result) entry first.

Fix: after each deferred agent.run() call, scan result.all_messages() for ToolReturnPart and patch every matching entry in all_tool_calls by tool_name (forward iteration, no early break). The tool_result stream event is also now emitted after execution with the actual output text instead of an empty string.

Stability + UX followups (commits 25b5c59, 6792369, f42d7c5, a519ccb, 412f9d6, 1e6c794)

A second wave of fixes addresses problems owners reported after the initial release: a disappearing chat message, a chat thread bloated with heartbeat noise, and a routing endpoint that returned a server error for a perfectly normal request.

The most visible bug: a user would type a message, hit send, and watch their own message vanish from the conversation a second later. The cause was in the frontend message loader — loadMessages called data.reverse() on the array returned by the chat history fetch, but that array was shared with React's state cache. Reversing in place mutated the cache, so the next render saw an empty or scrambled message list. The fix replaces the in-place call with [...data].reverse(), a non-mutating copy, in frontend/src/app/agents/[id]/chat/page.tsx and the three other chat-style pages that had the same pattern.

The second user-visible symptom: hosted-agent chat threads were filling up with Heartbeat: idle (5m) rows, sometimes hundreds per day per agent. These belonged in the operational activity log, not in the conversational history shown to the owner. The backend was writing them to owner_messages on every heartbeat. The fix routes them solely to agent_activity and adds a 10-second dedup window on identical user messages so that a double-click on the send button no longer creates two adjacent identical rows. A one-shot SQL cleanup removed 226 leftover heartbeat rows and 45 duplicate user messages already in the table.

Third: hitting GET /hosted-agents/available-models returned HTTP 500 instead of a model list. The route was declared as /{hosted_id} greedy-matching any string, so available-models was treated as a UUID, the parse blew up downstream, and the response was a stack trace. Twenty-four path parameters across the hosted-agents routes are now declared as hosted_id: UUID, which makes Pydantic reject non-UUID slugs with a clean 422 before the handler runs.

A larger refactor under the same theme: agent_service.heartbeat had grown into a 200-line method mixing message persistence, activity logging, runner sync, and metrics. It's now decomposed into eleven named private methods, each handling one phase, with all imports hoisted to the top of the module. Behaviour is unchanged; the function is now possible to read top to bottom.

On the testing side, the agent-runner integration suite for checkpoint and chat-history endpoints had been failing because the test client never sent the RUNNER_KEY header that the runner now requires, and because the patch target for the git service was pointing at an old module path. Both are wired correctly in conftest.py, restoring the suite to green.

Infrastructure: outbound traffic from sandbox containers to internal platform services was previously reachable through the Docker bridge. Host-level firewall rules now block sandbox sources from reaching internal IP ranges on private service ports, while still permitting outbound LLM API calls to the public internet. Verified end-to-end: a sandbox can reach the upstream model provider, but can no longer hit the runner control plane or the database directly. Separately, six leftover test agents from earlier debugging sessions and eleven orphaned workspace directories were removed from the production database and disk.

The full 21-check production regression suite (smoke, auth, registration plus heartbeat, hosted-agent end-to-end, files API, log scan for new errors) passed after the refactor — 21/21 green.


Русская версия

Workspace hosted-агента получил большую уборку, агенты теперь умеют парсить PDF и офисные документы из коробки, а операторы могут поднять лимит hosted-агентов на пользователя без пересборки backend-образа.

Что было сломано

Владельцы наблюдали как файлы пропадают из workspace их hosted-агента. Агент записывал что-то вроде .deep/contacts/note.md или опирался на засеянный при первом запуске .deep/memory/main/MEMORY.md, и при следующей синхронизации файл исчезал из дерева workspace. Многие владельцы решили что агенты молча теряют память между сессиями.

Hosted-агенты не могли парсить PDF, DOCX, XLSX, PPTX и изображения внутри своей песочницы — несмотря на то что базовый фреймворк уже поддерживал соответствующий инструмент-парсер.

Количество hosted-агентов на пользователя было захардкожено на единице. Поднять лимит на self-hosted установке требовало правки исходников и пересборки backend-образа.

Дополнительно проявлялась регрессия на свежеустановленных стендах: каждый новый hosted-агент падал с HTTP 503 на первом же сообщении в чате.

Первопричина

Эндпоинт /files в runner (API листинга workspace, который опрашивает backend) скрывал каждую директорию имя которой начинается с точки. Вместе с логикой prune-on-sync в backend — синхронизатор удаляет любую запись в базе которой нет в последнем листинге runner — это приводило к тому что засеянный .deep/memory/main/MEMORY.md и любые созданные агентом пути под .deep/* затирались при каждой синхронизации. Модель «по умолчанию скрыто» оказалась неверной: по-настоящему шумовыми являются только .deep/checkpoints, .deep/plans, .deep/todos.json, .venv, __pycache__ и node_modules. Всё остальное под .deep/ — содержательное состояние агента.

Инструмент-парсер (liteparse, набор фреймворка для PDF/Office/OCR) поставляется как Python-extra плюс Node-CLI из @llamaindex/liteparse. Образ runner собирался на python:3.12-slim без установленного Node.js и без подтянутого extra liteparse, так что флаг возможности _HAS_LITEPARSE оставался False и набор возвращал пустые результаты во время исполнения.

Регрессия с 503 на первом чате возникла из-за промежуточного исправления, которое начало записывать include_liteparse: true в сохраняемый agent.yaml. Модель DeepAgentSpec фреймворка отклоняет неизвестные ключи (extra='forbid'), поэтому загрузка новой спецификации поднимала Pydantic-ошибку валидации, а эндпоинт чата отдавал её как 503.

Исправление

  1. Раскладка путей workspace: префикс .deep/ убран. Память теперь живёт в memory/MEMORY.md, чекпоинты — в checkpoints/, планы — в plans/, todos — в todos.json. Изменения в agent-runner/main.py и backend/app/services/hosted_agent_service.py. Однократная миграция переписала и базу (agent_files.file_path), и смонтированные workspace на диске через cp -n — ни один агент данные не потерял.

  2. Модель видимости: чёрный список директорий по точке заменён явным набором шумовых имён (.venv, __pycache__, node_modules, .git). Хелпер _is_visible в agent-runner/main.py удалён; обходчик /files теперь возвращает всё остальное. UI workspace показывает то, что агент действительно записал.

  3. Набор liteparse: Dockerfile runner теперь устанавливает Node 20 и выполняет npm install -g @llamaindex/liteparse, а pyproject.toml подтягивает extra pydantic-deep[liteparse]. Runner включает набор передавая include_liteparse=True через overrides в DeepAgent.from_file — этот kwarg уходит во фреймворковую create_deep_agent напрямую, а не в строгую модель спецификации. В agent.yaml флаг больше не пишется, поэтому DeepAgentSpec(extra='forbid') валидируется успешно и регрессия с 503 устранена.

  4. Регулятор лимита hosted-агентов: MAX_HOSTED_AGENTS_PER_USER теперь рантайм-переменная окружения, проброшенная через deploy/docker-compose.prod.yml. Значение по умолчанию остаётся равным 1. Чтобы поднять его на self-hosted установке, задайте значение в .env.prod и перезапустите контейнер backend — пересборка образа не требуется.

Тесты

282 backend pytest, 106 agent-runner pytest. Две прежде существовавшие неудачи в backend, не связанные с этим релизом (test_list_issues_returns_200_not_500, test_commits_endpoint_passes_branch_param), и две прежде существовавшие ошибки коллекции в agent-runner (test_checkpoints.py, test_history_integration.py — Pydantic-ошибка валидации на этапе import) остаются и будут разобраны отдельным тикетом. Сквозная проверка на проде: только что созданный hosted-агент за один ход чата успешно вызвал write_file и liteparse, оба инструмента отработали, итоговый файл появился в дереве workspace, очистка прошла штатно.

Что нужно сделать владельцам

Ничего.

Если хочется поднять лимит hosted-агентов на пользователя на собственной установке, задайте MAX_HOSTED_AGENTS_PER_USER=N в .env.prod и перезапустите контейнер backend.

Дополнительное исправление (коммиты 585d737, 6839e80)

Эндпоинт /chat/stream возвращал tool_calls[].result как пустую строку, даже если инструмент действительно отработал. Когда hosted-агент использует инструмент execute (или любой другой, маршрутизируемый через DeferredToolRequests), потоковый runner разбивает выполнение на два этапа: agent.iter() выдаёт события вызова инструментов, затем отдельный agent.run(deferred_tool_results=...) фактически выполняет их и формирует финальный ответ. Срез new_messages() отложенного запуска содержит только ToolReturnPart (результат инструмента) и итоговый текст — никакого ToolCallPart. Поэтому _extract_response возвращал extra_tools = [], а записи в all_tool_calls так и не получали поле result. Второй сопутствующий дефект: потоковый цикл добавлял дублирующую запись для того же инструмента до того как цикл отложенного выполнения добавлял свою, из-за чего дедупликация при финальном слиянии выбирала непомеченную запись (без result) первой.

Исправление: после каждого отложенного вызова agent.run() выполняется обход result.all_messages() в поисках ToolReturnPart; каждая совпадающая по имени инструмента запись в all_tool_calls получает поле result (проход вперёд, без досрочного выхода). Потоковое событие tool_result теперь тоже отправляется после выполнения и содержит реальный вывод, а не пустую строку.

Стабильность + UX-доработки (коммиты 25b5c59, 6792369, f42d7c5, a519ccb, 412f9d6, 1e6c794)

Вторая волна правок закрывает проблемы, о которых владельцы написали уже после первичного релиза: исчезающее сообщение в чате, лента переписки заваленная heartbeat-шумом, и роутинговый эндпоинт возвращающий серверную ошибку на штатный запрос.

Самый заметный баг: пользователь набирает сообщение, нажимает отправить и через секунду видит, как его собственная реплика пропадает из переписки. Причина была во фронтенде, в загрузчике сообщений — loadMessages вызывал data.reverse() на массиве, полученном из истории чата, но этот массив одновременно лежал в кеше состояния React. Реверс на месте мутировал кеш, и следующий рендер видел пустой или перемешанный список. Исправление заменяет вызов на не-мутирующую копию [...data].reverse() — в frontend/src/app/agents/[id]/chat/page.tsx и в трёх других чат-страницах с тем же паттерном.

Вторая видимая проблема: треды hosted-агентов заваливались строками вида Heartbeat: idle (5m), иногда сотнями в день на каждого агента. Им место в операционном журнале активности, а не в переписке которую видит владелец. Backend записывал их в owner_messages на каждом heartbeat. Исправление маршрутизирует их только в agent_activity и добавляет окно дедупа в 10 секунд на идентичные пользовательские сообщения — двойной клик по кнопке отправки больше не создаёт двух соседних одинаковых строк. Однократная SQL-уборка удалила 226 застрявших heartbeat-строк и 45 дублирующих пользовательских сообщений из таблицы.

Третье: запрос GET /hosted-agents/available-models возвращал HTTP 500 вместо списка моделей. Маршрут был объявлен как /{hosted_id} и жадно матчил любую строку, поэтому available-models трактовался как UUID, парсинг падал ниже по стеку, а в ответе уходил трейсбек. Двадцать четыре path-параметра в роутах hosted-agents теперь объявлены как hosted_id: UUID — Pydantic чисто отклоняет не-UUID-слаги с 422 ещё до того как сработает обработчик.

Более крупный рефакторинг в том же духе: agent_service.heartbeat разросся до 200-строчного метода, в котором смешались сохранение сообщений, журнал активности, синхронизация с runner и метрики. Метод декомпозирован на одиннадцать именованных приватных функций, каждая отвечает за одну фазу; все импорты подняты на верх модуля. Поведение не изменилось — функция теперь читается сверху вниз.

По тестам: интеграционный набор agent-runner для эндпоинтов checkpoint и chat-history падал, потому что тестовый клиент не отправлял заголовок RUNNER_KEY, который теперь обязателен в runner, а патч git-сервиса указывал на старый путь модуля. Оба момента приведены в порядок в conftest.py, набор снова зелёный.

Инфраструктура: исходящий трафик из sandbox-контейнеров до внутренних сервисов платформы раньше проходил через Docker-bridge. Правила firewall на уровне хоста теперь блокируют sandbox-источники при попытке достучаться до внутренних IP-диапазонов на приватных портах сервисов, при этом исходящие вызовы LLM-API в публичный интернет остаются разрешены. Сквозная проверка: sandbox добирается до внешнего провайдера моделей, но больше не может напрямую попасть ни в control-plane runner, ни в базу. Отдельно: шесть тестовых агентов, оставшихся от прежних отладок, и одиннадцать осиротевших каталогов workspace удалены из продуктивной базы и с диска.

Полный набор из 21 продуктивной регрессионной проверки (smoke, auth, регистрация плюс heartbeat, сквозной hosted-agent, files API, проверка логов на новые ошибки) прошёл после рефакторинга — 21 из 21 зелёные.